# Seg faults in coadd task

**URL:** <https://www.rubin.community/t/seg-faults-in-coadd-task/8074>\
**Category:** Support\
**Created:** [November 15, 2023, 1:56am UTC](https://www.rubin.community/t/seg-faults-in-coadd-task/8074 "2023-11-15T01:56:45Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![tdboer](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/tdboer/32/2019_2.png) [@tdboer](https://www.rubin.community/u/tdboer)\
**Post date:** [November 15, 2023, 1:56am UTC](https://www.rubin.community/t/seg-faults-in-coadd-task/8074/1 "2023-11-15T01:56:45Z")

</div>

I am running into a frustrating issue while stacking together archive HyperSuprimeCam images for a piece of sky I am looking at.

For some small fraction of my coadd runs I get a bad termination error, which seg faults with exit code 139. In most cases, these stacks have four input visits, one of which is larger than the others (as in, one visit covers pretty much the whole tile, and the other 3 lesser fractions). If I remove the larger visit from the coadd input, things are working just fine.

I initially thought this might be a memory related issue, but I have run plenty of other deeper stacks with more input visits, and also still larger input fractions (as in, I have run stacks with tens of input visits, several of which cover the full sky tile, on the same machine, with the same setup, without issue).

Is there any way to get more info out of the runs, and/or run with some higher verbosity or debug, to try to get more of a handle on why this is happening. I have visually inspected the rogue piece of problem image, and found nothing strange that jumps out at me, so running out of things to try.

relevant log snippet

```auto
coaddDriver.assembleCoadd.detectTemplate INFO: Detected 26588 positive peaks in 11404 footprints to 5 sigma
coaddDriver.assembleCoadd.detectTemplate INFO: Detected 26588 positive peaks in 11404 footprints to 5 sigma
coaddDriver.assembleCoadd.scaleWarpVariance INFO: Renormalizing variance by 0.988779
coaddDriver.assembleCoadd.scaleWarpVariance INFO: Renormalizing variance by 0.988779
coaddDriver.assembleCoadd.detect INFO: Detected 7686 positive peaks in 946 footprints and 3301 negative peaks in 794 footprints to 5 sigma
coaddDriver.assembleCoadd.detect INFO: Detected 7686 positive peaks in 946 footprints and 3301 negative peaks in 794 footprints to 5 sigma
coaddDriver.assembleCoadd.scaleWarpVariance INFO: Renormalizing variance by 1.119721
coaddDriver.assembleCoadd.scaleWarpVariance INFO: Renormalizing variance by 1.119721
coaddDriver.assembleCoadd.detect INFO: Detected 7124 positive peaks in 1671 footprints and 3707 negative peaks in 1122 footprints to 5 sigma
coaddDriver.assembleCoadd.detect INFO: Detected 7124 positive peaks in 1671 footprints and 3707 negative peaks in 1122 footprints to 5 sigma

===================================================================================
= BAD TERMINATION OF ONE OF YOUR APPLICATION PROCESSES
= PID 23716 RUNNING AT ippc134
= EXIT CODE: 139
= CLEANING UP REMAINING PROCESSES
= YOU CAN IGNORE THE BELOW CLEANUP MESSAGES
===================================================================================
YOUR APPLICATION TERMINATED WITH THE EXIT STRING: Segmentation fault (signal 11)
This typically refers to a problem with your application.
Please see the FAQ page for debugging suggestions

```

---

<div class="post-metadata">

**Author:** ![timj](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/timj/32/10_2.png) [@timj](https://www.rubin.community/u/timj)\
**Post date:** [November 15, 2023, 3:57pm UTC](https://www.rubin.community/t/seg-faults-in-coadd-task/8074/2 "2023-11-15T15:57:20Z")

</div>

> [@tdboer](#):
>
> I am running into a frustrating issue while stacking together archive

What version of the software are you running? The reference to `coaddDriver` implies to me that you are running the old gen2 version of the pipelines software.

---

<div class="post-metadata">

**Author:** ![tdboer](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/tdboer/32/2019_2.png) [@tdboer](https://www.rubin.community/u/tdboer)\
**Post date:** [November 15, 2023, 5:11pm UTC](https://www.rubin.community/t/seg-faults-in-coadd-task/8074/3 "2023-11-15T17:11:28Z")

</div>

Indeed this was still using the gen2 version from v23, since this particular project was essentially pre-setup uding the old pipeline a few years back.

---

<div class="post-metadata">

**Author:** ![timj](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/timj/32/10_2.png) [@timj](https://www.rubin.community/u/timj)\
**Post date:** [November 15, 2023, 6:00pm UTC](https://www.rubin.community/t/seg-faults-in-coadd-task/8074/4 "2023-11-15T18:00:04Z")

</div>

> [@tdboer](#):
>
> gen2 version from v23

If you had a stack trace we might be able to give you some clues (and maybe point at a ticket where we fixed the problem) but at this point we are not planning to make any new v23 releases. If you get the segv in v26 using the gen3 pipelines then that’s a different story.

---

<div class="post-metadata">

**Author:** ![timj](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/timj/32/10_2.png) [@timj](https://www.rubin.community/u/timj)\
**Post date:** [November 15, 2023, 6:15pm UTC](https://www.rubin.community/t/seg-faults-in-coadd-task/8074/5 "2023-11-15T18:15:04Z")

</div>

> [@tdboer](#):
>
> with some higher verbosity or debug

The old documentation might help with log levels:

> **[Logging with command-line tasks — LSST Science Pipelines](https://pipelines.lsst.io/v/v23_0_0/modules/lsst.pipe.base/command-line-task-logging-howto.html)**

---

<div class="post-metadata">

**Author:** ![price](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/price/32/488_2.png) [@price](https://www.rubin.community/u/price)\
**Post date:** [November 15, 2023, 9:49pm UTC](https://www.rubin.community/t/seg-faults-in-coadd-task/8074/6 "2023-11-15T21:49:12Z")

</div>

The usual way forward from here is to guess which patch it was operating on at the time and run `coaddDriver.py` on just that patch with `--batch-type=none` under `gdb`, and get a stack trace when (if) it segfaults. If you can’t figure out which patch it is, you could run them all under that mode, but it would run serially.

Because LSST v23 was cut in the process of ditching the Gen2 middleware, you might do better rewinding a bit further and using hscPipe 8.5.3, where the Gen2 middleware has been extensively tested. It’s been a long time since I’ve tried this, but the installation instructions are:

```auto
    wget https://tigress-web.princeton.edu/~HSC/hscPipe8/newinstall.sh
    bash newinstall.sh
    # Answer "yes" to installing Anaconda unless you really know what you're doing
    # Source the appropriate file it tells you to, and then proceed to the next step
    
    # Install this release:
    eups distrib install hscPipe 8.5.3

```

---

<div class="post-metadata">

**Author:** ![tdboer](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/tdboer/32/2019_2.png) [@tdboer](https://www.rubin.community/u/tdboer)\
**Post date:** [November 16, 2023, 12:11am UTC](https://www.rubin.community/t/seg-faults-in-coadd-task/8074/7 "2023-11-16T00:11:19Z")

</div>

I certainly cant fault that. This is a pretty old project, so if it wasn’t nearly done I would try to migrate to gen3.

---

<div class="post-metadata">

**Author:** ![tdboer](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/tdboer/32/2019_2.png) [@tdboer](https://www.rubin.community/u/tdboer)\
**Post date:** [November 16, 2023, 12:11am UTC](https://www.rubin.community/t/seg-faults-in-coadd-task/8074/8 "2023-11-16T00:11:48Z")

</div>

Somehow I missed that page. I will see if anything pops up when running with better logging turned on. Thanks!
