# Topics for Butler conversations at the AHM

**URL:** <https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018>\
**Category:** Data Management\
**Tags:** butler\
**Created:** [August 12, 2016, 12:06am UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018 "2016-08-12T00:06:58Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![natepease](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/natepease/32/2086_2.png) [@natepease](https://www.rubin.community/u/natepease)\
**Post date:** [August 12, 2016, 12:06am UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/1 "2016-08-12T00:06:58Z")

</div>

A few of us have been discussing Butler stuff recently and have been putting off some conversations until the all hands meeting, hoping that we can find a moment to have an ad hoc face to face conversation and whiteboard stuff as needed.

To keep track of these topics, I created a list at [https://confluence.lsstcorp.org/display/DM/Butler+Topics+for+AHM+2016](https://confluence.lsstcorp.org/display/DM/Butler+Topics+for+AHM+2016)  
If you’re interested in any of these in particular, add your name under the topic and I’ll include you when we work on finding a time to meet.

If you’d like to add a topic, I think it will work if you just add it to your page (don’t forget to include your name under the topic).

---

<div class="post-metadata">

**Author:** ![FabioHernandez](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/fabiohernandez/32/82_2.png) [@FabioHernandez](https://www.rubin.community/u/FabioHernandez)\
**Post date:** [August 12, 2016, 7:37am UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/2 "2016-08-12T07:37:36Z")

</div>

I would like to understand the mechanisms currently provided by the Butler for using non-POSIX storage backends. Specifically, I’m interested in exploring the suitability of object stores (such as OpenStack Swift or Amazon S3-compatible) as repositories for LSST data.

I wonder if this topic will be addressed in one of the breakout sessions. If that’s the case, I would like to attend.

---

<div class="post-metadata">

**Author:** ![natepease](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/natepease/32/2086_2.png) [@natepease](https://www.rubin.community/u/natepease)\
**Post date:** [August 12, 2016, 3:45pm UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/3 "2016-08-12T15:45:11Z")

</div>

@FabioHernandez let’s make time to discuss it. I don’t think there will be dedicated sessions set up for these topics (@gpdf?), but we could use some of the time during the pairwise discussions, and/or meet informally any time during the week. If anyone else is interested (let me know), we can figure out a way to set up a more formal time.

I do remember the conversations we had re. using S3 servers at last year’s meeting in Bremerton.

There’s not any mechanism other than POSIX yet, but I’ve been thinking about it and talking about it a lot with Fritz & KT. It might be a helpful primer for you to read through the document at [https://confluence.lsstcorp.org/display/DM/Butler+Storage+and+Format+Refactor](https://confluence.lsstcorp.org/display/DM/Butler+Storage+and+Format+Refactor)

---

<div class="post-metadata">

**Author:** ![mtpatter](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/mtpatter/32/335_2.png) [@mtpatter](https://www.rubin.community/u/mtpatter)\
**Post date:** [August 12, 2016, 6:27pm UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/4 "2016-08-12T18:27:15Z")

</div>

Hi there. I have a little bit of experience from a previous project operating S3-compatible storage at the near petabyte scale (specifically, Ceph’s radosgw) as a data repository and would be glad to share knowledge/experience if that is helpful.

---

<div class="post-metadata">

**Author:** ![KSK](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/ksk/32/393_2.png) [@KSK](https://www.rubin.community/u/KSK)\
**Post date:** [August 12, 2016, 11:50pm UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/5 "2016-08-12T23:50:35Z")

</div>

Just to add to the non-POSIX conversation, there is a person here at UW trying to run the command line tasks on Spark and she ended up un-butling the tasks. I would be interested in knowing whether a backend suitable for Spark is even a reasonable thing to ask for. Note that I know next to nothing about how processing in Spark works.

---

<div class="post-metadata">

**Author:** ![natepease](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/natepease/32/2086_2.png) [@natepease](https://www.rubin.community/u/natepease)\
**Post date:** [August 13, 2016, 4:25pm UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/6 "2016-08-13T16:25:23Z")

</div>

> [@mtpatter](#):
>
> Hi there. I have a little bit of experience from a previous project operating S3-compatible storage at the near petabyte scale (specifically, Ceph’s radosgw) as a data repository and would be glad to share knowledge/experience if that is helpful.

Hey Maria, I think it would be very good hear about that. Will you be at the meeting all week?

---

<div class="post-metadata">

**Author:** ![natepease](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/natepease/32/2086_2.png) [@natepease](https://www.rubin.community/u/natepease)\
**Post date:** [August 13, 2016, 4:30pm UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/7 "2016-08-13T16:30:29Z")

</div>

> [@KSK](#):
>
> Just to add to the non-POSIX conversation, there is a person here at UW trying to run the command line tasks on Spark and she ended up un-butling the tasks. I would be interested in knowing whether a backend suitable for Spark is even a reasonable thing to ask for. Note that I know next to nothing about how processing in Spark works.

I’m only wikipedia-familiar with spark. I’d like to discuss this more. It feels a little funny, like an abstraction layer on top of an abstraction layer(?). But I guess using it to abstract a distributed dataset across multiple machines could be useful? Are you and/or that person available at the meeting this week? (I suppose you’ll be busy with the review through wednesday or thursday…)

---

<div class="post-metadata">

**Author:** ![ktl](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/ktl/32/1373_2.png) [@ktl](https://www.rubin.community/u/ktl)\
**Post date:** [August 13, 2016, 4:38pm UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/8 "2016-08-13T16:38:00Z")

</div>

My understanding is that Spark is intended to behave as if it is working on in-memory data, with its own data distribution and task assignment. In that case, there seem to me to be two alternatives: either have an essentially trivial Butler layer that just retrieves whatever data Spark already has, or change the model and use the Butler only to load data into Spark to begin with.

---

<div class="post-metadata">

**Author:** ![FabioHernandez](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/fabiohernandez/32/82_2.png) [@FabioHernandez](https://www.rubin.community/u/FabioHernandez)\
**Post date:** [August 13, 2016, 4:42pm UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/9 "2016-08-13T16:42:32Z")

</div>

> [@natepease](#):
>
> It might be a helpful primer for you to read through the document at [https://confluence.lsstcorp.org/display/DM/Butler+Storage+and+Format+Refactor](https://confluence.lsstcorp.org/display/DM/Butler+Storage+and+Format+Refactor)

Thanks for pointing me to the document. I will get familiar with it for next week’s conversation.

---

<div class="post-metadata">

**Author:** ![price](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/price/32/488_2.png) [@price](https://www.rubin.community/u/price)\
**Post date:** [August 13, 2016, 4:52pm UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/10 "2016-08-13T16:52:29Z")

</div>

> [@ktl](#):
>
> In that case, there seem to me to be two alternatives: either have an essentially trivial Butler layer that just retrieves whatever data Spark already has, or change the model and use the Butler only to load data into Spark to begin with.

Third option: use Spark as a parallelisation engine only. In that case, Spark would distribute `dataId`s and temporary products, and you would use the `Butler` in `Task`s for I/O the same way we always have — no need to change anything except the glue in-between the top-level `CmdLineTask`s.

---

<div class="post-metadata">

**Author:** ![mtpatter](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/mtpatter/32/335_2.png) [@mtpatter](https://www.rubin.community/u/mtpatter)\
**Post date:** [August 13, 2016, 8:35pm UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/11 "2016-08-13T20:35:31Z")

</div>

Yes, will be there all week.

---

<div class="post-metadata">

**Author:** ![ktl](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/ktl/32/1373_2.png) [@ktl](https://www.rubin.community/u/ktl)\
**Post date:** [August 15, 2016, 2:43am UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/12 "2016-08-15T02:43:25Z")

</div>

> [@price](#):
>
> Third option: use Spark as a parallelisation engine only.

Indeed that’s an option and was how I initially thought about it. It’s not clear to me how much we can do with temporary products in Spark without modifying the current CmdLineTasks, and I worry that much of the benefit of using Spark might go away in this mode.

---

<div class="post-metadata">

**Author:** ![KSK](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/ksk/32/393_2.png) [@KSK](https://www.rubin.community/u/KSK)\
**Post date:** [August 15, 2016, 6:20pm UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/13 "2016-08-15T18:20:44Z")

</div>

> [@ktl](#):
>
> I worry that much of the benefit of using Spark might go away in this mode.

This was one of my concerns as well.

Another thing that came up in this space is whether we should try harder to allow construction of objects from streams. Currently we don’t allow construction from byte streams for any of our low level objects. I don’t know if that was conscious decision or just an artifact of the capabilities of cfitsio.

---

<div class="post-metadata">

**Author:** ![ktl](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/ktl/32/1373_2.png) [@ktl](https://www.rubin.community/u/ktl)\
**Post date:** [August 16, 2016, 1:54am UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/14 "2016-08-16T01:54:57Z")

</div>

> [@KSK](#):
>
> construction of objects from streams

This is mostly a cfitsio issue. I think it will most likely go away (I can think of a couple of different ways to make it do so, at varying efficiencies) when we allow back-ends like S3.

---

<div class="post-metadata">

**Author:** ![natepease](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/natepease/32/2086_2.png) [@natepease](https://www.rubin.community/u/natepease)\
**Post date:** [August 18, 2016, 9:58pm UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/15 "2016-08-18T21:58:51Z")

</div>

> [@mtpatter](#):
>
> Hi there. I have a little bit of experience from a previous project operating S3-compatible storage at the near petabyte scale (specifically, Ceph’s radosgw) as a data repository and would be glad to share knowledge/experience if that is helpful.

Hey Maria, are you available to meet today after the last session, at 5 or a little after? Tomorrow after 11 would also work for me (my shuttle to the airport leaves around 4)

---

<div class="post-metadata">

**Author:** ![natepease](https://sea2.discourse-cdn.com/flex002/user_avatar/www.rubin.community/natepease/32/2086_2.png) [@natepease](https://www.rubin.community/u/natepease)\
**Post date:** [August 18, 2016, 10:02pm UTC](https://www.rubin.community/t/topics-for-butler-conversations-at-the-ahm/1018/16 "2016-08-18T22:02:01Z")

</div>

FYI @FabioHernandez and I are planning to test butler using a repository in Swift (IE not on the local filesystem). The working plan is [on confluence](https://confluence.lsstcorp.org/display/DM/Swift+Butler+Storage+Trial).
