Custom RDataSource without (or with limited) multithreading

Dear ROOTers,

I want to learn how to implement a custom RDataSource, using the real cases in which I am already using the trick of creating empty RDataFrame and populating column with a series of Define().

For me it is easier to start without supporting multi-threading, since this is tricky from the point of view of the real, “back-end” sources.

The first question is on the method GetEntryRanges() to be implemented. I guess that I do not need to distribute always evenly the entries among slots, so I can bring this to the extreme and assign all the entries to the first slot and none to the others. How do the RDataFrame slot engine react to an empty list? Is it calling anyway any of the other slot methods?

Then a question on GetColumnReaders(), which I am implementing it on the basis what you call “new API” (unfortunately the only example is the quite complex RNTuple). I understand that this is called for each slot, so each thread will have its own instance of RColumnReaderBase-derived class. This means that if the low-level data buffers are owned and managed by this class instead of the overall custom RDataSource (it seems that I can choose both options), these are replicated for each thread avoiding (maybe) data contention.

Some clarification of these two points would be really useful.

Thanks,
Matteo


The question is quite general, but it should be supported by:
Built for linuxx8664gcc on Mar 15 2025, 21:37:19
From tags/v6-34-04@v6-34-04
With c++ (Debian 12.2.0-14) 12.2.0


Hi Matteo,
Thank you for your question.
@vpadulan should be able to help.
Best,
Lukas

Dear @malfonsi79 ,

Thanks for reaching out and for taking interest in RDataFrame. Creating a custom RDataSource is totally supported, albeit as you mention the interface has been undergoing some changes that are not thoroughly documented yet because of their somewhat new nature. Given that the request for creating custom RDataSource is also not very high, I apologise if the experience hasn’t been smooth.

If I understand your question correctly, you are currently not foreseeing the support of multi-threading. As such, your GetEntryRanges method should just return the entire range of entries you want to process, given that it will be anyway processed by a single thread and so there is only one RDataFrame slot. Every other consideration about splitting the input dataset in chunks only really makes sense when considering multi-threading.

About GetColumnReaders, I guess you’re trying to implement the overload std::unique_ptr<ROOT::Detail::RDF::RColumnReaderBase> GetColumnReaders. That’s good, but here comes a bit of that interface change that I was referring to. I have on my task list to change the name of this method which I find a bit confusing, perhaps to just CreateColumnReader since the return value is one column reader only. For the rest, your description is correct: the various threads may call GetColumnReaders independently and then you can choose how to implement your column reader to different degrees of the dependency from its data source.

Since you’re embarking on this task, I would find it useful to discuss some details with you which could also help in developing better documentation. Let me know if you would be available for example to schedule a call.

Cheers,
Vincenzo

Dear Vincenzo,

thanks a lot for the offer, which I am very happy to take! I have just written a message to you.

Ciao,
Matteo