MC stitching¶
Several backgrounds are generated as a set of overlapping samples — an inclusive one plus samples enriched in a corner of phase space (jet multiplicity, boson pT, a decay channel, a generator-level filter). Stitching combines them into a single process: every event is assigned to exactly one bin, and the events of that bin are normalised with the cross-section and the event count of all samples that contribute to it.
Declaring a stitcher¶
A process declares a stitching processor in its processes.yaml entry:
TT:
processors:
- name: Stitcher
module: FLAF.Processors.MCStitchingTT
class: TTStitcher
config: FLAF/config/Processors/stitching_TT_decayMode.yaml
stages: [ AnaTuple, AnaTupleMerge ]
dependency_level:
AnaTuple: file
AnaTupleMerge: process
genInfo: [ TT ]
datasets:
- TTto2L2Nu
- TTtoLNu2Q
- TTto4Q
The config file lists the bins, each with a selection (a C++ expression) and a
crossSection (an expression over the cross-section database):
The bins must be orthogonal and exhaustive: an event that matches no bin aborts the job,
and an event that matches two is counted twice. When totalCrossSection is given, the sum of
the bin cross-sections is checked against it.
Several datasets, one physics point¶
An extension sample (_ext1) or a variant of the same sample carrying LHE weights adds
statistics to a point that is already there. Both belong to the same process, so each event
must be normalised with the combined event count — normalising each dataset on its own
would count the process twice. A stitcher with no bins does exactly that:
processors:
- name: Stitcher
module: FLAF.Processors.MCStitching
class: MCStitcher
useDatasetCrossSection: true
stages: [ AnaTuple, AnaTupleMerge ]
dependency_level:
AnaTuple: file
AnaTupleMerge: process
A stitcher must be declared at both stages
stages accepts exactly two values, AnaTuple and AnaTupleMerge — nothing else is ever
looked up, so a third name is silently ignored. Both are required: the AnaTuple stage writes
the stitcher's denominator into the anaCache, and the merge stage needs the same processor to
combine those caches again. Declared at AnaTuple alone, every merge of that process fails
with combineAnaCaches: processor <name> not provided for combining anaCaches.
dependency_level is read only for AnaTupleMerge, where process makes the merge wait for
the whole process — which stitching needs, since the denominator spans its datasets.
useDatasetCrossSection takes the cross-section from the dataset entry (both datasets point
at the same one) instead of a bin configuration, and the denominator is summed over every
dataset of the process. A point that has a single dataset is unaffected: the sum is over that
one dataset, which is what it was normalised with before.
Variables the bins select on¶
LHE_Vpt, LHE_NpNLO and friends are nanoAOD branches that the anaTuple keeps, so a bin can
select on them directly. Anything else — a gen-level decay mode, a dilepton mass, the
generator filter of a sample — has to be derived from GenPart/LHEPart, and those
collections are not part of the anaTuple.
This matters because the bin selections are evaluated twice:
| Stage | Frame | Purpose |
|---|---|---|
AnaTupleFileTask |
nanoAOD | sum the denominator of each bin into the report |
AnaTupleMergeTask |
merged anaTuple | give each event the cross-section and denominator of its bin |
The merge stage no longer has GenPart/LHEPart, so a derived variable must be stored in
the anaTuple by the analysis. Which gen-level information a process needs is declared as
genInfo next to its processors, and the analysis-level anaTuple definition turns that into
branches (DYInfo_*, TauTauInfo_*, TTInfo_*). A stitcher then prefers the stored branch
and falls back to the nanoAOD collections only when it is not there:
class TTStitcher(MCStitcher):
def defineVariables(self, df):
df = defineFromStoredOrExpression(
df,
"TT_n_leptonic_W",
stored="TTInfo_nLeptonicW",
expression="gen_process::tt::identify(...).nLeptonicW()",
prepare=_prepare,
)
return super().defineVariables(df)
Both stages therefore see the same value, computed once from the nanoAOD.
Adding a variable to a bin selection changes the anaTuple
A bin that selects on a new derived variable needs that variable stored, which means the
affected samples have to be produced again. A stitcher whose variable is missing from the
anaTuple fails in AnaTupleMergeTask with use of undeclared identifier 'GenPart_…'.
Adding a stitcher¶
- Put the gen-level identification in a self-contained header under
FLAF/include/GenProcess/and give it a test inFLAF/test/GenProcess/. - Subclass
MCStitcherand overridedefineVariableswithdefineFromStoredOrExpression, naming the branch the analysis stores. - Store that branch in the analysis anaTuple definition and declare the corresponding
genInfofor the process. - Cover it in the integration test. Every analysis runs two CI backgrounds — one t̄t and one
DY dataset — and each carries the same
processors:andgenInfo:as the analysis's realTTand DY process for that era, so the stitcher runs over the whole anaTuple → merge → histogram chain, which is where a missing branch shows up. If your stitcher replaces one of those, update the matching CI process; if it stitches something else, give that process the same treatment. One dataset per process is enough — each bin's denominator is summed over the very events that later read it back, so a bin no event falls into is never divided by.