Running the analysis¶
The pipeline itself — producing anaTuples, computing observables, filling and merging histograms,
plotting — is the standard FLAF chain. Follow the
FLAF full-workflow walkthrough for the
commands (InputFileTask → AnaTupleFileTask → AnaTupleMergeTask → HistTupleProducerTask →
HistFromNtupleProducerTask → HistMergerTask → HistPlotTask). This page collects the points
that are specific to HH→bb̄ττ.
Always pick the DeepTau version¶
Every stage that depends on τ identification must agree on the DeepTau version. Pass it consistently across the whole production:
ERA=Run3_2022
VER=v1_deepTau2p5 # version name encodes the DeepTau version (see Setup)
law run FLAF.Analysis.tasks.HistPlotTask \
--period $ERA --version $VER --workflow local \
--customisations deepTauVersion=2p5
If you omit deepTauVersion, the default (2p1) is used. Keep the --version name and the
deepTauVersion customisation in sync to avoid mixing productions.
Channels¶
HH→bb̄ττ is analysed in the three τ-pair channels — eτ, μτ and ττ. Channel selection is
driven by the analysis configuration (config/global.yaml); restrict or extend the channels there
or via your user_custom.yaml.
Choosing which variables to histogram¶
The set of variables produced by HistFromNtupleProducerTask/HistPlotTask is controlled by the
variables: list in your user_custom.yaml (or the --variables argument):
Some observables are computed in the cache step (AnalysisCacheTask, e.g. the
LegacyVariables/heavier quantities) rather than directly from the anaTuple. When you request such a
variable, LAW pulls in the cache task automatically — see
FLAF → Task reference. Listing
a short variables: set is the easiest way to keep test runs fast.
Stitched backgrounds: DY¶
DY is stitched from several samples, so each event is normalised with the cross-section of the bin it belongs to (MC stitching). The bins select on gen-level quantities that nanoAOD does not provide directly, so the anaTuple stores them:
| Branch | Stored for | Meaning |
|---|---|---|
DYInfo_flavor, DYInfo_mll |
DYto2Tau_M_50 |
flavour (11/13/15) and mass of the LHE dilepton pair |
TauTauInfo_passFilter |
DYto2Tau_M_50 |
the Z→ττ generator filter decision, the axis that stitches the filtered samples in |
TauTauInfo_vis_type{1,2}, TauTauInfo_vis_pt{1,2}, TauTauInfo_vis_abseta{1,2} |
DYto2Tau_M_50 |
the visible tau quantities the filter is made of, so its definition can be revisited without reprocessing |
TTInfo_nLeptonicW, TTInfo_wDecay{1,2} |
TT |
gen-level t̄t decay channel |
genTop_{pt,eta,phi,mass} |
TT |
last-copy top and anti-top, in this order; read by the top-pT reweighting |
genTop_b_{pt,eta,phi}, genTop_lep_{pt,eta,phi,mass}, genTop_lep_gen_kind |
TT |
the b quark and the W's charged lepton of each top, and that lepton's GenLepton::Kind (-1 for a hadronic W) |
Which of these a process gets is declared as genInfo next to its processors in
config/<era>/processes.yaml, where a stitcher selects on it or, for TT, where the
top-pT reweighting reads genTop_pt (every era) — adding a kind to a process that is
already produced means producing it again; AnaProd/genProcessInfo.py turns that into the
branches above. The genTop_* arrays are defined before the event selection
(defineGenVariables in AnaProd/anaTupleDef.py), because the reweighting is a shape weight whose
denominator sums over all events. TTInfo_* columns must stay scalars: the anaTuple stores all
columns that share a prefix as one collection. A process that stitches on one of these quantities without declaring genInfo fails
in AnaTupleMergeTask, where GenPart/LHEPart are no longer available.
t̄t is not stitched: TT takes only the three decay-channel samples (TTto2L2Nu, TTtoLNu2Q,
TTto4Q) in every era, which do not overlap. The inclusive TT/TT_ext1 samples of Run3_2022 and
Run3_2023BPix are not used: they carry no parton-shower weights (PSWeight has a single entry), on
which the parton-shower weight producer stops.
The integration test guards this. TestModel runs two backgrounds — custom_CI_Background_TT,
one t̄t dataset, and custom_CI_Background_DY, one DY→ττ dataset — and each carries the same
processors: and genInfo: as the real TT and DYto2Tau_M_50 process for that era
(DYtautauStitcher for 2022–2023BPix, the plain MCStitcher for 2024 onwards, which is what
those eras configure). The stitchers therefore
run over the whole anaTuple → merge → histogram chain in CI, which is exactly where a missing
gen-level branch shows up. Change one of the real processes and change its CI counterpart with
it.
Theory weights stored in the anaTuple¶
All MC is produced with the parton-shower ISR/FSR (PSWeight), PDF (LHEPdfWeight, 103 members)
and QCD-scale (LHEScaleWeight) weights, and t̄t with the top-pT reweighting
(genTop_pt, data_nlo). Each is a shape weight staged like the pileup weight: computed at
AnaTupleFileTask, where its inclusive sum enters the denominators, and turned into
weight_base_<variation>_rel at AnaTupleMergeTask. The merged anaTuple keeps weight_base and the
weight_base_*_rel branches and drops the per-member weights (anaTupleMerge_drop_columns in
config/global.yaml). They are not used in weights.yaml or the datacards yet; adding them there
needs no new anaTuple production.
config/Run3_2024/global.yaml (and 2025, 2026) inherits the analysis-wide corrections: block
through the anchor *corrections_default and overrides only btag (UParTAK4, no shape
calibration) and dy_hhbbtautau, so these weights, and any correction added to
config/global.yaml, apply to those eras too.
Columns taken from the central tree¶
config/global.yaml lists in anaTuple_shift_invariant_columns the anaTuple columns that no
systematic shift changes: event numbers and dataset metadata, generator, luminosity and
cross-section weights, the pileup and theory weights, pileup truth, the LHE record, gen jets, the
stitching information (DYInfo_*, TauTauInfo_*, TTInfo_*), genTop_* and the gen-level signal
leptons (genLepton{1,2}_*). FLAF stores them in
the central tree only and fills them in for events that only a shift (JES, JER, tau or lepton
energy scale) selected, after checking that every variation agrees, so such events enter the
shifted templates with their real weight_base instead of 0.
- Only event-level generator quantities belong there. Generator information attached to a
reconstructed object —
tau*_gen_*, jet flavour labels,nJetFromGenHbb— follows the selected object and changes under the shifts; the fuse step stops withColumn '…' is declared shift-invariant but differs in …if one is listed. - Every input of
weight_base(generator, luminosity, cross-section and shape weights, the stitching variables) must be listed, or an event selected only by a shift gets weight 0 there. - A shifted tree reaches these columns through its
Centralfriend, whichHasColumnsees andGetColumnNames()does not list. A column that code downstream looks up withGetColumnNames()must stay off the list: the generator boson of the bosonic recoil correction (recoil_GenBoson_*) is invariant but is not listed for that reason. - Changing the list changes the anaTuple layout, so it goes with a new anaTuple production.
Signal points with more than one sample¶
Some GluGlutoHH points are produced more than once — an _ext1 extension, or a variant
carrying LHE weights. Both are declared and both are used, so the point gets all the
statistics; the GluGlutoHHto2B2Tau process therefore carries the *ext_processors
stitcher, which normalises the point with the summed event count instead of counting it
twice (MC stitching). Which samples
exist differs per era — Run3_2022EE, for instance, has only the LHE-weighted variant of
kl = 2.45 and no plain one.
Quick stack plots¶
For a fast look at distributions (outside the full HistPlotTask styling), the analysis ships a
helper script. Edit the paths/variable names at the top to match your run, then:
Publishing plots¶
To share plots through a personal interactive web browser, see Interactive plot browser.
Statistical interpretation¶
Once histograms exist, continue to Statistical inference for datacards, limits and diagnostics.