Processes¶
Everything process-specific in DSProd lives in a process customization module. The tasks stay generic; a module teaches DSProd how to turn a compact process configuration into concrete production points, how to obtain each point's gridpack, and how to render the CMSSW gen fragment.
The interface — the ProcessCustomization abstract base class — lives in DSProd itself
(dsprod/processes/base.py). The models (plugins + cards + fragments) live in a separate
submodule, DSProdModels, mounted at models/.
That tree is content, not a Python package — it has no __init__.py anywhere; DSProd loads each
plugin.py straight from its path. The reference model is models/X_HH/ — resonant X→HH, with
the bbWW final states supplied as gen fragments.
The interface¶
A process module subclasses ProcessCustomization, sets a unique name, and implements three
required methods:
from dsprod.registry import register_process
from dsprod.processes.base import ProcessCustomization, Point, GridpackSpec
@register_process
class MyProcess(ProcessCustomization):
name = "My_Process" # the prod_setup `process:` key
def enumerate_points(self, process_cfg: dict) -> list[Point]:
"""Expand the setup's `points:` (e.g. a mass scan) into concrete Point objects."""
def gridpack(self, point: Point, era: str) -> GridpackSpec:
"""Return how to *generate* the gridpack (used only when it is not already stored)."""
def gen_fragment(self, point: Point, era: str) -> str:
"""Render the CMSSW gen fragment for this point/era and return its path."""
The @register_process decorator registers the class under its name. get_process(name)
walks the models tree, loads every plugin.py, and then resolves the name from any
production setup.
Gridpacks: locate-or-generate¶
A point never names its gridpack. Its canonical location in the
DSProdGridpacks store comes from
gridpack_rel_path(point): ImportGridpack copies it to fs_default if
the store has it, and MakeGridpack generates it otherwise. So the plugin provides two things:
- where the gridpack lives — override
gridpack_rel_path(point, era)to return the path (relative to thegridpacksstore) mirroring the model's own directory convention. Its first level names the hard process the gridpack contains — which is what the model is named after too: if the decays are applied by the gen fragment, one gridpack serves every final state, soX_HHstores underX_HH/and shares the tarball across all of them; - how to generate it —
gridpack()returns aGridpackSpec(generator="MadGraph5_aMCatNLO", cards_template=...), andrender_gridpack_cards(point, out_dir)writes thegenproductionsinput cards (proc_card/run_card/customizecards/extramodels) intoout_dirand returns the processNAME.
Optional overrides¶
point_name, gridpack_name, gridpack_rel_path, xsec, filter_efficiency, and validate
have sensible defaults and can be overridden when a process needs them.
A plugin lives in the models submodule and imports DSProd's framework classes, so it is
usable only inside a DSProd checkout (not as a standalone library). It resolves its cards and
fragment relative to its own location (os.path.dirname(__file__)), so a model is self-contained.
Model layout¶
Inside models, models are organized by process → generator → center-of-mass energy,
with the energy as the last level so the plugin and the process/generator tooling above it are
shared across energies. For X_HH:
models/ (the DSProdModels submodule, mounted at models/)
└── X_HH/ # process (matches the plugin name)
├── README.md # process docs + links to the original sources
├── filters/ # (optional) final-state filters, shared across generators/energies
└── MadGraph5_aMCatNLO/ # generator (matches genproductions_scripts/bin + GridpackSpec.generator)
├── plugin.py # the ProcessCustomization, @register_process (shared across energies)
├── scripts/ # (optional) prodcard-generation scripts (e.g. parametrized in mX)
├── models/ # (optional) custom generator models, when not centrally available
└── 13p6TeV/ # center-of-mass energy (LAST level)
├── cards/ # genproductions cards, one directory per production mode
│ └── <production_mode>/ # proc_card, run_card, customizecards, extramodels
└── fragments/ # CMSSW gen fragments — one per final state
Model directories and the names inside them follow the DAS tokens of the central samples the
model reproduces (GluGlutoRadiontoHHto2B2Vto2B2JLNu_M-800 → production mode GluGlutoRadion,
final state 2B2JLNu), so a setup point reads like the dataset it produces.
Only plugin.py, cards/, a gen fragment, and the READMEs are required; filters/, scripts/,
and models/ appear only when a model needs them. The plugin resolves its energy-specific inputs
via com_energy(era) and fills in the per-point values (mass, …) when it renders the cards.
DSProd walks this tree and loads every plugin.py, so a model becomes available simply by
adding its directory — there is no central registration list, and no __init__.py, to edit.
Adding a process¶
In the DSProdModels repository:
- Create
<process>/<generator>/with aplugin.py— aProcessCustomizationsubclass decorated with@register_processand a uniquename. - Add a
<process>/<generator>/<comEnergy>/folder with thecards/and thefragments/for each energy you produce, and document the process in<process>/README.md.
Then in DSProd, write a production setup with process: <name> and its
points:, and advance the models submodule pointer. No task code changes — the graph is
generic.