8.8.8. ModeSelector#
ModeSelector is a neural network classifier for hadronic FEI \(B\) meson
tagging. The FEI assigns a signal probability (sigProb) to each \(B\)
candidate independently, using only that candidate’s own decay products.
ModeSelector instead uses the full set of FEI \(B\) candidates in the event
simultaneously (all candidates with sigProb > 0.001) to predict which
physics category the event belongs to and which specific FEI decay mode most
likely produced the tag \(B\).
When to use ModeSelector#
ModeSelector builds directly on the FEI: it was trained on FEI skims and requires them as input.
The selection on the ModeSelector output then supersedes any additional sigProb cut.
Its main strength is separating \(B^0\) and \(B^+\) events, which substantially reduces crossfeed in analyses that do not fully reconstruct the signal side.
The gain in performance depends on the signal mode: analyses with more inclusive signal sides benefit the most, because the less constrained the signal reconstruction is, the more sensitive it is to crossfeed from the wrong \(B\) type.
Algorithm#
Two-stage network#
ModeSelector runs two neural networks in sequence on each event.
All B candidates -> [feature matrix] -> category network -> B0 / B+ / continuum
|
v
main network -> 139-class output -> BplusScore
The category network (3 outputs) classifies the event as \(B^0\), \(B^+\), or continuum based on the full set of candidates.
The main network (139 outputs) then predicts which specific FEI decay mode produced the tag \(B\). The 139 classes are:
classes 0-135: signal modes, where the class index (
input_id) directly encodes the \(B\) type, decay mode indexdmID, and particle/antiparticle sign. Thus the 136 classes are (32 neutral FEI modes + 36 charged FEI modes) \(\times 2\).class 136:
bad_tag(combinatorial background)class 137:
cross_deltaC1(\(BB\) crossfeed from the wrong \(B\) charge)class 138:
continuum
Each input_id slot can in principle be occupied by multiple FEI
candidates (same \(B\) type, decay mode, and charge). When this happens,
only the candidate with the highest sigProb is kept per slot
(deduplication). The input feature matrix has 9 blocks of per-candidate kinematic and
quality variables (FEI signal probability (\(B\) and daughters), vertex fit quality (\(B\) and daughters), D* veto mass differences, deltaE, cosTBTO)
across 136 deduplicated slots. To these, 12 event-level variables are
appended: 8 event-shape quantities (sphericity, thrust, thrust axis, aplanarity, Fox-Wolfram R2 + 3 harmonic moments) plus the total number of candidates, the input_id of the
highest-sigProb candidate, the highest-sigProb input_id from the other
\(B\) type, and the experiment number. Sparse columns which do not encode
information are excluded from the active network inputs. The main network additionally
receives the category network outputs and its predicted \(B\) type as
inputs.
D* veto#
The D* veto identifies FEI candidates where the \(B\) was reconstructed with a \(D\) meson directly from the \(B\) decay vertex, but the true decay actually proceeded through a \(D^*\) with a soft pion or \(\pi^0\) that was not reconstructed by the FEI.
These misidentified decays are a significant source of crossfeed: the missing soft pion carries information about the \(B\) flavour and charge, so its absence leads to wrongly identified \(B\) type.
Detection is based on the delta mass difference of a reconstructed \(D^*\) hypothesis:
A small \(|\Delta m_{\mathrm{diff}}|\) indicates a likely \(D^*\) decay. The D* veto produces two variables per \(B\) candidate:
Dstp_deltaMassDiff: delta mass difference for the \(D^{*+}\) hypothesisDst0_deltaMassDiff: delta mass difference for the \(D^{*0}\) hypothesis
These are used as input features and are also stored as candidate ExtraInfo for offline selection.
Outputs#
The following variables are written by the module:
Event-level (EventExtraInfo):
BplusScoreThe main signed event-level score. The sign encodes the predicted \(B\) type: positive for \(B^+\), negative for \(B^0\). The magnitude is the maximum main-network probability over candidates present in the predicted sector (background classes excluded). A score close to \(\pm 1\) indicates high confidence; a score near 0 indicates a background-like event.
On events where the predicted sector has no reconstructed candidate, a fallback score is used based on the sum of predicted-sector outputs which encode the predicted \(B\) type.
modeSelector_catBp,modeSelector_catB0,modeSelector_catContCategory network probabilities for \(B^+\), \(B^0\), and continuum. The predicted sector is \(B^+\) when
modeSelector_catBp > modeSelector_catB0.
Candidate-level (ExtraInfo):
modeSelector_eqSigProbThe main-network probability assigned to a candidate’s specific decay mode. Written only on candidates in the predicted sector (\(B^+\) if
BplusScore > 0, \(B^0\) ifBplusScore < 0). This can serve as a mode-aware complement tosigProb.modeSelector_rankLocal sector ranking written on deduplicated representative candidates. In the predicted sector, candidates are ranked by
modeSelector_eqSigProb; in the non-predicted sector, bysigProb. The two sectors are ranked independently.
Best candidate selection#
Because ModeSelector uses additional information with respect to FEI, its
best candidate (rank 1 by modeSelector_eqSigProb in the predicted sector)
can differ from the one selected by sigProb alone, though in practice the
two agree when the FEI already assigns a clearly dominant signal probability.
The mode probability from the main network uses the full event context rather
than a single candidate’s features, and its training dataset better reflects
the composition expected in data.
modeSelector_rank is written on all deduplicated candidates, in both
sectors. The ranking criterion differs by sector:
Predicted sector (
modeSelector_rank == 1): candidate with the highestmodeSelector_eqSigProb.Non-predicted sector (
modeSelector_rank == 1): candidate with the highestsigProb. SincemodeSelector_eqSigProbis not defined for the non-predicted sector, onlysigProbis used for ranking there.
The non-predicted sector contains candidates reconstructed as the \(B\)
type opposite to what the category network predicts for the event. These
candidates form a crossfeed-enhanced sample that can serve as a
calibration or control sample. To restrict the analysis to the predicted
sector only, require BplusScore > 0 for \(B^+\) analyses or
BplusScore < 0 for \(B^0\) analyses.
How to use#
After running the FEI, apply candidate preselections, build the rest of
event and continuum suppression, then add event shape variables and call
modeSelector.modeSelector(). The D* veto reconstruction and the neural
network evaluation are both set up by this single call.
The preselections Mbc > 5.23, -0.15 < deltaE < 0.1, and
cosTBTO < 0.9 were applied during training and are also used in the FEI
calibration. They must be reproduced at inference time:
import modularAnalysis as ma
import modeSelector
track_mask = "[[dr < 2] and [abs(dz) < 4] and [pt > 0.2] and [thetaInCDCAcceptance==1]]"
ecl_mask = ("[[[[clusterReg==1] and [E>0.080]] or [[clusterReg==2] and [E > 0.03]] "
"or [[clusterReg==3] and [E > 0.06]]] and [clusterNHits > 1.5] "
"and [abs(clusterTiming) < 200] and [thetaInCDCAcceptance==1]]")
roe_mask = ("cleanMask", track_mask, ecl_mask)
for b in ['B+:feiHadronic', 'B0:feiHadronic']:
# Preselections must match the training setup and FEI calibration
ma.applyCuts(b, '[Mbc > 5.23] and [-0.15 < deltaE < 0.1]', path=my_path)
# Build rest of event and continuum suppression (required for cosTBTO)
ma.buildRestOfEvent(b, path=my_path)
ma.appendROEMasks(b, [roe_mask], path=my_path)
ma.buildContinuumSuppression(b, 'cleanMask', path=my_path)
ma.applyCuts(b, 'cosTBTO < 0.9', path=my_path)
# Some event shape variables are required by ModeSelector
ma.buildEventShape(
allMoments=False,
cleoCones=False,
jets=False,
collisionAxis=False,
harmonicMoments=True,
foxWolfram=True,
sphericity=True,
thrust=True,
path=my_path,
)
# Add ModeSelector (D* veto reconstruction is included automatically)
modeSelector.modeSelector(
bp_list='B+:feiHadronic',
b0_list='B0:feiHadronic',
output_variable='BplusScore',
path=my_path,
)
# Best candidate selection: keep the rank-1 candidate per list
for b in ['B+:feiHadronic', 'B0:feiHadronic']:
ma.applyCuts(b, 'extraInfo(modeSelector_rank) == 1', path=my_path)
Models are loaded from the conditions database. The payload names are not given by the user: they are derived from the contract version this release implements, so the training is selected by the globaltag alone. Prepend the recommended ModeSelector performance globaltag in addition to the analysis globaltag:
basf2.conditions.prepend_globaltag('<performance globaltag>')
The payload names can still be given explicitly with payload_cat_model and
payload_main_model, which is the way to load the models of an older contract
version that this release still supports. Doing so emits a warning, since the
training is then no longer selected by the globaltag. A payload built for a newer
contract version, or for one that is no longer supported, is rejected with a fatal
error rather than used. See Contract version of a weightfile for the scheme.
The following variables are available after running modeSelector().
Event-level outputs are in EventExtraInfo and accessed via
eventExtraInfo(...):
BplusScore: main output score; positive for \(B^+\), negative for \(B^0\)modeSelector_catBp,modeSelector_catB0,modeSelector_catCont: category network outputs
Candidate-level outputs are in ExtraInfo and accessed via extraInfo(...):
modeSelector_eqSigProb: mode-aware (equalized) signal probability (predicted sector only)modeSelector_rank: local sector candidate rank
Training#
Analysts do not normally need to retrain the networks. This section is provided for reference.
Training is performed on run-dependent MC (MCrd) using FEI-skimmed events.
FEI calibration factors are applied as per-event weights during training,
so the network is conditioned on a sample composition that matches data.
The training sample includes both \(BB\) events and continuum, with
continuum downweighted to 25% relative importance compared to \(BB\).
MC truth matching is applied to assign labels, making use of
mostcommonBTagPDG and mostcommonBTagDeltaP.
The category network is trained first. The main network is trained afterwards, receiving both the feature matrix and the category network output as inputs, so the two stages are coupled. The architecture is fully connected with ReLU activations.
All training scripts and configuration are in analysis/scripts/modeSelector/:
config.py: training configuration, feature definitions and used FEI calibration weightstraining/train.py: training pipeline for both networkstraining/convert_to_onnx.py: export trained models to basf2 MVA weightfiles using ONNX
See analysis/scripts/modeSelector/README.md for full details.
Functions#
- modeSelector.modeSelector(bp_list, b0_list, payload_cat_model=None, payload_main_model=None, output_variable='BplusScore', cat_model_path=None, main_model_path=None, addDstarVetoReco=True, training_mode=False, skip_nn_evaluation=False, store_fei_calib_weight=False, debug=False, debug_max_events=10, path=None)[source]#
Add ModeSelector neural network evaluation to a basf2 path.
This function applies a two-stage neural network to FEI B meson candidates to compute an improved signal probability score. The network considers information from all B candidates in the event.
The main output score is stored in EventExtraInfo using output_variable. Auxiliary category and per-candidate mode outputs are stored with the fixed modeSelector_* names.
- Parameters:
bp_list (str) – B+ meson particle list name to process. Example: ‘B+:feiHadronic’
b0_list (str) – B0 meson particle list name to process. Example: ‘B0:feiHadronic’
payload_cat_model (str) – Conditions DB payload name for the category model. Used only when cat_model_path is None. None (default) uses the name derived from the contract version this release implements, and the training is chosen by the performance globaltag. Passing a name always emits a warning that the default payload is not used. A payload built for an older contract version runs with that contract’s behaviour if the version is in config.SUPPORTED_CONTRACT_VERSIONS (with a warning); a newer or unsupported version is fatal.
payload_main_model (str) – Conditions DB payload name for the main model. Used only when main_model_path is None. Same rules as payload_cat_model.
output_variable (str) – Name of the ExtraInfo variable for the output score. Default: ‘BplusScore’
cat_model_path (str) – Path to the basf2 MVA weightfile for the category network, as produced by convert_to_onnx.py (a .root file, not a raw .onnx file). If None, loads from conditions database.
main_model_path (str) – Path to the basf2 MVA weightfile for the main network, as produced by convert_to_onnx.py (a .root file, not a raw .onnx file). If None, loads from conditions database.
skip_nn_evaluation (bool) – If True, skip loading and evaluating the neural networks and fill deterministic placeholder outputs instead. Intended for debugging or timing studies and emits a warning at module initialization.
training_mode (bool) – If True, skip NN inference and instead expose per-event features and MC truth as EventExtraInfo/ExtraInfo (
modeSelector_feat_XXXX,modeSelector_tr_*,modeSelector_trainSigInputId) for a variablesToNtuple call in the steering script to dump. Seeanalysis/examples/modeSelector/produceTrainingInputs.py. Default: False.addDstarVetoReco (bool) – Whether to add D* veto reconstruction before the NN. Default: True
debug (bool) – If True, print the feature vector and network outputs for the first few events. Intended for comparing against a reference implementation. Default: False
debug_max_events (int) – Number of events to print when debug is True. Default: 10
store_fei_calib_weight (bool) – If True, compute and store modeSelector_feiCalibWeight in EventExtraInfo using the reco path (truth-compatible tag PDG and DeltaP < DELTA_P_THRESH). Returns NaN when reco conditions are not met. Requires mostcommonBTagPDG and mostcommonBTagDeltaP to be defined. Meaningful only on MC. Default: False.
path (basf2.Path) – The basf2 path to add the module to.
Notes
Candidate-level
ExtraInfo(predicted sector only):modeSelector_eqSigProb: main-network probability for the candidate’s specific decay mode.modeSelector_rank: sector-local rank. Predicted sector ranked bymodeSelector_eqSigProb; non-predicted sector ranked bysigProb.
Event-level
EventExtraInfo:BplusScore(or the name given byoutput_variable): signed score, positive for B+ prediction, negative for B0.modeSelector_catB0,modeSelector_catBp,modeSelector_catCont: category network probabilities.
The D* veto reconstruction is added automatically (
addDstarVetoReco=True). PassaddDstarVetoReco=Falseto skip it if already added separately.
- modeSelector.addDstarVeto(particleLists, path: Path = None, deltaMassDiffCut: tuple = (-0.02, 0.02), dMassCut: tuple = (-0.03, 0.03), writeExtraInfo: bool = True, skipTreeFit: bool = True)[source]#
Add D* veto reconstruction to the path for B meson particle lists.
This function reconstructs D* candidates by combining D mesons (first daughter of B candidates) with soft pions or pi0s from the Rest of Event, to identify cases where the FEI reconstructed B -> D X but the true decay was B -> D* X.
- For B candidates with D0 as first daughter:
D*+ -> D0 pi+ (from ROE)
D*0 -> D0 pi0 (from ROE)
- For B candidates with D+ as first daughter:
D*+ -> D+ pi0 (from ROE)
The veto builds its own ROE on private particles that wrap the B candidates, so it neither uses nor creates an ROE related to the B candidates. An ROE built by the user on the input lists, before or after this function, is independent of the veto.
- Parameters:
particleLists (str or list) – Name(s) of B meson particle list(s) (e.g., ‘B+:feiHadronic’ or [‘B+:feiHadronic’, ‘B0:feiHadronic’])
path (basf2.Path) – The basf2 path to add modules to.
deltaMassDiffCut (tuple) – Cut on deltaMassDiff (D* mass diff - true mass diff) in GeV.
dMassCut (tuple) – Cut on D and D* mass deviation (dM) in GeV.
writeExtraInfo (bool) – Whether to write ExtraInfo to particles.
skipTreeFit (bool) – If True, skip the vertex TreeFit (significant speedup). The deltaMassDiff will use InvM-based computation instead of fit-based, and chiProb will not be available (stored as NaN). Candidates are ranked by abs(deltaMassDiffInvM) instead of chiProb. Default: True
The following ExtraInfo fields are added to B candidates:
- For D0 daughter (D*+ and D*0 veto):
Dstp_deltaMassDiff: Delta mass difference for D*+ -> D0 pi+
Dstp_chiProb: Vertex fit chi2 probability for D*+ (NaN if skipTreeFit)
Dst0_deltaMassDiff: Delta mass difference for D*0 -> D0 pi0
Dst0_chiProb: Vertex fit chi2 probability for D*0 (NaN if skipTreeFit)
- For D+ daughter (D*+ veto only):
Dstp_deltaMassDiff: Delta mass difference for D*+ -> D+ pi0
Dstp_chiProb: Vertex fit chi2 probability for D*+ (NaN if skipTreeFit)