Open Data CERN
  • Login
  • Help
    Discussion forum
    Search tips
  • About
    CERN Open Data
    ALICE
    ATLAS
    CMS
    DELPHI
    JADE
    LHCb
    OPERA
    TOTEM
    Glossary

ATLAS multi-process simulation for ML-based jet flavour tagging (JetSet2)

ATLAS collaboration

Cite as: ATLAS collaboration (2026). ATLAS multi-process simulation for ML-based jet flavour tagging (JetSet2). CERN Open Data Portal. DOI:10.7483/OPENDATA.ATLAS.XEVX.LJJ2

Dataset Derived Simulated Datascience ATLAS 13TeV, 13.6TeV pp CERN-LHC


Description

Flavour tagging, the identification of jets originating from bottom and charm quarks and from hadronically decaying tau leptons, is essential for many physics analyses at the ATLAS experiment. This dataset, released for public use, can be used to train and evaluate machine learning models for jet flavour tagging. It is the second ATLAS flavour-tagging open dataset (JetSet2), the successor to the JetSet dataset (record 93940, arXiv:2505.19689), and is the data on which the tagger scaling study ATL-SOFT-PUB-2026-002 is based.

The dataset contains approximately 11 billion fully labelled jets from 29 simulated samples covering 7 physics processes: top quark pair production, QCD dijet production (including b-filtered samples), Z' production, associated VH production with H to bb, cc and tau tau, and gamma* to tau tau, from the ATLAS Run 2 (13 TeV) and Run 3 (13.6 TeV) simulation campaigns. Each jet carries its kinematics, truth labels, up to 40 inner-detector tracks, 10 electrons, 50 particle-flow objects and 5 truth heavy-flavour hadrons, the scores of the ATLAS GN2v01 and GN3EPCLV01 taggers, and a sample identifier (DSID). Files are in HDF5 format with one structured dataset per object type.

Unlike JetSet, the files are released after preprocessing: every jet is kept exactly once and the training files carry a per-jet weight that balances the flavour classes and their kinematics, so no resampling step is needed. The files are direct inputs to the ATLAS tagger training framework salt (code, paper).

  • train/: 198 files of about 50 million jets (about 70 GB) each, 9,859,015,723 jets in total. All samples are interleaved in every file, so any single file is a representative training subset.
  • test/: 111 files grouped by physics process, 1,096,231,479 jets in total, each with up to 10 million jets. The test files additionally carry the scores of TN25-86M, the 86-million-parameter tagger from the scaling study, whose model files are provided in the accompanying repository.

The accompanying repository documents every variable, the samples, the file layout, and provides the preprocessing and training configurations, the TN25-86M model, and example notebooks.

Related datasets

The first ATLAS flavour-tagging open dataset, of which JetSet2 is the successor

ATLAS $t\bar{t}$ simulation for ML-based jet flavour tagging (JetSet)

Dataset characteristics

10955247202 entries. 309 files. 14.4 TiB in total.

How were these data generated?

Jets were extracted from ATLAS DAOD_FTAG1 derivations of the Run 2 (mc20) and Run 3 (mc23) simulation campaigns with the ATLAS training-dataset-dumper (tag 26-06-03_jetset2-prod), keeping the full set of GN3 tagger input variables. The dumps were preprocessed with umami-preprocessing v0.3.2 in split-and-reweight mode: events were split by event number into training (90%) and test (10%) sets, every variable of every object was kept, a 6-class ghost-association flavour label and the sample DSID were added to each jet, and the training jets were given a weight that maps each flavour's (pT, |eta|) spectrum onto the mean spectrum across flavours. The test files were then scored with the TN25-86M tagger using salt 0.13. The exact configurations are in the accompanying repository.

training-dataset-dumper

umami-preprocessing

Preprocessing and training configurations

How can you use these data?

The accompanying repository documents the file layout and every variable, and provides the preprocessing and training configurations (salt), the TN25-86M model in ONNX and checkpoint form, and example notebooks for reading the files and evaluating the in-file tagger scores. The files are plain HDF5 and can be read with h5py; salt, umami-preprocessing, atlas-ftag-tools and puma are installed with 'pip install salt-ml umami-preprocessing atlas-ftag-tools puma-hep'. If this dataset is used in a publication, please cite this record together with the ATLAS note ATL-SOFT-PUB-2026-002.

JetSet2 documentation, configurations, model and examples

ATL-SOFT-PUB-2026-002

JetSet (volume 1) record

arXiv:2505.19689

General ATLAS Open Data citation policy


      

Files and indexes

Disclaimer

These open data are released under the Creative Commons Zero v1.0 Universal license.

Logo CC0-1.0

Neither the experiment(s) ( ATLAS ) nor CERN endorse any works, scientific or otherwise, produced using these data.

This release has a unique DOI that you are requested to cite in any applications or publications.

ALICE experiment
ATLAS experiment
CMS experiment
DELPHI experiment
JADE experiment
LHCb experiment
OPERA experiment
PHENIX experiment
TOTEM experiment
© CERN, 2014–2026 ·
Terms of Use ·
Privacy Policy ·
Help ·
GitHub ·
Twitter ·
Email
Powered by Invenio
Open Data Portal v1.2.1
CERN