Process → category migration (row-normalised, %)
B-like purity—
C-like purity—
S-like purity—
G-like purity—
bc
sg
light / other
per jet j:
event score:
The toy above models each jet's seven-vector as a softmax over logits Af·Sim[f][·] plus Gumbel noise — i.e. a multinomial-logit response with a hand-set flavour-confusion structure (b↔c, s↔g↔u/d). It is not the note's GNN, but it has the same two degrees of freedom that matter: per-flavour separation, and which flavours bleed into which.
Three things to try. (1) Slide s down to zero: the S-like category fills with gluon and light-quark jets and the H→s̄s row goes flat — this is the whole story of why B(H→s̄s) is a limit and not a measurement. (2) Slide b up to maximum: B-like purity saturates but δμb̄b barely moves, because b-tagging was never the bottleneck. (3) Switch on L/M/H purity subdivision and watch the highest-purity bins carve out a nearly background-free b̄b region — that single move is worth more than any amount of extra tagger performance.
Check yourself — why summing the two jets' scores is not obviously right
The note forms K = K₁ + K₂ and takes the argmax. An alternative is the product (a likelihood ratio for "both jets are flavour f"), or a genuine 2-jet joint classifier. Under what conditions does the sum lose information relative to the product, and why might the note's choice still be defensible?
The sum is dominated by whichever jet is more confident: an event with K₁ = (0.95 b) and K₂ = (0.20 b, 0.45 g) still lands with B = 1.15 > G = 0.5 and is called B-like, even though the second jet actively prefers gluon. The product 0.95×0.20 = 0.19 versus 0.0×0.45 would flag the inconsistency. So the sum is insensitive to disagreement between the jets, which is precisely the signature of the dangerous backgrounds — Z(bb̄) recoiling against something else, or g→bb̄ splitting giving one b-like jet.
The defence is threefold. First, for a genuine H→f f̄ decay the two jets are the same flavour, so the product's extra discriminating power applies mostly to backgrounds you are already suppressing by mass. Second, the product is numerically brutal: one jet with a score of 10⁻⁴ kills the event, and tagger scores that far into the tail are not calibrated. Third, and most importantly, the note does not use the categorisation to measure — it uses it to partition, and then fits templates within each partition. Any orthogonal partition gives an unbiased answer; a better one just gives a smaller variance. That is why you can play with the sliders above and never introduce a bias, only lose sensitivity.
Everything the note concludes, in one table
Expected precision (%) on σ(ZH)×B(H→jj) at 68% CL, √s = 240 GeV, L = 10.8 ab⁻¹ (four detectors). Note 23 of the paper. Two variants of the ν̄ν analysis are quoted because they were done independently.
| Channel | b̄b | c̄c | gg | s̄s | ττ | ZZ | WW |
| Z → ℓℓ | 0.60 | 3.47 | 1.93 | 220 | 2.54 | 7.65 | 1.49 |
| Z → qq | 0.32 | 3.52 | 3.07 | 410 | 110 | 50 | 8.74 |
| Z → νν̄ (I) | 0.35 | 2.06 | 1.01 | 100 | 10.6 | 11.4 | 1.28 |
| Z → νν̄ (II) | 0.33 | 2.27 | 0.94 | 140 | 21.8 | 19.8 | 1.89 |
| Combined (I) | 0.21 | 1.56 | 0.85 | 89 | 2.46 | 6.24 | 0.95 |
| Combined (II) | 0.21 | 1.66 | 0.80 | 105 | 3.97 | 10.1 | 1.16 |
Split by production mode and energy
Relative uncertainty (%) on σZH×B and σνeν̄e×B, 68% CL. Table 24. Note the ordering flip in b̄b between the two energies.
| 240 · ZH | 240 · νeν̄eH | 365 · ZH | 365 · νeν̄eH |
| H → b̄b | 0.21 | 1.89 | 0.41 | 0.67 |
| H → c̄c | 1.61 | 19.4 | 3.13 | 3.49 |
| H → s̄s | 120 | 990 | 360 | 290 |
| H → gg | 0.80 | 5.50 | 2.21 | 2.66 |
| H → WW | 1.17 | 15.6 | 3.18 | 5.36 |
| H → ZZ | 9.94 | 130 | 26.0 | 37.1 |
| H → ττ | 3.67 | ∞ | 11.0 | 24.2 |
Inputs you will want at hand
| 240 GeV | 365 GeV |
| Integrated lumi | 10.8 ab⁻¹ (4 IPs, 3 years) | 3.0 ab⁻¹ (5 years) |
| σ(ZH) | ≈ 200 fb — near the L·σ maximum, not the σ maximum (which is at ~255 GeV) | ≈ 125 fb |
| σ by Z decay | qq̄H 136.35 · νν̄H 46.2 · eeH 7.17 · μμH 6.76 fb | qq̄H 84.36 · νν̄H 53.94 · eeH 7.39 · μμH 4.19 fb |
| W-fusion | ≈ 6.1 fb (15% of ZH→νν̄H) | ≈ 29.2 fb (117% of ZH→νν̄H) |
| Z-fusion (eeH) | 0.40 fb vs 6.76 fb Z(ee)H | 3.20 fb vs 4.19 fb |
| Beam energy spread | 222 MeV per beam | — |
| Beamspot | σz 0.64 mm · σx 9.8 μm · σy 25.4 nm | — |
| Detector | Modified IDEA as simulated |
| Tracking | Si vertex (innermost layer R = 1.2 cm) + low-mass drift chamber (1.6% X₀ transverse) + Si micro-strips, in 2 T. Beam pipe R = 1 cm, 0.67% X₀. 100% efficiency assumed above 100 MeV, |η| < 2.56. |
| ECAL | σE/E = 3%/√E ⊕ 0.2%/E ⊕ 0.5% — dual-readout crystals, solenoid moved behind the ECAL so it does not degrade this |
| HCAL | σE/E = 30%/√E ⊕ 5%/E ⊕ 1% — dual-readout lead/fibre |
| ℓ / γ ID | 99% above 2 GeV, |η| < 3; mis-identification neglected entirely |
| Jets | Durham exclusive kT, N = 2 (or 4), E-scheme; flavour tagging by GNN with 7 outputs summing to unity, trained on Z(νν̄)H |
Higgs branching ratios assumed
| Decay | B (%) | Status at the LHC | FCC-ee combined |
| b̄b | 58.24 | observed, HL-LHC proj. 4.4% (ATLAS+CMS) | 0.21% |
| WW | 21.37 | observed | 0.95% |
| gg | 8.187 | no attempt | 0.85% |
| ττ | 6.272 | observed | 2.46% |
| c̄c | 2.891 | searches only, sensitivity ≈ SM | 1.56% |
| ZZ | 2.619 | observed | 6.24% |
| s̄s | 0.024 | unfeasible | limit at 1.6 × SM |
Where this note is softer than its abstract
01 · No experimental systematics anywhere in the leading channels
The ℓℓH and ν̄νH(I) fits state explicitly that there are no systematic uncertainties, only Monte-Carlo statistics. The ν̄νH(II) analysis adds 0.1% (signal) / 5% (background) log-normal normalisations and then reports that the sidebands constrain the backgrounds to "better than a percent level". Every headline number is therefore a statistical floor. Flavour-tagging efficiency calibration, fragmentation modelling, and the jet energy response for b- versus gluon-jets are all absent, and each of them is plausibly at the 0.3–1% level for b̄b. Use the correlated-floor slider on the combination tab to see what that does.
02 · Table 1 contains unphysical placeholder samples
The rows ν̄νH(bd̄), ν̄νH(bs̄), ν̄νH(sd̄), ν̄νH(cū) — and their eeH/μμH counterparts — carry σ = 1000 fb, five times the entire ZH cross-section, for flavour-violating Higgs decays that are forbidden at any observable rate. These are clearly generator bookkeeping placeholders that were never renormalised, and they appear in the figure legends (llH(bd), llH(bs), llH(sd), llH(cu)). If any of them entered a yield table at face value it would swamp the analysis. Nothing in the results suggests they did, but their presence means the sample bookkeeping was not audited.
03 · Table 18 quotes the wrong luminosity
The 365 GeV ν̄νjj results are headed "√s 365 GeV, integrated luminosity 10.8 ab⁻¹" while the run plan, Table 2 and every other table use 3.0 ab⁻¹ at that energy. Either the table header is a copy-paste error, or the numbers are optimistic by √(10.8/3.0) = 1.9. Since Table 24's 365 GeV column (ZH b̄b ±0.41%) is better than Table 18's ±0.73%, the inconsistency is not merely cosmetic.
04 · The s-tagger is doing work no experiment has ever demonstrated
Strange tagging relies on identified high-momentum kaons — in IDEA, on dE/dx or cluster counting in the drift chamber plus time-of-flight. The Delphes parameterisation encodes an assumed performance; there is no fragmentation systematic, no assessment of how the s-tag responds to a gluon jet that fragments into strange hadrons, and no variation of the strangeness-suppression parameter in the hadronisation model. The resulting limit — B(H→s̄s) < 1.6 × SM — is also being placed on a target whose SM value carries a large uncertainty from ms, which the note itself flags but does not propagate.
05 · "gg" is not a clean parton label
A gluon-tagged category is contaminated by g→b̄b and g→c̄c splittings inside jets, and conversely H→b̄b events radiate hard gluons. The seven-output tagger was trained with the true flavour "based on the generated decay of the Higgs boson" — i.e. the label is the decay mode, not the parton reaching the detector. That is the right label for this measurement, but it means the tagger has partly learned to identify the Higgs decay via its global event shape, and its performance will not transfer to any other sample. It also makes the confusion matrices in Figures 7 and 14 optimistic if the parton shower model is wrong.
06 · S/√(S+B) is quoted as "significance" in the yield tables
The per-category numbers in parentheses in Tables 4 and 9 use S/√(S+B), which is neither the asymptotic discovery significance nor the quantity entering the profile likelihood. It systematically overstates categories where B ≫ S and understates highly pure ones. It is a sorting aid, not a result — but it is presented adjacent to the yields as if it were one.
07 · Four identical detectors is an assumption, not a design
10.8 ab⁻¹ presumes four interaction points each with IDEA-like performance. The note tells you the scaling honestly (halving the IPs costs a factor 1.7 in luminosity, not 2, because of the machine layout), but the headline 0.21% becomes ~0.28% for two IPs before any systematic is added — and the current FCC-ee baseline has fluctuated between two and four IPs during the design study.
08 · Internal inconsistency in the ν̄ν selection table
Table 8's selection row reads "70 < mvis < 150 GeV, 60 < mmiss < 220 GeV" while the text specifies mmiss in 50–140 GeV and the fit is performed over 60–140 GeV. The efficiencies quoted (86–92%) are only reproducible with one of these. Small, but the reader cannot reconstruct the cutflow.
Harder problems
Problem 1 — the ℓℓ efficiency dip on b̄b
In the ℓℓH cutflow, the "no additional leptons with p > 25 GeV" cut removes 5% of H→b̄b events but essentially none of H→gg or H→s̄s. The note remarks it "might be considered to veto additional isolated leptons instead". Quantify the effect from first principles and say what the right fix costs.
B hadrons decay semileptonically with B(b→ℓν X) ≈ 10.5% per lepton species, and with the b→c→ℓ cascade the inclusive rate of a lepton from a b-jet is ~20%. With two b-jets and a spectrum that puts a few percent of those leptons above 25 GeV (the b-jets carry ~60 GeV each, and the lepton takes a soft fraction), you land at a few percent per event. The observed 5% is entirely consistent.
The proposed fix — vetoing only isolated leptons — recovers most of it, because leptons from heavy-flavour decay sit inside the jet cone. The cost is that the veto's original job was to suppress H→ττ with a leptonic τ and ZZ→ℓℓℓℓ, whose leptons are isolated, so isolation is exactly the discriminant you want. In practice the fix is nearly free, worth ~5% in b̄b yield, i.e. ~2.5% in δμ — small compared to the 40% gap identified in the Fisher exercise, which is why nobody bothered. Recognising which inefficiencies are worth fixing is the skill; this one is not.
Problem 2 — the exclusive-2 clustering choice
Every analysis here forces the event into exactly N jets with the exclusive Durham algorithm and uses d2,3, d3,4 as discriminants. Why is exclusive clustering the natural choice at an e⁺e⁻ collider and pathological at a hadron collider, and what specifically do d2,3 and d3,4 buy in this analysis?
Exclusive clustering asserts that every reconstructed particle belongs to one of N jets. At an e⁺e⁻ collider that is nearly true: there is no underlying event, no pile-up, no beam remnant, and the total visible energy is √s by construction. At the LHC the same statement is false at the level of tens of GeV per event, which is why hadron colliders use inclusive algorithms with a fixed R.
di,i+1 is the kT distance at which the event transitions from i+1 to i jets, d = 2 min(Ei², Ej²)(1−cos θij). It is therefore a continuous measure of "how badly does this event want to be more than two jets". H→gg and H→WW*→4q produce large d2,3; H→b̄b produces small d2,3 but non-negligible d3,4 from b-decay substructure; a leptonic τ leaves d2,3 = 0 exactly, which is why the note's cut is literally d2,3 > 0 — a topology veto disguised as a kinematic one. Feeding log(d) rather than d to the network is not cosmetic either: the distribution is roughly log-uniform over three decades.
Problem 3 — reconstruct the ν̄ν sensitivity from the yield table
In the ν̄νH categories, the three b̄b bins hold S = 61 092 / 55 486 / 101 039 signal on totals of 199 630 / 88 233 / 111 892. Estimate δμb̄b by hand and compare with the note's 0.35%. Then explain why the b̄bhigh bin, with 90% purity, contributes only about half the information.
Σ S²/T = 61092²/199630 + 55486²/88233 + 101039²/111892 = 18 695 + 34 891 + 91 249 ≈ 144 800, plus a few hundred from the leakage bins. δμ = 1/√144 800 ≈ 0.26%, against the note's 0.35%, the difference again being MC statistics and the seven-parameter fit (compare Table 14: 0.26% without MC stat, 0.33% with — the counting estimate reproduces the no-MC-stat number almost exactly).
The high-purity bin contributes 91 249 of the 145 000, i.e. 63%. Its information density is S²/T = S·(S/T) = S × purity, so information scales as yield times purity, not purity alone. The low-purity bin has 61 092 events at 31% purity → 19 000; the high bin has 101 039 at 90% → 91 000. Doubling purity in a bin is worth exactly as much as doubling its yield. That symmetry is the reason the note's L/M/H subdivision (which changes neither total yield nor total purity, only their correlation across bins) improves the result at all: it moves yield into the high-purity end of the distribution.
Problem 4 — the s̄s limit, done properly
The note reports a 95% CL upper limit of 1.6 × SM on B(H→s̄s) from the combination, with δμ ≈ 89–105%. Show that these two statements are consistent, and identify the statistical subtlety that makes "1.6 × SM" a slightly odd thing to quote.
For a Gaussian likelihood centred on μ = 1 with σμ ≈ 0.9, the one-sided 95% upper limit is μ < 1 + 1.64σ ≈ 2.5. To land at 1.6 you need σμ ≈ 0.37, or an expected limit computed on the background-only Asimov set (μ = 0), where the limit is 1.64σ ≈ 1.5. The 1.6 is therefore a background-only expected limit, while the 89–105% are signal-injected uncertainties on μ. Both are legitimate; they answer different questions and should not be read as the same number in different units.
The oddity: quoting a limit in units of the SM prediction when that prediction is itself poorly known. B(H→s̄s) ∝ ms(mH)², and ms carries a several-percent uncertainty which is amplified by the running and squaring. The note flags this in the introduction ("a large uncertainty related to the limited knowledge of the mass of the strange quark") and then quotes a limit in exactly those units. The physically meaningful statement is a limit on κs·ysSM, or better, on the absolute partial width.
Problem 5 — design the measurement you would actually do
You have been asked to improve B(H→c̄c) beyond 1.6%. You may spend effort on exactly one of: (a) charm tagging, (b) jet energy resolution, (c) more luminosity, (d) the categorisation scheme. Argue for one using the numbers on this page.
Work out what limits it. In the ℓℓ channel the c̄c categories hold S ≈ 2 960 on T ≈ 8 244 — purity 36%, and the contamination is dominated by ZZ background (4 587 events) rather than by other Higgs decays (≈ 100 events). So c̄c is background-limited, not tag-limited.
(c) more luminosity scales as 1/√L and is the least efficient use of anything. (b) jet energy resolution is nearly irrelevant, because the discriminant is mrecoil from the leptons, which never touches the jets. (a) charm tagging helps but you are fighting Z→cc̄ inside ZZ, which has the same flavour content as your signal — better c-tagging does not separate them at all; only the recoil mass does. That leaves (d): the ZZ background is separable from ZH by the Z polar-angle distribution (1/σ dσ/dcos θ rises steeply toward |cos θ| = 1 for ZZ and is flat for ZH — the note plots exactly this in Fig. 3 and feeds it to the network) and by the recoil lineshape. Finer categorisation in cos θℓℓ and tighter recoil binning attack the actual limitation. In the ν̄ν channel the same reasoning points instead at the 2D (mvis, mmiss) template granularity.
The general lesson: look at what is in the denominator of S²/(S+B) before you improve anything.