ATLAS ANA-HIGP-2024-12  ·  draft 1.0, 26 July 2026  ·  interactive re-derivation

Reading the Higgs width off a 40 MeV wobble — and what the draft doesn't tell you

The Higgs is 4 MeV wide and the detector is 1.6 GeV wide, so the lineshape is useless — except that the resonance interferes with the continuum underneath it, and the cross term drags the peak sideways by an amount that grows as √ΓH. This letter turns that drag into a limit. Every panel below rebuilds a piece of it from the paper's own Table 1 and checks itself against the paper's own published limits. One number in the draft contradicts its own table, and the observable the whole method rests on is never quoted.

Fig. 01

Physics

Display

Expected events / GeV vs mγγ [GeV]
Net interference
as % of signal
Peak shift
Paper, ΓSM−2.05 %
signal gg-interference qg-interference gg + qg paper's value
Eq. 1 — ggγγ= sig mγγ2 mH2+iΓHmH +bkg
cross term — 2Re[sig mγγ2 mH2+iΓHmH bkg] = antisymmetric × Re + symmetric × Im
Because ΓH is four hundred times smaller than the mass resolution, the two pieces of the cross term collapse onto distributions. The symmetric piece becomes a δ-function at mH — after smearing, a copy of the signal's own lineshape, scaled negative: it removes about 2 % of the events. The antisymmetric piece becomes a principal value 1/(m2 − mH2) — after smearing, an odd function that integrates to zero but moves events from one side of the peak to the other. That is the whole measurement.
smeared — I(m)= μΓHΓHSM NI[R(x)+ βH(x)], x=mmH
with — H(x)=PV R(t)xt dt2σ F(xσ2) F = Dawson function
peak shift — δ=βNI σ2π NS ∝ σ: a blunter detector sees a bigger shift
R(x)
the detector resolution function — the paper's double-sided Crystal Ball. The signal, and the symmetric half of the interference, both have exactly this shape.
H(x)
its Hilbert transform: odd, zero-integral, peaking at x = 1.31σ and falling as 1/x. This is the smeared principal value, and it is what tilts the peak.
β
how much dispersive there is per unit absorptive — equivalently the relative phase of the two amplitudes. The paper never gives this, or the shift it implies. The fit panel infers it.
NI
the integrated interference yield: Table 1's −60.3 (gg) and +2.4 (qg) in the low-pT category.
μ
signal strength, floated in the fit. It is what makes the rate uninformative and the shape everything.
Worked example, low-pT category, at the SM width. NI = −60.3 + 2.4 = −57.9 events on NS = 3045, so the rate change is −57.9/3045 = −1.90 %. With σ = 1.60 GeV and β = 0.505 the shift is 0.505 × (−57.9) × 1.60 × √(2π) / 3045 = −0.0385 GeV = −38.5 MeV. Set the sliders to those values and the readouts land on both numbers.
Check yourself — why does a worse detector see a bigger shift?
Drag the resolution slider from 1.6 GeV to 3.2 GeV at fixed β. The apparent peak shift doubles. A detector that measures the mass worse reports a larger interference shift. Why is that not absurd — and why does the measurement still get worse, not better?

Both halves follow from the fact that the dispersive term is scale-free.

Why the shift grows. The true interference distortion is a principal value, 1/(m2 − mH2), which has no width of its own — it is equally large at every distance from the resonance, in the sense that x·H(x) → 1 for any x. Convolving it with a resolution of width σ produces a curve whose extremum sits at x = 1.31σ and whose height is 0.765/σ. Meanwhile the signal peak height is NS/(σ√2π), also falling as 1/σ. So the ratio of dispersive amplitude to peak curvature — which is what sets where the maximum of the sum lands — carries one net factor of σ. Algebraically: near the peak the sum is NSR(x) + A·H(x), and expanding to first order, R′(x) ≈ −x·R(0)/σ2 while H′(0) = 1/σ2. Setting the derivative to zero, the σ−2 cancels and δ = A/(NSR(0)) = Aσ√(2π)/NS. One factor of σ survives, from R(0).

Why the measurement still degrades. The quantity that matters is the shift divided by the statistical uncertainty on the fitted peak position. For a peak of NS events on a locally flat background of density b, that uncertainty goes as σ3/2√b/NS: one power of σ because the peak is wider, and half a power because a wider peak admits more background into the region that constrains it. So the significance of the shift scales as σ/σ3/2 = σ−1/2. Blunter detector, bigger shift, worse measurement — by half a power.

The toy confirms the exponent empirically: holding the expected stat-only limit fixed at the paper's 34 MeV and re-solving for the shift at each resolution gives −38.5 MeV at 1.6 GeV, −74 MeV at 2.4 GeV and −108 MeV at 3.0 GeV, i.e. shift ∝ σ1.64. The naive argument above predicts σ1.5; the extra 0.14 comes from the Crystal Ball tails and from the finite [105,160] GeV fit window, neither of which is scale-free.

The judgement call. This is why the paper's silence on its own mass resolution matters. Because shift and resolution are degenerate along σ1.64, no reader can convert the published limit into a statement about the interference calculation without knowing both. A single line — "the interference shifts the peak by −X MeV at the SM width for our σeff of Y GeV" — would make the letter self-contained. Whether that omission is an oversight or a deliberate choice not to quote a number that depends on the parton shower is not something the draft lets you tell.

What the paper doesn't publish

Systematics (paper's values)

Asimov truth

Fitting…
−2 Δln L vs ΓH [MeV]
68 % UL, stat only
95 % UL, stat only
68 % UL, total
95 % UL, total
stat only + peak position total paper's limits
Eq. 4 — Fc(mγγ)= μKSNScfSc +μΓHΓHSM [NI,ggcfI,ggc +NI,qgcfI,qgc] +NBcfBc
Eq. 5 — t(ΓH)=2ln L(ΓH,μ^^, θ^^) L(ΓH^, μ^,θ^) 68 % at t = 1, 95 % at t = 3.84
The panel builds this model bin by bin over 105–160 GeV in both categories with the Table 1 yields, constructs an Asimov dataset at the paper's truth, and profiles seven parameters — the signal strength, and a normalisation and two shape parameters for the continuum in each category — plus the peak position and the interference normalisation when their systematics are switched on. Nothing is fitted to the paper's answer except one number: the mass shift at the SM width, tuned once against the expected stat-only 68 % limit. Every other limit on the readout is then a prediction.
Check yourself — set β to zero and the limit disappears entirely. Why?
In the distortion panel, set the dispersive fraction β to 0. The interference is still there — it still removes 2 % of the signal events. Yet the fit returns no limit on ΓH at all, at any width. Where did a 2 % effect go?

Into μ. This is the single most important structural fact about the measurement, and it is the reason the paper can claim independence from the Higgs couplings.

With β = 0 the interference contributes exactly the signal's own lineshape, scaled by −0.0205√(μΓ/ΓSM). The model's resonant part is then μNSR(x) [1 − 0.0205√(Γ/ΓSM)/√μ], which is a single number multiplying a single fixed shape. Any change in Γ can be undone by a compensating change in μ — the likelihood has a flat direction, and the profile likelihood ratio is identically zero along it. The toy shows this: at β = 0, t(1000 MeV) = 0.16, essentially flat, and no crossing of t = 1 exists.

That flat direction is the same one the paper describes in §1 in coupling language. The on-shell rate goes as cg2cγ2H, so the rescaling c → κc, ΓH → κ4ΓH leaves it untouched. The interference goes as cgcγ, one power of each, so it scales as κ2 = √(Γ/ΓSM) at fixed rate — that is where the square root in Eq. 4 comes from. But scaling alone is not enough: a rate change and an interference change are the same thing if they have the same shape. The degeneracy is broken only by the dispersive term, which has a shape no amount of μ can imitate, because it is odd and μ multiplies something even.

Consequence for reading the paper. Table 1's interference yields — the −60.3 and −62.9 — carry almost none of the sensitivity by themselves. They fix the normalisation of a distortion whose shape is what gets measured, and the paper never prints that shape's amplitude. A reader who tries to sanity-check the 250 MeV expected limit from Table 1 alone will fail, not because the table is wrong but because the necessary number is not in the letter.

The judgement part. Whether β is closer to 0.5 or to 1.5 is a genuine theory question — it is the relative phase between the resonant amplitude and the continuum box, and it depends on the order of the calculation and on the pT selection. The toy can only say what the paper's sensitivity requires given an assumed resolution, and that is a curve in the (shift, σ) plane, not a point. The third sub-tab draws that curve.

Check yourself — the toy reproduces seven limits after tuning one number. Is that impressive?
One free parameter is fixed on the expected stat-only 68 % limit (34 MeV). The toy then predicts 87.9 / 52.4 / 116.0 / 64.4 / 250.9 / 93.4 / 322.4 MeV where the paper has 91 / 52 / 115 / 61 / 250 / 90 / 319. How much of that agreement is real, and how much is arithmetic that had to work?

Some of it had to work; the interesting part did not.

What was nearly guaranteed. The three 68/95 pairs are not independent. Once the toy has the right likelihood shape, one crossing determines the other. And the shape is forced: with only a statistical uncertainty, t is exactly quadratic in √Γ, because the interference amplitude is linear in √Γ and everything else is Gaussian. So "expected stat-only 95 % = 87.9 against 91" is really a test of that functional form, not of the toy's physics. The paper's own numbers pass the same test: from Γ̂ = 4.1 and the 34 MeV crossing, the √Γ-Gaussian predicts 90.0 where the paper prints 91.

What was not guaranteed, and this is the real result. Switching the observed Asimov on — moving the truth from 4.1 to 12.0 MeV — changes both crossings in a way nothing was tuned to, and the toy gives 52.4 / 116.0 against 52 / 115. More strikingly, adding the two systematics at the paper's published sizes takes the expected 95 % limit from 88 to 251 MeV where the paper has 250. That is a factor of 2.9 inflation reproduced to 0.4 %, and it is not a shape identity: it is the specific way a multiplicative uncertainty on an amplitude that itself scales as √Γ blows up at large Γ, which is exactly the mechanism §7 describes in one sentence and never quantifies.

What genuinely does not work. Group by group, the toy misses Table 2 badly: it implies an experimental contribution of 50.7 / 139.4 MeV against the paper's 43 / 119, and a theory contribution of 34.8 / 165.7 against 50 / 264. The totals agree and the decomposition does not, because Table 2 is built by progressively fixing groups on the observed likelihood while the toy adds one nuisance at a time to a stat-only baseline. Those two operations agree only when the nuisances are uncorrelated with each other and with Γ, and here the interference normalisation is strongly correlated with Γ by construction. So: the toy reproduces the paper's answer and does not reproduce the paper's accounting of it, and the second failure is honest information about how much a table like Table 2 can be read as a set of independent contributions. It cannot.

The judgement part. One tuned parameter against seven predictions sounds like a strong test, but the seven are not seven degrees of freedom — closer to three. A sceptical reading is that the toy demonstrates the model class is right (√Γ scaling, log-normal amplitude systematic, Asimov asymptotics) without demonstrating that any individual ingredient is right. That is a fair reading, and it is why the resolution slider matters: it shows that a whole family of (shift, σ) pairs fits equally well, so the toy is genuinely not pinning down the physics the paper omitted.

Uncertainty groups (Table 2)

Contribution to the limit on ΓH [MeV]
68 % half-width
95 % half-width
95 % CL upper limit
Paper (observed)319 MeV
theory experimental statistical paper's Table 2
the shape of t — t(ΓH)= (ΓH ΓH^)2 σ02+f2ΓH f = multiplicative error on the amplitude
√ΓH
the natural variable. The interference amplitude is linear in it, so with only statistical errors the likelihood is an ordinary parabola here and nowhere else.
σ0
everything that is a constant uncertainty on the amplitude: the data statistics, the mH constraint, the energy scale. These do not grow with ΓH.
f
the fractional uncertainty on the interference normalisation. Because it multiplies an amplitude that already grows as √ΓH, its absolute effect grows with ΓH — the denominator opens up, the parabola flattens, and the 95 % crossing runs away.
Fit this two-parameter form to the paper's own published limits and it closes exactly, because two numbers determine two parameters. The check is that it closes with the same parameters for the expected and observed sets: f = 29.1 % and σ0 = 5.32 from (61, 250), against f = 27.8 % and σ0 = 5.41 from (90, 319). And that the independent binned toy, which knows nothing about these numbers, returns f = 27.1 %.
Check yourself — the KI prior is 39 %, so why does the fit behave like 27 %?
§6 puts a 39 % uncertainty on the gg-interference normalisation and calls it the largest single contribution to the error on ΓH. Three independent routes — the paper's expected limits, the paper's observed limits, and the binned toy — all say the interference amplitude behaves as though it carried about 27–29 %. Where do the missing twelve points go?

Two places, and they are worth separating because only one of them is physics.

Six points are a definition. A normalisation systematic quoted as "39 %" is implemented as a log-normal, not as 1 + 0.39θ. The reason is structural: a linear Gaussian on a normalisation passes through zero at θ = −2.56, and an interference amplitude that can vanish makes ΓH formally unbounded. The toy demonstrates this — with a linear constraint the expected 95 % limit came out at 610 MeV against the paper's 250, while the 68 % limit was barely affected, which is the signature of a tail pathology rather than a physics error. Switching to the log-normal (1.39)θ fixes it, but that form has a log-slope of ln(1.39) = 0.329, so "39 %" is really 32.9 % in the exponent that the likelihood actually sees.

The remaining six points are the fit constraining its own nuisance. The 39 % is a prior, and the data push back on it. Two handles do the pushing. First, the qg-interference is not scaled by KI — it is a separate, LO-accurate contribution of opposite sign — so the total distortion is not simply proportional to the nuisance, and a large pull on KI changes the gg:qg balance in a way the two categories see differently. Second, and more importantly, KI and ΓH enter the amplitude in the fixed combination KI√ΓH only for the gg piece, while the mH nuisance acts on the peak position rather than the amplitude; the three-way correlation structure leaves the profiled interval narrower than the prior alone would give.

Numerically: prior 39 % → log-slope 32.9 % → profiled 27.1 % in the toy, against 27.8 % and 29.1 % inferred from the paper's two sets of published limits by an entirely different route. Three numbers within two points of each other from three independent calculations is the strongest single consistency check in this file.

The judgement part. None of this makes the 39 % too conservative. It covers the gap between an NLO matrix element and an NNLOsv calculation, which is a real and poorly controlled difference, and the paper is right that it dominates. What the exercise does show is that "the largest single contribution to the total uncertainty" is being reported at its prior width while the fit is using something meaningfully tighter — so a reader who mentally rescales the limit by taking the 39 % at face value will over-correct. Whether that distinction belongs in a letter, as opposed to a supporting note, is a matter of taste.

Interference split between categories

Which categories enter the fit

Interference-to-signal ratio, and what each category buys
I/S, pT < 30
I/S, pT ≥ 30
95 % UL, stat only
vs merged
pT < 30 GeV pT ≥ 30 GeV combined Table 1 / paper's claim
The sentence that does not match the table

§4, line 135: "The low-pTγγ category contains events with pT < 30 GeV, where the size of the interference contribution relative to the signal is largest". Table 1, four lines later, gives |Igg|/S = 60.3/3045 = 1.98 % at low pT and 62.9/2527 = 2.49 % at high pT. Including the qg term: 1.90 % against 2.23 %. The relative interference is larger in the high-pT category on either definition, and the toy fit agrees — run each category alone and the high-pT one gives the better limit (137 MeV against 170 MeV) despite having 17 % fewer signal events. Either the sentence or the table needs fixing before submission.

Check yourself — the split into two categories is worth about 1 %. Should it be there?
Press "Merged" and watch the 95 % limit move from 87.7 MeV to 89.1 MeV. The pT categorisation — one of the two observables §1 says the measurement is built on — buys 1.6 %. Is the categorisation pointless, and if not, what is it actually for?

It is nearly pointless for sensitivity, and it is not there for sensitivity.

Why the gain is so small. Splitting a dataset helps when the sub-samples have different signal-to-distortion ratios, because then the fit can play them against each other. Here they barely differ: −1.90 % against −2.23 %, a contrast of 17 %. The information a two-bin split adds over a merged fit scales with the square of that contrast weighted by the statistical power of each bin, and 17 % squared is 3 %. Getting 1.6 % out of it is about right. The toy confirms the categories are otherwise statistically independent: combining the two single-category amplitude precisions in quadrature, 1/√(1/5.722 + 1/5.042) = 3.779, against 3.782 measured from the simultaneous fit. To three decimal places the fit is doing nothing but adding two independent measurements.

What the split is really for. §4 says it plainly, in a clause that is easy to skim past: the boundary sits at 30 GeV "because predictions of the qg-interference obtained with different parton-shower algorithms diverge significantly above this value". The categorisation is a systematics containment device, not a sensitivity device. It isolates the region where the least-trusted component of the model is least trusted, so that a nuisance parameter can absorb it without contaminating the region that carries the measurement. §6 backs this up with the migration uncertainties: 15.8 % on the qg-interference in the low-pT category and 23 % from the parton-shower subtraction scheme — large numbers attached to a small component, which is exactly what you would build a category to quarantine.

The internal tension this exposes. §1 lines 60–62 say the sensitivity is "accessed through two observables: the modification of the lineshape and the different interference-to-signal fractions in event categories". Three lines earlier the same paragraph says the sensitivity "derives entirely from the interference-induced distortion". The toy sides with the second sentence: 98 % of the information is lineshape. The first sentence oversells the categorisation, and it does so in the paragraph a reader is most likely to quote.

The judgement part, and a caveat on my own number. The toy gives both categories the same mass resolution and the same dispersive-to-absorptive ratio, differing only in yields. In reality the shift should differ between them — it depends on the interference kinematics, which is the whole reason the two categories exist — and a genuine difference would raise the categorisation's value above 1.6 %. How far above, I cannot say from the letter, because the per-category lineshapes are not published. So read "1.6 %" as a lower bound derived under an explicitly stated simplification, and the paper's own "at most 17 %" for a finer binning as the other end of the plausible range.

The paper's tables, transcribed

Table 1 — expected yields in mγγ ∈ [105, 160] GeV at the SM hypothesis (ΓH = ΓHSM, μ = 1). Signal includes KS = 1.4; gg-interference includes KI = 1.64. The last two columns are computed here, not in the paper.

ProcesspTγγ < 30 GeVpTγγ ≥ 30 GeVTotal/ signal, low/ signal, high
Signal (ggF)304525275572
gg-interference−60.3−62.9−123.2−1.98 %−2.49 %
qg-interference+2.4+6.5+8.9+0.08 %+0.26 %
Net interference−57.9−56.4−114.3−1.90 %−2.23 %
Continuum background6.7 × 1055.1 × 1051.18 × 106

Table 2 — decomposition of the uncertainty on ΓH, in MeV. The quadrature column is computed here: the groups close to 0.5 %, and each limit is exactly Γ̂H = 12.0 MeV plus the entry.

Source68 % CL95 % CL68 % check95 % check
Theory5026450.2264.0
  NNLOsv K-factor50263
  Others423
Experimental4311942.9119.6
  mH (from H→ZZ*)3597
  Energy scale2466
  Energy resolution413
  Others519
Statistical40103
Total7730777.1307.4

Results, as published. "Half-width" is the Table 2 entry; the limit is Γ̂H + half-width for the observed column and ΓSM + half-width for the expected one.

ConfigurationObs. 68 %Obs. 95 %Exp. 68 %Exp. 95 %Toy 95 % (exp.)
Statistical only52115349187.9
Total9031961250250.9
In units of ΓSM = 4.07 MeV22×78×15×61×

Everything the toys need that the paper does not give

QuantityStatus
Mass shiftNever quoted, at any width. It is the observable the entire measurement rests on. The fit panel infers −38.5 MeV at ΓSM for a 1.6 GeV resolution, degenerate with resolution along σ1.64. Literature: ≈−70 MeV inclusive at LO (Martin), O(−200 MeV) under pT cuts (Dixon–Li).
mγγ resolutionOnly "O(1–3) GeV" in §1. No per-category value, no DSCB parameters, though §5 says six of them were fitted. The toy uses 1.60 / 1.55 GeV with typical ATLAS tails and exposes both as sliders.
Interference phaseNot given. Equivalent to the dispersive-to-absorptive ratio β, and therefore to the shift above.
Non-ggF signal§3 includes VBF, VH, tt̄H, bb̄H and tH as non-interfering signal; Table 1 lists none of them. The cross-check that Table 1 is ggF-only: 48.61 pb × 2.27×10−3 × 140.1 fb−1 = 15 460 produced, so ε = 5572/15 460 = 36 %, a normal H→γγ efficiency; and 48.61/1.4 = 34.7 pb is the NLO ggF cross-section, confirming KS is NLO→N3LO on ggF alone.
Spurious signalThe criterion is stated (§5), the value is not. This is the standard H→γγ background-modelling systematic.
Background shapeFunctional form given, parameters not — they are fitted to data, so this is reasonable, but it means the local S/B under the peak cannot be checked. The toy assumes a continuum falling as m−4.5, giving S/B ≈ 4 % in ±2σ.

Critique

1 — §4 contradicts Table 1 on which category has the larger relative interference

The text says low-pT; the table says high-pT, by 1.98 % against 2.49 % on the gg term and 1.90 % against 2.23 % on the net. Running each category alone through the fit gives 170 MeV and 137 MeV respectively, confirming the table. This is the one error in the draft that would change a reader's understanding of the analysis design, and it sits in the sentence that justifies the design.

2 — §1 claims two observables; the second contributes about 2 %

Lines 60–62 say the sensitivity is accessed "through two observables: the modification of the lineshape and the different interference-to-signal fractions in event categories". Merging the two categories in the toy costs 1.6 % of the limit. Three lines earlier the same paragraph says the sensitivity derives "entirely" from the lineshape distortion, which is what the numbers support. The categorisation is a systematics-containment device — §4 says so, in the clause about parton-shower divergence above 30 GeV — and describing it as a second observable oversells it.

3 — the mass shift is never quantified

An interference measurement whose sensitivity comes entirely from an induced peak shift should state the size of that shift, at least at the SM width, together with the effective mass resolution it is measured against. Neither appears. The consequence is concrete: no reader can convert ΓH < 319 MeV into a statement about the underlying interference calculation, and no reader can check the result against Martin or Dixon–Li without redoing the analysis.

4 — Table 2's columns are half-widths presented as if they were limits

The caption says "the impact on the 68 % CL uncertainty and on the 95 % CL is given". The entries are neither limits nor symmetric uncertainties: they are the amount each group adds to Γ̂H = 12.0 MeV to reach the corresponding one-sided limit. They do close in quadrature to 0.5 %, which is worth stating explicitly since it is the only thing that makes the table readable as a decomposition — and even then, the toy shows that reproducing the total does not mean the groups can be read independently.

5 — the 39 % KI prior is reported at a width the fit does not use

Called "the largest single contribution to the total uncertainty in ΓH". Implemented as a log-normal it has a log-slope of 32.9 %, and after profiling the fit behaves as though the amplitude carried 27–29 % — a figure recoverable three independent ways, including from the paper's own expected and observed limits. The 39 % is the right prior; quoting it without the profiled width invites the reader to rescale the limit incorrectly.

6 — the injection-test bias does not appear in Table 2

§6 quotes biases "of the order of 5 % of the injected value" at ΓH = 400 MeV, i.e. about 20 MeV, included in the likelihood as additional nuisance parameters. Table 2's smallest experimental subgroup is "Others" at 4 MeV (68 %). Either the bias is much smaller than the 5 % headline in the region that matters, or it is not in the table.

7 — "at most 17 %" is used to dismiss an improvement that is larger than most of the systematics

§4 rejects finer pT binning and the mass-measurement categories because they "improve the expected sensitivity by at most 17 %", on the grounds of "reduced reliability of the interference modelling". Seventeen per cent on a 250 MeV expected limit is 43 MeV, comparable to the entire experimental uncertainty. The reliability argument may well be right, but it is asserted rather than quantified, and the asymmetry — a number for the gain, no number for the cost — is conspicuous.

8 — small bookkeeping

The luminosity appears as 140.1 fb−1 in §3, "140 fb−1" on both figures, and the mH constraint is imported from a 139 fb−1 measurement; nothing tells the reader which was used for Table 1. Table 1 omits the non-interfering production modes that §3 puts in the fit — harmless for the sensitivity, as the toy shows, but it makes the table impossible to reconcile with a cross-section calculation without the check above. §7 quotes "approximately 22 (15) and 78 (61) times the SM width" using ΓSM = 4.07 MeV while the Asimov is built at 4.1 MeV.

Problems

Problem 1 — derive the √ΓH scaling from scratch
Show that at fixed observed signal rate the interference scales as √(μΓHHSM), starting from the Breit–Wigner. Then explain why the naive answer — that the interference is ΓH-independent — is also correct, and why the two statements do not conflict.

The naive answer first, because it is the one that is actually calculated. Take the resonant term of Eq. 2 and integrate over m2 in the narrow-width limit. The signal piece is |Msig|2/[(m2 − mH2)2 + ΓH2mH2], and ∫ dm2 of the Lorentzian is π/(ΓHmH). So S ∝ |Msig|2H ∝ cg2cγ2H, the familiar statement that a narrower resonance is a taller one. The symmetric interference piece carries an explicit ΓHmH in its numerator, which cancels the π/(ΓHmH) exactly: I ∝ cgcγ, with no ΓH at all. The antisymmetric piece integrates to zero but its amplitude is likewise ΓH-independent, since the principal value has no scale. So at fixed couplings, changing the width changes the signal and leaves the interference alone.

Now impose the experimental constraint. We do not observe the couplings; we observe a rate, and the rate is fixed to its measured value by floating μ. Setting S = μSSM means cg2cγ2H = μ cg,SM2cγ,SM2HSM, so cgcγ = cg,SMcγ,SM√(μΓHHSM). Since I ∝ cgcγ, we get I/ISM = √(μΓHHSM), which is Eq. 4.

Why they do not conflict. They are answers to different questions. "The interference is width-independent" holds at fixed couplings; "the interference grows as √ΓH" holds at fixed observed rate. The measurement lives in the second frame because the rate is data. The invariance the paper exploits — c → κc, ΓH → κ4ΓH — is exactly the statement that moving along the first frame's flat direction traces out the second frame's square root: under it, I → κ2I and ΓH → κ4ΓH, so I ∝ ΓH1/2.

What this buys the analysis, and what it costs. It buys total independence from the coupling values: any enhancement of ΓH, whether from invisible decays or from a rescaling of everything, moves the interference the same way. It costs a square root, and that is expensive. A limit on the amplitude translates into a limit on ΓH that is the square of it, which is why a measurement sensitive to a few per cent distortion ends up quoting a limit at 78× the SM width.

Problem 2 — how big would the dataset have to be to reach the SM width?
Using the toy's decomposition, estimate the integrated luminosity at which this method would reach ΓH < 10 MeV at 95 % CL. Do it twice: once with the theory uncertainty as published, once with it removed. Comment on which of the two answers is meaningful.

Set up the scaling. Statistical uncertainty on the amplitude scales as 1/√L, so σ0,stat(L) = 3.79 × √(140/L) MeV1/2. The systematic pieces do not scale. Using the parametrisation t = (√Γ − √Γ̂)2/(σ02 + f2Γ) with an expected Γ̂ = ΓSM = 4.07 MeV, the 95 % limit solves (√Γ95 − 2.02)2 = 3.84(σ02 + f2Γ95).

Target. Γ95 = 10 MeV means √Γ95 = 3.162, so the left side is (3.162 − 2.018)2 = 1.309, and we need σ02 + 10f2 = 1.309/3.84 = 0.341.

With theory as published. f = 0.271, so 10f2 = 0.734 — already more than 0.341 on its own, with σ0 set to zero. There is no luminosity that gets there. The multiplicative uncertainty on the interference normalisation imposes a floor: solving (√Γ − 2.018)2 = 3.84f2Γ with f = 0.271 gives √Γ(1 − 1.96 × 0.271) = 2.018, i.e. √Γ = 4.28, Γfloor = 18.3 MeV. Infinite data, 18 MeV — four and a half times the SM width.

With theory removed. Set f = 0 and keep only the constant terms. The experimental part of σ0 is √(5.592 − 3.792) = 4.11 MeV1/2, which also does not scale with luminosity. That alone gives 4.112 = 16.9 ≫ 0.341: still no solution. Push the mH constraint and the energy scale down as well — plausible over a long run, since mH improves with data — and only then does σ02 = 0.341 become reachable, requiring 3.79√(140/L) = 0.584, i.e. L = 140 × (3.79/0.584)2 = 5.9 ab−1. That is roughly twice the full HL-LHC dataset, in one experiment, with every systematic driven to zero.

Which answer is meaningful. Neither as a projection, both as a diagnosis. The honest reading is that this method is not a route to the SM width and was never going to be — it is three orders of magnitude away, and the square root in Eq. 4 means closing a factor of 78 in ΓH requires closing a factor of 8.8 in the amplitude. What the exercise does establish is where the wall is: the theory normalisation, not the data. That is precisely the paper's own conclusion, and this is the arithmetic behind its closing sentence about "motivating improved calculations of the interference".

Caveat, flagged as judgement. The floor calculation treats f as a fixed multiplicative constant. In reality a future NNLOsv-matched-to-parton-shower calculation would reduce f, and the floor moves as Γfloor ≈ ΓSM/(1 − 1.96f)2 — at f = 0.10 the floor is 6.3 MeV, at f = 0.05 it is 5.0 MeV. So the method's asymptotic reach is a direct, and quite steep, function of one theory number.

Problem 3 — why is the observed limit worse than the expected one, and by exactly how much?
The observed 95 % limit is 319 MeV against an expected 250 MeV. §7 attributes this to Γ̂H = 12.0 MeV being above the SM value assumed in the Asimov. Verify that quantitatively, and say whether a 12.0 MeV best fit is in any tension with the SM.

Verify the shift. Use t(Γ) = (√Γ − √Γ̂)2/(σ02 + f2Γ). From the expected pair (61, 250) with Γ̂ = 4.1: solving the two crossing conditions gives f = 0.291 and σ0 = 5.32. Now hold those fixed and move Γ̂ to 12.0. The 95 % crossing solves (√Γ − 3.464)2 = 3.84(5.322 + 0.2912Γ), which gives √Γ = 17.9, Γ = 320 MeV against the published 319. The 68 % crossing similarly gives 91 against 90. §7's one-sentence explanation is exactly right and the numbers close to within a MeV.

Now the tension question, which is more interesting. A best fit of 12.0 MeV against an SM of 4.07 MeV looks like a factor of three, and means nothing. The likelihood is parabolic in √Γ, not in Γ, and the relevant distance is (√12.0 − √4.07)/σ0 = (3.464 − 2.018)/5.41 = 0.27σ. The observed amplitude is a quarter of a standard deviation above the SM. There is no tension whatever, and the apparent factor of three is entirely an artefact of quoting a squared quantity.

Why this matters for reading such papers. Any measurement whose observable is an amplitude but whose reported parameter is that amplitude squared will do this: small upward fluctuations look like large parameter excursions, and limits degrade fast. It is the same arithmetic that makes the observed limit 28 % worse than expected from a 0.27σ upward fluctuation. A reader who sees "Γ̂H = 12.0 MeV, three times the SM prediction" and reaches for a physics interpretation has been misled by the parametrisation, not by the authors.

The judgement part. The paper reports Γ̂H = 12.0 MeV without an uncertainty and without a compatibility statement. Given how easy the misreading is, a parenthetical — "0.3σ from the SM prediction" — would cost one clause and remove the trap. Whether the omission matters depends on the audience; for a Phys. Lett. B letter that will be read by non-specialists, I would argue it does.

Problem 4 — design the experiment that would make this method competitive
Off-shell measurements give ΓH = 3.6+2.1−1.6 MeV, two orders of magnitude better than this letter's 319 MeV. Given that, construct the strongest defensible argument for doing the interference measurement at all — then construct the strongest argument against it.

For. The off-shell method measures a ratio of rates at two very different virtualities and converts it into a width using an assumption that the couplings are the same in both regimes. That assumption is not a technicality: it fails in exactly the models one would most want to test, including any with new states between 125 GeV and the 2mZ threshold that modify the gg→H loop off-shell, or with momentum-dependent effective couplings. A dimension-six operator with a derivative structure changes the off-shell rate and not the on-shell one, and the off-shell determination absorbs that into ΓH. The interference method has no such loophole: μ is profiled, the couplings cancel identically by the scaling argument above, and the sensitivity comes from a lineshape distortion measured entirely within the on-shell peak. It is a genuinely different systematic, and a 319 MeV bound with no assumptions is not made redundant by a 5 MeV bound with one.

There is a second, sharper version of the argument. The two methods disagree in a specific, testable way if the Higgs sector is non-minimal. Suppose ΓH is enhanced by decays to invisible states while the on-shell couplings are SM. Off-shell and on-shell rates are both unaffected in the ratio the standard method uses, so it returns the SM width; the interference method sees the enhanced width directly through the amplitude. A disagreement between the two is therefore informative, and you cannot have a disagreement without both measurements.

Against. The method is 78× from the SM and, as Problem 2 shows, has a hard floor around 18 MeV set by the NNLOsv normalisation — four and a half times the SM width, at infinite luminosity, with the theory error as published. So it will never measure ΓH; the best it can do is exclude gross enhancements. But gross enhancements are already excluded by direct searches for invisible and undetected decays, which reach branching fractions of a few per cent and therefore widths of order 4.07/(1 − 0.03) ≈ 4.2 MeV under mild assumptions. The window where the interference method is the unique probe — SM-like couplings, SM-like visible branching fractions, and a width between roughly 20 and 300 MeV, meaning a 98–99 % invisible branching fraction that direct searches somehow missed — is very narrow and quite contrived.

What would change the balance. A calculation of the gg-interference at NNLOsv matched to a parton shower, which the paper explicitly calls for. Dropping f from 27 % to 10 % moves the floor from 18 MeV to 6.3 MeV and the achievable HL-LHC limit into the range where the method starts constraining models rather than gross pathologies. That is a theory investment, not an experimental one, and it is the honest headline of this letter: the measurement is currently limited by a calculation that could be improved.

Which side I come down on, flagged as judgement. The "for" case is sound on complementarity and weak on practical reach; the "against" case is sound on reach and weak on the assumption-freedom point, because assumption-free bounds have a way of mattering later. On balance the measurement is worth publishing and would be hard to justify repeating without the improved calculation. The paper's own conclusion says almost exactly this, which is to its credit.

Problem 5 — the spurious-signal test, and what it cannot catch here
§5 chooses the background function by the spurious-signal criterion: the signal yield fitted on a background-only template must be small compared with its statistical uncertainty. Explain why that criterion, designed for a rate measurement, is not sufficient for this analysis, and what the analogue would be.

What spurious signal tests. Fit the full signal-plus-background model to a background-only template built from simulation or from a data-driven mixture. Any signal yield you get out is a bias injected purely by the mismatch between the true background shape and the analytic one. Requiring it to be small compared with the statistical uncertainty on the signal yield bounds the mismodelling bias on the rate.

Why that is the wrong quantity here. This analysis does not measure a rate — μ is profiled and, as the β = 0 exercise shows, a pure rate change carries no information about ΓH at all. What it measures is an odd distortion of the peak: an imbalance between the number of events just below and just above mH. A background function can be perfectly adequate on the symmetric criterion and still have a wrong local curvature, and a wrong local curvature under a peak is precisely a spurious shift. Concretely: the exponentiated second-order polynomial exp(p0m2 + p1m) has a fixed sign of curvature over the whole range, so if the true continuum has an inflection near 125 GeV, the fit will absorb it by moving the peak.

The analogue that should have been run. A spurious-shift test: fit the full model to a background-only template with the interference normalisation free, and require the fitted ΓH — or better, the fitted peak displacement — to be small compared with its statistical uncertainty. The paper comes close with the injection tests of §6, but those inject at ΓH = 400 MeV with one component taken from simulation, which tests the analytic signal and interference models against their templates; it does not test the background function against a plausible alternative continuum shape. The two failure modes are different and only one is covered.

How big could it be? Order-of-magnitude: the measurement is sensitive to a shift of about 40 MeV per unit √(Γ/ΓSM), and the statistical uncertainty on the peak position is roughly σ/√Neff with Neff ≈ S2/B ≈ 55722/118000 ≈ 260 in a ±2σ window, giving about 100 MeV. A background mismodelling that displaced the peak by even a third of that — 33 MeV — would move Γ̂H by roughly (33/40)2 ≈ 0.7 in units of ΓSM, small against the 12.0 MeV best fit but not negligible against the 4.07 MeV SM value. So the effect is probably genuinely small here; the objection is that the paper does not demonstrate it.

Judgement. This is the weakest of the critiques in this file, because the standard H→γγ spurious-signal machinery is well understood and the collaboration will have looked at more than the letter reports. But a letter whose entire result is a peak displacement should state a bound on spurious displacement, not only on spurious yield, and the published text does not.