The measurement: adaptive draw allocation — two rollouts for everyone, two more where the early ranking is unreliable. Offline the rule is impeccable: a trigger built on winner instability and the point's own noise, frozen on five other levels, keeps 94% of the 2→4 saving at 69% of the cost against an independent reference. Online — nothing, twice: −503 on its own noise, −202 [−1438, +1068] on common random numbers, at +29% compute.
The cause: the criterion asked for a point above the line between oracle-2 and oracle-4, and the endpoints themselves are indistinguishable: o4−o2 = +254 [−919, +1441] on 32 paired seeds under fully shared noise. The medians differ twofold (3124 against 7106 — the distribution is bimodal around clearing the level); the means do not. The line was strung between two statistically indistinguishable points.
The verdict: pre-registering a criterion without first drawing the noise floor is a methodological mistake, recorded as one. Online at this scale certifies only effects the size of the planner itself (+3800 over the policy); the fine economics of draws is measurable offline, on the stored matrices. What remains are the instruments: the CRN harness (arms play byte-identical games until a genuine policy difference) and the offline calibration stand.