Third flake found in the same hunt. `timing variance stays bounded across
mismatch positions` took two single measurements of 50 000 iterations each and
compared their ratio, which is one GC pause away from a false alarm — it failed
about one run in fourteen on a loaded machine.
Two things were wrong with the measurement, not the property:
The first loop through that code pays for JIT compilation the second one does
not, so whichever side ran first was biased upward. The old comment admitted it
— "allow 2x variance for JIT/noise" — which is a tolerance covering a
measurement artefact rather than the thing under test. There is a warm-up now.
And a single pair has no defence against a scheduler hiccup. Five interleaved
pairs compared by median throw the pause out instead of the property.
The threshold is unchanged at 3, deliberately. An early-exit compare takes
~256x longer to reach a mismatch in the last byte than the first, on every
round; nothing here weakens what the test catches.
Verified that it still catches what it is for: with `constantTimeEqual`
temporarily replaced by an early-exit loop the test fails, and passes again
with the real one restored. A security test nobody has watched fail is a
security test nobody knows works.
A security test that cries wolf under load is worse than none — it teaches
people to re-run until green, and then a real regression looks like the noise.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014489bKUtUEY1Zgs9xN9mt7