benchmark scores have a shelf life
Contamination, saturation, and Goodhart pressure give every leaderboard figure a decay curve. Institutions should read benchmark scores the way they read market data — a number, a date, and a version, or nothing.
The memo that reaches your model-selection committee cites a leaderboard score to one decimal place. The score was measured on a model release two generations old, against a benchmark version that has since been revised, before the last training-data crawl — which almost certainly swept the test set into some corpus somewhere. Nobody in the room knows any of this, because the number arrives the way numbers arrive in institutions: stripped of its date, its version, and its conditions. Everyone treats it as current.
The same committee would never accept a bond price quoted that way. Someone would ask, before the sentence finished, as of when. A benchmark score deserves the same reflex, because it is the same kind of object — a perishable measurement with a decay curve. Three forces drive the decay, and they compound.
saturation turns ranking into noise
At the top of a well-used benchmark, frontier models cluster within a few points of the ceiling, and the gaps between them shrink below the run-to-run variance of the harness itself — the prompt template, the sampling settings, the parser that decides whether an answer counts. A gap that flips sign when you re-run the evaluation is not a ranking; it is noise dressed as one. When a scale is nearly exhausted, the honest reading is that the models are indistinguishable there — which is information, but not the kind a procurement memo wants to carry. The memo wants a winner, and the saturated benchmark will obligingly produce one, a different one each run.
a contaminated score measures memory
Test sets are published. Training corpora are crawled. The two meet — not through fraud, mostly, but through entropy: public text flows downhill into training data, and every crawl between a benchmark’s publication and a model’s cutoff raises the odds. Once the questions are in the corpus, the score stops measuring capability and starts measuring recall. The tell is a model that reproduces the answer key but fails a paraphrase of the same problem — it knew the answer, not the reasoning. Which is exactly the failure the harnesses we build from a firm’s own decision history are designed to catch, and a public leaderboard cannot: your past credit memos were never on the open web, so a model cannot have memorized its way through them.
the measure degrades because it matters
The third force is the oldest one. The moment a number drives purchasing decisions and launch announcements, it gets optimized directly — training data selected for adjacency to the test, prompts tuned against it, checkpoints chosen by it. None of this is cheating in any narrow sense; all of it quietly decouples the number from the capability it was built to proxy. The measure degrades precisely because it matters, and the better-known the benchmark, the stronger the pressure. The scores most likely to appear in your selection memo are the ones most degraded by having appeared in everyone else’s.
A benchmark score is a price, and nobody who trades on prices quotes one without a timestamp.
read a score like a price
The buyer’s discipline follows directly. A score is a triple — a number, a date, and a benchmark version — and a score missing either of the last two is not incomplete evidence; it is not evidence. Prefer evaluations with rotating or held-out test sets, where contamination has less time to accumulate. And for any figure that reaches a committee, insist on three answers: what exactly was measured, when, and against which model release — because “the model” in a vendor deck is a moving target that has usually moved since the measurement.
Then there is the reason stale numbers survive all of this: they once appeared in a deck. Decks outlive their evidence. The slide gets copied forward into the next quarter’s version, the caveat doesn’t, and a score measured on a retired model keeps circulating as institutional fact long after the leaderboard itself has moved on. The epistemic problem has an institutional memory problem underneath it, and the second is the one you can actually fix.
The fix costs one line. Put an as-of date and a version on every benchmark figure in every selection memo, and treat any figure without one the way a trader treats an undated price — as something you are not allowed to act on. A dated score is a measurement. An undated score is not evidence; it is nostalgia.