Field notes — August 2026

The market beat our model. Here is the sealed record.

Most forecasting operations tell you how good their model is. This summer we spent five days and 9,009 paid API credits finding out whether ours is worth anything at all, measured against the hardest benchmark that exists for football probabilities: the prices of a liquid betting market.

It isn’t. Not yet. The market beat our model by about 0.010 RPS across 217 tournament fixtures, and when we let a preregistered procedure choose the best possible blend of the two, it chose to be 95% market and 5% us — a mix that is statistically indistinguishable from ignoring our model entirely.

We are publishing that result with the same machinery we would have used to publish a win, because we committed to doing so before we knew the answer. This piece explains what we tested, how we made it hard to fool ourselves, what we got wrong along the way, and why a company whose product is forecasts would tell you any of this.

−0.010RPS, market advantage
217fixtures, three tournaments
95%market share of best blend
v1–v6hash-chained locks

Where we started

A good record raises a harder question.

Flamsteed ran a fully recorded live forecast of the 2026 World Cup: 104 matches, every probability issued before kickoff from a point-in-time information set, every number committed to a public repository before the result existed. That record was good. The model named Spain champion, priced her at 50.7% in the final she won, had all four semifinalists in advance, and called the modal outcome in 25 of 32 knockout matches. Its stated probabilities matched observed frequencies to within about 1.5 percentage points at every threshold we checked.

A good record raises a harder question. Calibration tells you a forecaster is honest with itself; it does not tell you the forecasts contain information someone else doesn’t already have. The betting market is the natural referee, because its prices aggregate the views of everyone with money at stake. So we asked the uncomfortable question directly: if we had anchored our forecasts to de-vigged bookmaker prices, would they have been more accurate?

How we made it hard to cheat

Retrospective tests are easy to rig, usually by accident.

You choose the sample after seeing results, tune the method until it works, and report the version that survived. We built the test to make each of those moves either impossible or visible:

  • Preregistration, sealed by hash. The analysis plan — the exact accuracy gate, the sample, the statistical family, even the tie-breaking rules — was frozen and bound into a chained sequence of cryptographic locks before the scoring run. Any later change to any locked document breaks the chain in public. There are six locks; each amendment is dated and explains itself.
  • Point-in-time market data, paid for. Every bookmaker snapshot was archived under its own content hash at a fixed cut-off — 08:30 UTC, at least 30 minutes before any kickoff — so no price seen by the test could postdate the information it was compared against.
  • Ninety-minute settlement. Bookmaker match odds settle at ninety minutes. A knockout match that finished level and was decided in extra time or on penalties is a draw for scoring purposes, and our pipeline refuses to score any fixture whose ninety-minute result it cannot verify.
  • A decision rule fixed in advance. The gate was written down before the data arrived: a mean improvement of at least 0.002 RPS with bootstrap support of at least 80%, both required, with the decision belonging to the gate and not to anyone’s judgment after the fact.

One honesty constraint outranks all of that. Because the tournaments in the sample had already been played, no amount of hashing can prove the analysts were blind to the outcomes. An external adversarial review made exactly this point, and it is correct. This result is therefore development-grade evidence, not confirmation. A confirmatory version — the same test, sealed before a tournament that has not happened yet — is being designed now. Its terms will be published in full before it runs, and if the design turns out not to be runnable at a sample size this programme can actually reach, we will publish that finding instead.

What the test found

The gate passed decisively — for the market.

The market-anchored forecasts beat our model by a mean of 0.01018 RPS, with bootstrap support of 0.995 against the 0.80 requirement. The preregistered gate passed decisively. And because the winning forecast was 95% bookmaker, passing the gate means something bracing: on these matches, the market outforecast us.

The texture matters as much as the headline. Our model won 42% of individual fixtures outright. It lost the aggregate through a small number of catastrophic misses — the ten worst fixtures account for roughly 88% of the whole deficit. The single worst was Cameroon against Brazil in 2022, where the market kept meaningful probability on the upset and our model did not, and Cameroon won. The clearest pattern we found, on an independent sample, is that the model’s losses concentrate in the matches where it disagrees most sharply with the market, in either direction. When it strays far from the price, it is usually the one that is wrong.

Two things the result does not say. It does not say the model contains zero information: pure market beat the 95% blend by 0.00018 with an interval straddling zero, so “no detectable difference” is the honest phrasing, and the comparison that would settle it was not part of this design. And it says nothing about betting. No edge is claimed in either direction, and nothing here is betting advice.

What we got wrong

The failure is more instructive than the result.

The test itself, run through the sealed pipeline, held up. Our follow-up analysis of where the model loses did not, and we think the failure is more instructive than the result.

The exploratory code was written outside the sealed pipeline, without tests, and an adversarial review by an independent model found it had mislabelled extra-time matches, dropped exactly the fixtures whose ninety-minute outcome was most certain, and understated its own uncertainty. We repaired it. A second review found the repair itself wrong in four further ways. The third computation is the one now published, together with every withdrawn claim and all three sets of numbers. One candidate finding cleared a conventional significance bar after the corrections, and we declined to certify it anyway, because the decision rule had been fixed after a near-miss was visible — and a rule adopted once you have seen the data cannot certify anything.

Discipline that lives in infrastructure survives; discipline that lives in intentions does not.

The pattern is the lesson. Every error occurred in unsealed, untested, in-the-moment analysis. The pipeline that carried the preregistered result — built slowly, tested, and reviewed — never produced a wrong number across five days and three independent examinations.

Why publish this

That sentence, with something at stake.

Our methodology page has said from the start that a negative result is published as a result, not hidden as a failed experiment. This is that sentence with something at stake.

It also happens to be the strongest evidence we can offer for the thing we actually sell. Flamsteed’s product is not a claim of superhuman accuracy — the market just demonstrated why no honest forecaster in this sport should make one. The product is a verified record: probabilities issued before the fact, archived immutably, scored against reality in public, with the losses reported in the same type size as the wins. Anyone can check every number in this piece — the preregistration, the lock chain, the verdict, the full correction history — in the repository’s decision log.

The model stays as it is. The forecasts you see are unchanged. And whatever becomes of the confirmatory test, it will be published either way — the design, the decision rule, the venue, and the result if there is one — because a test whose terms are settled after the fact is not a test.