๐ช๐ธ Spain won. ๐
I called it in part 2c, back in June, at 35.3% โ the model’s boldest, most market-contrarian pick of this whole series. Two and a half weeks later, in part 2d, I was explaining why the model had changed its mind and put ๐ฆ๐ท Argentina in front instead. The trophy is now in Madrid anyway. This is the piece I promised on the day the series started: grading the homework, in full, now that there’s nothing left to simulate.
The whole tournament, one chart
Every market_comparison.csv this project ever wrote is sitting in Git history, so instead of describing the arc from memory I went and reconstructed it โ 42 forecast snapshots, pre-tournament through the final whistle, tracking the model’s own title probability for the teams that mattered most.
The numbers, at the three moments that count:
| Team | Pre-tournament | Post-group stage | Result |
|---|---|---|---|
| ๐ช๐ธ Spain | 35.3% | 24.7% | Champions |
| ๐ฆ๐ท Argentina | 23.0% | 31.8% | Runners-up |
| ๐ซ๐ท France | 12.7% | 22.2% | Lost in the semi-final |
| ๐ด๓ ง๓ ข๓ ฅ๓ ฎ๓ ง๓ ฟ England | 6.0% | 5.7% | Lost in the semi-final |
๐ฆ๐ท Argentina’s 31.8% wasn’t a blip โ it held up as the model’s top pick through the entire group stage and Round of 32, and most of the Round of 16 besides. The lead flipped back to ๐ช๐ธ Spain at one specific moment: their Round-of-16 win over ๐ต๐น Portugal, which sent ๐ช๐ธ Spain from the low-20s to 32.4% in a single update. ๐ฆ๐ท Argentina never regained top spot. Then the semi-final win over ๐ซ๐ท France sent ๐ช๐ธ Spain’s number past 50%, and it stayed there for the final. The model’s very first instinct from June โ ๐ช๐ธ Spain, comfortably โ turned out to be right. It just took a three-week detour through being wrong to get there.
Grading the knockout stage
Thirty-one modelled knockout matches, from the Round of 32 to the final itself.1 The model picked the correct winner in 26 of them โ 83.9%. It’s more demanding to ask for the actual scoreline: the model’s top-5 most-likely scores for a fixture contained the real result 21 times out of 31 โ 67.7%.
Five matches went against the model’s pick:
| Match | Model favoured | Probability | Actual winner |
|---|---|---|---|
| R32 โ Germany vs Paraguay | ๐ฉ๐ช Germany | 66.5% | ๐ต๐พ Paraguay |
| R32 โ Netherlands vs Morocco | ๐ณ๐ฑ Netherlands | 62.5% | ๐ฒ๐ฆ Morocco |
| R32 โ Australia vs Egypt | ๐ฆ๐บ Australia | 59.5% | ๐ช๐ฌ Egypt |
| R16 โ Brazil vs Norway | ๐ง๐ท Brazil | 57.6% | ๐ณ๐ด Norway |
| R16 โ Switzerland vs Colombia | ๐จ๐ด Colombia | 63.3% | ๐จ๐ญ Switzerland |
None of the five is a blowout misprediction โ every losing favourite sat somewhere in a 57.6%โ66.5% band, never the 80%+ territory that would mean the model got genuinely fooled. That’s roughly what a well-calibrated forecaster should produce over thirty-one matches: a model that never got a 60/40 call wrong would be overconfident, not accurate.
The literal penalty shootout
This series has been promising to call its finale “the penalty shootout” since part 2d, and the actual final โ 1โ0, settled in extra time โ never needed one. But four knockout matches elsewhere in the bracket did, and the pattern is too clean to leave out: every single shootout of the tournament was won by the side Elo rated the underdog.2
| Match | 90-minute score | Model favoured | Shootout winner |
|---|---|---|---|
| ๐ฉ๐ช Germany vs ๐ต๐พ Paraguay | 1โ1 | ๐ฉ๐ช Germany (66.5%) | ๐ต๐พ Paraguay |
| ๐ณ๐ฑ Netherlands vs ๐ฒ๐ฆ Morocco | 1โ1 | ๐ณ๐ฑ Netherlands (62.5%) | ๐ฒ๐ฆ Morocco |
| ๐ฆ๐บ Australia vs ๐ช๐ฌ Egypt | 1โ1 | ๐ฆ๐บ Australia (59.5%) | ๐ช๐ฌ Egypt |
| ๐จ๐ด Colombia vs ๐จ๐ญ Switzerland | 0โ0 | ๐จ๐ด Colombia (63.3%) | ๐จ๐ญ Switzerland |
Four for four. The simulator, as documented back in part 2c, doesn’t model penalties as their own event โ level knockout matches resolve on the same Elo-implied coin as the rest of the tie. That’s an honest modelling choice, not a hedge, and this year the coin landed on the underdog every time it was actually flipped for real. Small sample, and I wouldn’t bet on the pattern repeating โ but for a series that promised you a penalty shootout, this is the closest the data gets to delivering one.
Model vs market: who won the argument
Disclaimer: This section discusses betting odds for the purpose of statistical comparison and analysis. It is not intended to promote gambling or serve as betting advice. Please gamble responsibly and be aware of your local laws and age restrictions.
The spine of part 2c was a single number: ๐ช๐ธ Spain’s +19.3pp edge over Polymarket, the strongest model-vs-market disagreement this blog had ever published. Part 2d found it had shrunk to +14.2pp as both sides moved toward each other. Here’s the full arc โ model’s title probability minus the market’s, for the same four teams, over the same 42 snapshots:
๐ช๐ธ Spain and ๐ฆ๐ท Argentina sit above the zero line for almost the entire tournament โ the model liked both more than the market did, from the very first snapshot to deep into the knockout stage. ๐ด๓ ง๓ ข๓ ฅ๓ ฎ๓ ง๓ ฟ England and ๐ซ๐ท France sit below it almost as consistently โ the market was the more optimistic side on both of them, and stayed that way as ๐ซ๐ท France’s own run to the semi-final pulled its market price up faster than the model’s. The numbers at the three moments this series called out:
| Snapshot | Model | Market | Edge |
|---|---|---|---|
| Pre-tournament | 35.3% | 16.0% | +19.3pp |
| Post-group stage | 24.7% | 10.5% | +14.2pp |
| Eve of the final | 55.8% | 58.2% | โ2.4pp |
By kickoff of the final, the argument had not just closed โ it had flipped. The market ended up more confident in ๐ช๐ธ Spain than the model was. The +19.3pp headline that opened this series didn’t survive contact with three months of actual football, but the side of it that turned out closer to reality โ “๐ช๐ธ Spain, undervalued” โ was the model’s, all along.
One for the historians
๐ช๐ธ Spain’s women won the World Cup in 2023, and the men’s side already had one from 2010 โ which already made them the second nation ever, after ๐ฉ๐ช Germany, to hold titles on both sides. What 2026 adds is the rarer thing: because ๐ช๐ธ Spain’s women are still the reigning champions when the men lift this trophy, ๐ช๐ธ Spain becomes the first nation in history to hold both titles at the same time. ๐ฉ๐ช Germany has won both โ four men’s titles, two women’s โ but never in the same window: their last men’s title (2014) came seven years after their last women’s title (2007), by which point the women’s crown had already passed to Japan.
And going into 2030 โ hosted jointly by ๐ช๐ธ Spain, ๐ต๐น Portugal and ๐ฒ๐ฆ Morocco, with three centenary matches in ๐บ๐พ Uruguay, ๐ฆ๐ท Argentina and ๐ต๐พ Paraguay โ ๐ช๐ธ Spain will be the first host nation in men’s World Cup history to enter as the defending champion. No prior host has ever managed it โ hosts are picked years ahead of time, the championship changes every four years, and the two have simply never lined up before now. The women’s side got there first, twice, and it hasn’t gone well either time. The ๐บ๐ธ USA were reigning champions from 1999 when they hosted 2003 โ and lost the final to ๐ฉ๐ช Germany. ๐ฉ๐ช Germany then won again in 2007, hosted 2011 as two-time defending champions โ and lost in the quarter-final. Two attempts, two defending-champion hosts, zero repeat titles. Maybe a bad omen for ๐ช๐ธ Spain in 2030.
Last thing, because it’s the best line of the whole tournament and nowhere else in this piece fits it: the bronze medal match, ๐ซ๐ท France 4โ6 ๐ด๓ ง๓ ข๓ ฅ๓ ฎ๓ ง๓ ฟ England, is the highest-scoring third-place game in World Cup history, beating ๐ซ๐ท France’s own 6โ3 win over ๐ฉ๐ช West Germany in 1958. Bukayo Saka scored a hat-trick. It’s ๐ด๓ ง๓ ข๓ ฅ๓ ฎ๓ ง๓ ฟ England’s best World Cup finish since they won the whole thing in 1966 โ and it isn’t even the headline of their tournament.
Final whistle โฝ
Three articles, ten million simulations each time the bracket reshuffled, and a model that spent the middle third of the tournament confidently backing the wrong finalist before its very first instinct turned out to be the right one. ๐ช๐ธ Spain are champions, 83.9% of the knockout bracket went the way the model said it would, and the +19.3pp argument that started this series ended up inverted by the time it mattered. That’s a good outcome for a rating system: not infallible, not embarrassed, just โ mostly right, for reasons that shifted under it the whole way through.
Thanks for reading along since June. See you for the next 48-team fever dream in 2030.
All the code, data snapshots and figures for this article live on GitLab. The full match-by-match grading table is here; the full 42-snapshot trajectory data is here.
The bronze medal match sits outside the simulator’s structured pipeline โ it isn’t part of the single-elimination bracket the model tracks โ so it’s excluded from the 31-match grading and covered separately, below. ↩︎
“Underdog” here means the side the model gave less than 50% to win the tie inside 90 minutes; the shootout itself isn’t separately modelled, so this is a real-world coincidence the simulator has no mechanism to have predicted. ↩︎