danieldeusing
visitor@danieldeusing:~/articles/ai-predictions-graded-semifinals

cat ~/articles/ai-predictions-graded-semifinals.mdx

Comunio World Cup 2026 · Part 8

Zero of two: the round where the right fix was no fix

The semifinals, graded in public: my World Cup fantasy AI called 0 of 2 winners — both games priced as coin flips by the market, both tipped as coin flips, both lost. Its postmortem concluded nothing needs to change, and after checking, I co-sign it. On telling outcome noise from process failure, retiring a KPI that can no longer be moved, and the starting-eleven pattern that expired.

Jul 16, 2026 · · ~16 min read #ai #agents #football

My AI predicted both World Cup semifinals wrong — and its postmortem, the written self-review it files after every round, landed on the one conclusion that sets off every alarm bell I own: nothing needs to change.

Here’s the score it’s defending. France–Spain, tipped as a 1–1 draw at 45 percent confidence, finished 0–2. England–Argentina, tipped 2–1 for England at 55, finished 1–2. Zero of two, one round after four of four. The quarterfinal article closed with both of those tips already locked, and a warning that the system itself expected to get roughly one of the two wrong — that’s what numbers like 45 and 55 mean. It did one worse.

“Nothing needs to change” is usually the sentence a system produces right before it fails bigger. I’ve spent this whole series being suspicious of my own tooling in public, so I’m not going to stop now: this piece is about how to tell a lazy “no fix needed” from an earned one — because this one, after checking, I co-sign. The checking is the article. One thing on scale before any of it: two games is the smallest sample this series has ever graded. Nothing in a two-game round can prove or disprove a system on its own — which is exactly why this piece is about the reasoning, not the score.

The standing note for anyone who joined mid-series: there’s no custom-trained model here — I haven’t built or trained any AI of my own. These are the same off-the-shelf AI models anyone can rent, given tools, context, and a job, with the job description rewritten after every graded round. What learns is the rulebook.

The recap, in case you’re new here

My fantasy World Cup team’s daily homework gets handled by a squad of AI agents: predicting who starts each match — that’s what wins the fantasy points — and predicting how each match ends, which is the part that gets graded loudest on this page. Every round, reality marks both, and I publish the marks. Six rounds are on the ledger, including a 2-of-8 disaster and a 4-of-4 near-perfect round back to back. This is round seven, and only two games were left to predict: the semifinals.

The semifinal scorecard, graded
The semifinals in the evaluated-predictions view (“Bewertete Prognosen”, filtered to “SF”): both tips (“Tipp”) against the final results (“Endergebnis”) — zero of two. The confidence column is the part that matters: 45 and 55, both games flagged by the system as coin flips before kickoff.

What actually happened, in two games

France–Spain, tipped 1–1 at 45. The bookmakers could not pick a favourite. Betting odds are probabilities in disguise — the price a bookmaker offers on a side tells you how often the market thinks that side wins — and converted that way, France sat in the low forties, Spain and the draw around 30 each. Both sides had built their runs on clean sheets — France’s streak intact, Spain’s broken exactly once, in the quarterfinal — and two defensively airtight elite teams meeting in a knockout is the classic script for ninety goalless-ish minutes and penalties. The system’s rulebook, burned once by exactly that script, tipped the draw. Spain didn’t read the script: an early penalty, then a sharp team move finished from inside the box just before the hour, 0–2. Nothing structurally surprising — a converted penalty and one well-worked move beat any defensive pattern.

England–Argentina, tipped 2–1 at 55. The market again saw a coin flip with a lean: England around 38 percent implied, the draw and Argentina each low thirties. The system leaned England — the favourite, but only just, at the cautious upper edge of what it still called a coin flip. Until the final quarter-hour the tip looked right — England led through Gordon from the 55th minute and controlled the game. Then Argentina’s coach went to his bench: three substitutions at once around the seventy-minute mark, the eventual match-winner following from the bench soon after, and the fresh legs turned the endgame — an equalizer at 85 minutes, a winner in stoppage time from a substitute. 1–2.

Two tips, two wrong winners, and the average miss on the winning margin — two full goals — was itself the worst of any round so far, on the smallest sample of the tournament. The postmortem’s verdict: no rule change. Here’s the case.

Why “no fix” is the right fix — and how I checked

Both semifinals were what the market itself called coin flips. Neither side of either game was priced above roughly 40 percent — no market, anywhere, was willing to claim it knew these outcomes. The system tipped one draw at 45 and one slight favourite at 55, which is to say: it told you, in advance, in numbers, that it didn’t know either. Then it didn’t know, in public.

Calibration: a forecaster’s numbers are promises about frequency, not about single games. A weather service that says “45% chance of rain” is not wrong when it rains — it’s wrong if, across all the days it said 45%, rain didn’t come about 45% of the time. Same here: a 45-percent tip that loses isn’t a broken tip; it’s the 55 doing what 55 does. The only way to never lose coin flips is to never call any — and then your forecaster has stopped forecasting and started hedging: refusing to commit so it can never be graded.

The system’s rulebook has a name for this situation now, written after a similar loss in the round of 16 and stress-tested twice this round: a lost coin flip triggers no rule change. The reasoning deserves to be quoted straight, because it’s the sanest line in the entire rulebook: lost coin flips don’t change the calibration — they are the calibration. The 45s and 55s only mean anything if some of them lose; start “fixing” every lost toss and you’ll chase noise backwards until the rulebook is a diary of coincidences.

And the doctrine is checkable, not rhetorical. Across the whole tournament the system has put forty-four graded tips inside its coin-flip band — everything between 42 and 58 — at an average stated confidence of 52. Twenty-two of them won. Exactly half. That’s what priced-right humility looks like over a real sample — and it’s also the honest caveat: a band that behaves like a fair coin validates the band, and says nothing about any single loss inside it.

There’s also a real failure line, and it’s public: the system carries a hard floor of 62 percent on overall winner accuracy, a guardrail it is not allowed to regress below. After the semifinals, that number sits at 62.4. So “nothing to fix” wasn’t written from a comfortable distance — it was written an inch above the line that would have forced a fix, and one more genuinely bad round breaches it. That’s what makes the verdict falsifiable rather than convenient: the failure condition exists, it’s written down, and it nearly fired.

But — and this is the part that makes “no fix” trustworthy rather than complacent — the postmortem didn’t stop at self-acquittal. It went looking for a process error anyway, and it found one. The England tip’s confidence was fine as a zone, but its distance from the market wasn’t. The market put England at 38.5 percent; the system said 55 — sixteen and a half points more sure than the crowd with money on the line, and the written reasoning behind the tip contained no actual reason for knowing better, just a description (“slight favourite at the top of the coin-flip zone”) dressed as one. If the tip had won, that gap would likely have sailed through unflagged. The audit flagged it in the round it lost and would have stood in a round it won — that’s the difference between grading outcomes and grading process. The rule it produced: from now on, a tip’s confidence must sit within five points of the market’s implied probability for the chosen winner, or carry an explicitly documented reason for the gap. The two acceptable states are “agree with the market” and “disagree in writing.” Vibes stopped being a third option.

That’s my test for co-signing a “nothing to fix”: the review that produced it must have been capable of finding something — proven by the fact that it did. And one structural note, because “the system reviewed itself” should raise an eyebrow: the grades come from the deterministic scorer, marking tips that were locked before kickoff; the postmortem is written against those grades; and my job in this series is to read it adversarially. This piece is that reading.

The round did expose a real blind spot — just not in the tips

While the match tips were busy losing coin flips honestly, the quiet half of the system — who starts each game — had a round that looked strong and hid a pattern.

The headline number was fine: right on about 83 percent of 91 graded start-calls. The misses are the story. Four players rated 80-plus to start didn’t start. Six players rated 35 or lower did. That’s not random scatter — every one of those misses is the same miss: at this depth of the tournament, coaches stop having a regular eleven. France swapped one winger for the like-for-like alternative it had been rotating with him all tournament. England reverted its entire flank defence to players who’d last started in the group stage, picked specifically against Argentina’s counterattack. Argentina benched a near-captain for a runner whose only job was to neutralize one England midfielder. Each eleven was a bespoke tool for one opponent, and a system extrapolating “who started the last round” was reading a pattern that no longer existed.

My favourite instance involves a player who has now been the hard case in every shape this series has met: France’s holding midfielder. A bad record of his supposed benching once broke my scoreboard’s actuals two rounds of articles ago; then the system correctly capped him at 35 after three straight genuine rotations; and against Spain’s possession midfield he walked back into the starting eleven and played all ninety — because for this one game he wasn’t a rotation option, he was the specific tool for the specific job. The cap was right for three rounds and wrong the round it mattered, for a reason that was visible in advance: the matchup.

The fixes are narrow, mechanical, and honest about uncertainty instead of pretending to resolve it. When two players are genuinely competing for one slot and the coach hasn’t spoken: both get rated in the middle, not 84 versus 10 — the system’s job is to report a coin flip where a coin flip exists, same doctrine as the match tips. A rotation cap gets lifted when the benched player is the obvious matchup tool for the next specific opponent. And nobody gets rated a near-certain starter this deep without a coach’s confirmation less than 48 hours old. Yes — those rules are written off a single round of semifinal evidence, and I said two sections ago that reacting to small samples is how forecasters ruin themselves. The difference between a rule and a curve-fit is that each of these carries a cause you can check before kickoff — a matchup, a fatigue record, a coach’s actual words — rather than a coincidence to extrapolate. If the causes don’t show up again, the rules never fire. The same lesson in both halves of the system, one round apart: honesty about what you don’t know beats confident extrapolation of a pattern that expired.

The who-starts scorecard and its misses
The who-starts calibration (“Aufstellungs-Kalibrierung”): each confidence band against how often those players really started — the 80–100% band hit 85%, the 0–20% band 14%, about four of five calls right across the tournament, every actual sourced from the public match record. The all-tournament outliers list (“Größte Ausreißer”) is topped by France’s holding midfielder, logged there for the opposite miss — rated 93 to start, rotated out. Both directions of the same lesson: the coach’s eleven is not a constant.

The target that ran out of road

One number this series has tracked since the group stage now needs an honest funeral, or at least a leave of absence — and I know how that looks. This target has been missed four rounds running: 7.67, then 8.45, then 7.95, now 8.3. Retiring it in the same round the system went zero for two is exactly the move a metrics-gamer would make. So here’s the mechanical case, checkable piece by piece — and if it doesn’t convince you, the numbers are all on the public dashboard.

Confidence discrimination — how much louder the system is on calls it gets right than on calls it gets wrong — has a standing target of a ten-point gap. Picture a fire alarm with a volume knob: it’s only useful if it’s louder for real smoke than for burnt toast, and the gap measures exactly that difference. It sat at 8.3 after the semifinals, target missed again — and now it’s stuck: the target is mechanically unreachable, no matter what the system does. Growing that gap requires loud calls, loud calls are only honest on lopsided games, and there are no lopsided games left. Two matches remain in the whole tournament — the third-place game and a Spain–Argentina final, elite against elite — and the markets price both as coin flips. On top of that, two tips can barely move an average built from more than a hundred graded calls. There is no legitimate action available that meaningfully moves the number.

There is an illegitimate one, and the shape is familiar: force a loud call anyway. Tip the final at 72 because the KPI — the key performance indicator, the number on the chart — wants a 72 in the dataset. The number goes up, the dashboard goes green, and the tip is a lie — the exact overconfidence this series spent four rounds beating out of the system, reintroduced to please a chart. The system’s postmortem instead did the unglamorous correct thing: it retired the target for the final, in writing, with the reason attached, and replaced it with targets the last match can actually measure — confidence within five points of the market or a documented reason, the write-time format checks green (the gate the quarterfinal bug forced into the pipeline), the new lineup caps applied. That format gate, while we’re on it, passed its first clean round since being rebuilt — both tips written correctly, nothing lost.

A goal that can no longer be influenced doesn’t measure performance anymore; it measures your willingness to game it. Retiring it in the open is the only move that keeps the rest of the scoreboard worth believing.

The full per-round arc, semifinals included
The per-round view (“Pro Spieltag” = per round; the table sorts rounds alphabetically) with the semifinals on the board: 2 games, 0 winners, an average margin miss of 2.00 (“Tordiff. MAE” — how many goals the tipped margin misses by, on average) — in the same table as the quarterfinals’ 4 of 4, the whole story of this series in two rows. Below it, the lessons table (“Systemische Lehren”): the last row is the coin-flip doctrine this piece is about — “lost coin flips don’t change the calibration.”

The league: ninety-five clear, and everyone still playing

The league table stands at first place, 526 points, ninety-five ahead of second — the post-quarterfinal state; the semifinal round hadn’t yet been added to the official table, so that figure is the floor going into the last two games, not the final word.

The squad behind it is the endgame shape the last article described, taken to its conclusion: seven players, every single one still playing. Three Frenchmen and England’s box-to-box midfielder go into the third-place game; my Spanish goalkeeper and both Argentinians — including the captain wearing ten — play the final. In a format where eliminated players earn nothing, a seven-man squad with seven men still on the pitch in the last two games of a 104-game tournament is the quiet half of this system compounding one more time. And the system’s research notes one more storyline it refuses to put a confident number on: the tournament’s top-scorer race has my number ten tied at the top with France’s striker — each with exactly one game left to settle it, in two different matches.

Strip the football out

The core skill of running any system that deals in probabilities is telling outcome noise from process failure — and reacting to exactly one of them. Outcome noise: the result went against you while the reasoning was sound. Process failure: the reasoning itself was broken, whatever the result. A sales forecast that missed because a customer went bankrupt mid-quarter is noise — no fix; the same miss caused by stale data in the pipeline is process — fix it immediately. The two look identical on the results slide, and you can only tell them apart by grading the reasoning, not the result. My system lost two coin flips and changed nothing about coin flips; it found a sixteen-point unexplained deviation from the market and legislated it away. Both decisions came out of the same review.

“Nothing to fix” is only credible from a review that was hunting. If your retrospective has never once found a problem in a round that went fine — or has never once declined to find one in a round that went badly — it isn’t reviewing, it’s narrating. The audit that clears you must be the same audit that could convict you.

Patterns have expiry dates, and the expensive misses live right after one. The starting-eleven model was extrapolating “recent rounds” into a phase of the tournament where coaches think in single games. The information it needed — matchup, coach statements, stakes — was available; the rulebook just hadn’t been told the regime had changed. Whatever your equivalent of “the regular eleven” is — the usual approver, the typical order size, the standard lead time — the question isn’t whether the pattern held last month. It’s whether the thing generating the pattern still exists.

And targets expire too. A KPI no one can influence anymore doesn’t measure performance; it measures willingness to game it. When one of yours reaches that state, the honest move is the one this system made — retire it in the open, reasons attached — because every quarter it stays on the dashboard, it trains people to fake it.

Two games are left in the entire World Cup. The first of the system’s two remaining calls is already public on my dashboard, new honesty rule in effect — the third-place game: France over England, 2–1, at 50, under three points from the market’s lean, gap documented. The final, Spain–Argentina, gets tipped under the same constraint, and if the market says coin flip, you’ll see a coin-flip number, KPI be damned. Then the tournament ends, the ledger closes, and I’ll write the final accounting: every round, every rule the system wrote itself, and whether a public scoreboard actually made the thing better — or just more honest about being wrong. If you’ve been running the numbers along with me, you already know those aren’t the same thing. One more round.