danieldeusing
visitor@danieldeusing:~/articles/ai-predictions-graded-final-ledger

cat ~/articles/ai-predictions-graded-final-ledger.mdx

Comunio World Cup 2026 · Part 9

The final ledger: everything my World Cup AI got right, got wrong, and actually earned

The World Cup is over: Spain champions, my AI-managed fantasy team first by 104 points — and this closing piece audits the whole experiment. All eight report cards revisited: 103 graded match tips (62.1% winners, a floor held by 0.14), 984 graded lineup calls (81.8%, cleanly calibrated), every bug and overcorrection, what the title actually earned, where the AI added real value and where it merely matched the market — and what I'd tell you before you run the same play on a problem of yours.

Jul 20, 2026 · · ~19 min read #ai #agents #football

The World Cup is over. Spain beat Argentina 1–0 in extra time — goalless through ninety minutes, decided by one substitute finishing another substitute’s pass — and my system’s last tip of the tournament was Spain by a goal, at a stated confidence of 44 percent: 2.3 points more optimistic about Spain than the betting odds were, with that gap justified in writing the way its final rule demanded. Winner right, margin right, scoreline missed by the width of an extra-time goal.

And the fantasy league this whole experiment was built around ended the same day: first place, 654 points, 104 clear of second, out of eleven managers.

This is the last piece in the series, and it’s the one I’ve been writing toward the whole time: the full audit. Every graded round, every bug, every rule the system wrote itself, what the experiment actually earned, where it wasted my time, and what I’d tell you if you wanted to run the same play on a problem of yours. The previous pieces graded the system round by round; this one grades the experiment.

One final repetition of the note that frames everything, freshly worded one last time: no model was built or trained here. I rented the same off-the-shelf AI models available to anyone, gave them tools, data and a job, and rewrote the job description after every graded round. That job description — a rulebook the system largely wrote about itself — turns out to be the most valuable thing the tournament produced. More on that at the end.

The series, in one place

Eight pieces led here. Each one was written before the next round was played, which means the record can’t be retrofitted — every claim sat in public before reality graded it. That was the point.

  • The setup: a squad of AI agents doing my fantasy World Cup homework — researching players, predicting who starts, tipping results, proposing transfers. It’s also where the practical questions live: what this costs to run, how it’s wired together, what data it touches.
  • Matchday one: first grades. Who-starts strong, result tips weak — 46 percent winners.
  • Matchday two: the tips got sharp — 75 percent — and I explained the scoring rebuild that made the grades trustworthy.
  • Matchday three: the data-trust crisis. I discovered I couldn’t trust my record of who had actually played, and spent the round scrubbing the scoreboard instead of celebrating it.
  • Round of 32: the round with the most exact scores — and the most expensive single tip of the tournament, a 76-percent favourite knocked out on penalties.
  • Round of 16: the worst round — 2 of 8 — taken apart into a format bug, an overcorrected safety rule, and fairly lost coinflips.
  • Quarterfinals: the best round — 4 of 4 — nearly erased by a second format bug that silently dropped every tip from the grading.
  • Semifinals: 0 of 2 on two honest coin flips, and the round I co-signed a postmortem that said “nothing needs to change.”

Then the final weekend of the tournament: the third-place game — tipped France 2–1 at 50, finished France 4–6 England, a ten-goal circus in which both coaches rested half their teams and the system’s rotation model got run over — and the final, which the system called correctly and quietly.

The full per-round record
The complete per-round table (“Pro Spieltag”; rows sorted by name, not date). In tournament order, winners called correctly: 46% → 75% → 71% → 73% → 25% → 100% → 0% → 0% → 100%. Below it, the system’s own closing lessons table (“Systemische Lehren”) — the rulebook this series kept writing, now marked as carry-forward knowledge.

The scoreboard, closed

The two halves of the system finished the tournament like this.

The flashy half — match results. 103 tips graded: 64 winners called correctly, 62.1 percent. Fifteen exact scorelines. The average miss on the winning margin: 1.17 goals. The season-long guardrail — overall winner accuracy must never fall below 62 percent — survived by 0.14 points, which is about as close as a floor can be held while still being held.

And the number this series chased hardest and never caught: confidence discrimination. Measured in percentage points, it’s the gap between the system’s average confidence on calls it got right and on calls it got wrong — a fire alarm should be louder for real smoke than for burnt toast, and the gap is how much louder. The target was a ten-point gap; the tournament closed at 8.14. Short, and the postmortem — the written self-review the system files after every round — finally named the structural reason instead of promising to try harder: by the system’s own banding, 46 of its 103 tips sat in the coin-flip zone, close to half the tournament, and there is no honest way to sound confident about a coin flip.

The quiet half — who starts. 984 graded start-calls across the tournament: right on 81.8 percent, and its best round arrived last — 90 percent across the final’s fifty graded start-calls, after the deep-knockout lineup rules from the semifinal piece went live. The number I’m genuinely proudest of needs one concept first.

Calibration table: the simplest lie detector a forecaster can face. Collect every prediction, sort them into bands by how confident they were, then check how often each band actually came true. A well-calibrated forecaster’s 30-percent claims come true about 30 percent of the time — no more, no less. It’s hard to game on a big sample: you can fake one confident call, but faking five bands agreeing with reality at once, over hundreds of calls, is close to impossible.

Here is that table for the who-starts model, over 984 calls: when the system said 0–20 percent, players started 12.6 percent of the time. At 20–40, they started 24.8 percent. At 40–60: 45 percent. At 60–80: 67 percent. At 80–100: 85.6 percent. Five bands, five honest numbers — a forecaster whose percentages mean what they say.

One disclosure this audit owes you, because it sits in the same data file: the start-call’s weaker sibling — will this player see any minutes, including off the bench — finished at 75.8 percent, with rougher calibration in the low bands. The transfer decisions keyed on the start call, which is both the sharper question and the better-behaved metric; but a closing audit that showed only its best table wouldn’t be one.

The who-starts calibration, final state
The closed who-starts book (“Aufstellungs-Kalibrierung”): 984 graded calls, every actual sourced from the public match record, stale predictions excluded. Each confidence band lands within a few points of what it promised — the closest thing to a clean bill of health this series can show.

What went well

The quiet half carried its share of the league — and I have to be careful about how big that share was. Fantasy points come from players who actually play, and a who-starts model right four times in five, feeding buy-sell decisions every round, kept every squad slot pointed at players who would actually be on the pitch — the precondition all other points stood on. But the first article in this series planted a distinction this audit has to honor: “will he be on the pitch” and “will he play well” are different questions. The system graded the first. Picking Messi in the first place — the performance side — never got a public scorecard, and a fair share of the 654 belongs to that ungraded judgment and to plain tournament luck. What I can claim is narrower and still worth having: the graded half made sure the ungraded bets were always on the field to pay off.

Humility turned out to be measurable. Across the whole tournament the system put 46 tips inside its coin-flip band — everything between 42 and 58 confidence, average 52. By the official scoreboard, 22 of them won: 47.8 percent, a hair under the coin-flip ideal. And the single tip separating that from exactly half is a familiar one — the Portugal call from the round of 16 that had the right winner in it and stays graded wrong because of the format bug, on the principle this series refused to bend. Even the band’s near-miss of perfection is honest. That is what “I don’t know, and I’m telling you so in numbers” looks like when reality grades it, and it held from the group stage to the final.

The endgame economics worked. In a knockout fantasy format, a star on an eliminated team earns nothing and loses value daily. From the round of 16 on, I converted the squad ruthlessly into survivors — the system recommended, I pulled the trigger — eliminated nations sold, money concentrated, a deliberate five-million overpay for guaranteed semifinalist minutes — until all five remaining players stood on the pitch in the final itself: three Spaniards, two Argentinians. Every point the closing games offered a five-man squad was a point mine was positioned to earn.

The final tip was the system at its grown-up best. Spain by one goal, 44 percent confidence — 2.3 points above the probability implied by the betting odds. Odds are percentages wearing different clothes: the price a bookmaker offers converts directly into how often the market thinks a side wins, and Spain’s price said about 42. The divergence was documented under the rule written after the semifinals. The clean-sheet logic called the goalless ninety minutes exactly. No forced loud call to make a dashboard look brave. If you’d shown me that tip before the group stage — market-anchored, margin-capped, reasoning attached, confidence honest to the decimal — I wouldn’t have believed this system wrote it.

The last tip on the board
The final in the evaluated-predictions view (“Bewertete Prognosen”, filtered to “F”): tipped 2–1, finished 1–0 after extra time — orange on the scoreboard, meaning winner and goal difference both correct. The confidence: 44, an honest number for an elite-against-elite final the market itself couldn’t split.

Grading every round in writing — knowing it would be published — did more than any single fix. I can’t run the counterfactual of a private version of this series, so take it as conviction rather than measurement: every round that went out with its misses attached made the next round’s homework better. Accountability wasn’t a nice-to-have on top of the system; it was load-bearing.

What went wrong — the honest catalogue

Four kinds of mistakes, in increasing order of how much they taught me.

Data I trusted and shouldn’t have. The matchday-three crisis: my record of who had actually played — the ground truth every grade depended on — was partly wrong, and I’d been grading against it for two rounds. Players logged as benched had played; an injury had been recorded as a rotation. The fix was dull and total: every actual sourced against the public match record, stale predictions excluded, suspensions counted. It cost a full round of work and it was the best investment of the tournament, because every number after it meant something.

Format bugs at the boundaries. Twice, the pipe between “what the predictor wrote” and “what the grader read” broke. In the round of 16, tips written in a self-contradicting format cost two correctly-called winners — and stayed graded wrong, because a record that contradicts itself is wrong no matter what was meant. In the quarterfinals, tips stored as the wrong data type silently vanished from the grading entirely, and the system’s best round briefly didn’t exist. Two different bugs, one lesson learned twice: validation belongs at the point of writing, it must check the container as well as the contents, and anything that reads records must fail loudly rather than skip quietly.

Rules that overshot, then rules for the overshoot. The 76-percent favourite that went out on penalties produced a hard confidence cap; the cap produced a round where every tip mumbled in the same narrow band and a France side the market had at five-in-six got tipped like a toss-up. Fixing overconfidence manufactured underconfidence. It took two more rounds of tuning — zones anchored to market prices, a loud end reopened only for genuine mismatches — before confidence carried information again. Every correction casts a shadow shaped like the opposite mistake, and I now budget for the shadow up front.

Patterns extrapolated across regime changes — stretches where the rules of the world quietly change while the model keeps assuming yesterday’s version of it. The system kept assuming the world it had measured was the world it was in. Coaches who rotated in the group stage stopped rotating in the semifinals — or reverted to group-stage lineups precisely when it mattered. The third-place game inverted every seriousness assumption the knockout rules were built on: both coaches rested stars, motivation replaced ratings, and a 2–1 tip met a 4–6 scoreline. And one metric — the discrimination target — kept being chased long after the tournament’s structure had made it unreachable, until the honest move was to retire it in the open rather than force a fake loud call to feed it. Patterns expire. The expensive misses live just after the expiry date.

There’s also a mistake catalogue about the articles, kept by the review process that checked each piece before publishing: a caption that claimed something its screenshot didn’t show, a goal described as long-range that was actually finished from inside the box, a “four substitutions at once” that was actually three plus one. Each was caught before it shipped — but each one started as me trusting a summary instead of a source, which is the same disease the system had, in the author.

And one omission this audit itself almost repeated, which would have been the worst kind of mistake for an audit to make. The subsystem the very first article opened with — the intelligence file on my rivals’ bidding behaviour in the blind transfer auction — never got a public scorecard at all. The ledger behind it exists: 442 tracked purchases across the league, 85 percent of them above list price, a median premium around 18 percent. It quietly informed my bids all tournament. But I never built a grader for it, so no one, including me, knows how good it actually was. The pattern deserves naming, because it’s the sharpest lesson in this section: the two subsystems that got graded got relentlessly better; the one that didn’t got quietly ignored. Public grading pulls effort toward itself. That is its power, and it is also a bias to manage.

What it earned

The literal answer: a fantasy league title. 654 points, 104 ahead of second, first place at every checkpoint this series published. My gap to second place at each round I published it: 11 points after the group stage, down to 5 after the round of 32 — the closest it ever got — then 55, then 95, and 104 at the final whistle. The knockout rounds are where it broke open, because my squad kept playing while other squads’ players went home. Also first on squad value at the close. What that’s worth in money: nothing worth mentioning. Bragging rights in a league of eleven managers — a currency I happily accept.

The real answer has three parts.

A rulebook that outlives the tournament. The system ends the World Cup with a written, tested, self-authored operating manual: confidence zones tied to betting prices; special handling for ultra-defensive opponents and tired legs; hard checks that a tip is written in a readable format before it counts; lineup caps for the late knockout rounds; a doctrine for losing coin flips without panicking; and an honesty rule — stay within five percentage points of the betting market, or justify the gap in writing. In business clothes, it’s the document every team running a model should keep: when to trust it, when to override it, and what would make you switch it off. Its own closing note marks the file as carry-forward knowledge for the domestic season that starts next. The tournament is gone; the rules it paid for are not.

The final league table
The closed table (rival names masked — they’re real people): 654 points (“Punkte”), 104 clear of second, first on total worth (“Gesamtwert”). The gap to second place across the published checkpoints: 11 after the group stage, 5 after the round of 32, then 55, 95, and 104 at the end.

A set of engineering lessons that had nothing to do with football. Test the pipeline, not the patch. Silent failure is the expensive kind. Feed every metric a perfect day before you steer by it — simulate your best possible outcome and check which way the number moves. React to process failures, never to outcome noise — fix broken reasoning even when the result was fine, and leave sound reasoning alone even when the result hurt. Write the failure line before you need it — the 62-percent floor did more disciplinary work than any target ever did. Retire dead metrics in public. Every one of those was bought with a specific, documented mistake somewhere in the eight pieces above.

A track record. Not of being right — of being checkable. One hundred and three tips and almost a thousand lineup calls, all graded against sources, all published before the fact, misses louder than hits. I’ve sat in enough meetings where an AI system’s quality was asserted rather than shown to know what that record is worth.

So was it useful? The honest verdict

Split it the way the system splits.

The match-tip half: discipline, not edge. Sixty-two percent winners sounds respectable until you ask the uncomfortable question this series owes you: would just siding with the betting market’s favourite have done about as well? For the rounds where the market’s price sits logged as a clean number next to every tip — the quarterfinals onward, eight tips — the answer is plainly yes: the system disagreed with the market’s favourite exactly once, backing a draw over a slight France lean in the semifinal, and that tip lost. The earlier rounds carried the market in the written reasoning rather than as a clean number, but the pattern reads the same there: the tips leaned where the odds leaned. Its value was never an edge over the market. Its value was behaving well: staying calibrated, refusing fake confidence, documenting divergence, not chasing lost coin flips. That’s worth something — it’s what kept the flashy half from ever torching the league position the quiet half was building — but I wouldn’t hire a system like this to out-predict a betting market, and this series is the evidence of why.

The who-starts half: genuine value, for a structural reason. No public market with real money changing hands prices whether a Comunio-relevant squad player starts a group game — no crowd of professionals with skin in the game whose answer you can copy. The question is unpriced, data-rich, repeatable, and gradeable. That’s exactly the shape of problem where an AI system doing tireless daily homework earns its keep, and 984 graded calls at 81.8 percent with straight calibration is what earning it looks like. The general rule I take from this, and the single most transferable sentence in the series: point AI at questions nobody is already pricing, and grade everything. In a company, the unpriced questions look like this: which supplier will actually deliver on time next month; which invoice line has drifted from the quoted price; which machine gets the next unplanned stop. No market answers those for you — your own data can. Where a market or an expert consensus already exists, the honest ceiling is usually “match it with discipline.” Where none exists, the field is open.

And underneath both halves, the experiment’s quietest, most durable payoff: the operating discipline itself — public grades, adversarial postmortems, floors, calibration checks — would improve most systems I’ve seen in the wild, with or without any AI in them.

What I’d do differently, and what I’d tell you

If I ran this again, in order:

  1. Verify ground truth — the record of what actually happened — from the first day, not from the round it breaks. Grading against unverified actuals was the single most dangerous thing the system did, because it silently poisoned everything downstream.
  2. Fail loud from the first line of code. Both format bugs were survivable; their silence is what made them expensive. A pipeline step that swallows bad input and produces a cleaner-looking result is a time bomb with a pleasant face.
  3. Write the floors and the kill-criteria before the first prediction. The 62-percent floor and the “what would make me stop” lines did more for honesty than any post-hoc reflection, because they were written before anyone knew whether they’d be met.
  4. Expect the value in the boring half. The impressive-sounding capability drew all the attention and delivered market-level performance; the unglamorous one paid every week. In a company those halves have names too: the flashy half is the demo that impresses a meeting; the boring half is the invoice matching, the delivery-date tracking, the reconciliation nobody films. Budget attention accordingly.
  5. Let the system write its own postmortems — then read them adversarially. The rulebook works because it’s specific and self-authored; it’s trustworthy because a sceptical human co-signs it only after checking, and because twice in this series the right response to the postmortem was to refuse it.
  6. Keep samples humble. A knockout round is eight games, then four, then two, then one. Half the discipline of this series was attaching “and that means little on its own” to every number that deserved it — including the flattering ones.

The last entry in the system’s own lessons file isn’t about football at all. It marks the rulebook as the starting point for the domestic league season — same discipline, same public grading, a longer season and a transfer market that runs every week of the year. The scoreboard resets; what was learned doesn’t. That’s the trade this whole experiment was quietly about: one World Cup in exchange for a set of rules about data, confidence and honesty that I’ll be using long after the third-place game is forgotten. Spain are world champions, my league table says 654, and the most valuable artifact of the tournament is a text file of mistakes with rules attached. I’d make that trade again — and if there’s a repetitive, checkable, unpriced question sitting in your own work, it might be the same trade waiting for you. You know where to find me.