Rules run blind against every qualifying Phish show from 1994 to 2026 — 394 of them, the complete population rather than a sample. Each test predicts a show that is night three or later of a consecutive run, using only what was knowable before it: that year's play counts up to the day before, minus every song already burned earlier in the run. Rebuilt 29 July 2026, after the earlier 100-show sample was lost and proved impossible to reconstruct.
Every model is a deterministic rule executed blind — no hand-picking, no knowledge of the setlist. Play counts come only from shows strictly before the target date, so nothing leaks backwards. The burn list is the actual set of songs played on earlier nights of that same run. The panel was rebuilt on 29 July 2026 and is now the complete population — every show from 1994 to 2026 that is night 3 or later of a consecutive-night run, 394 of them, rather than a 100-show sample that could not be reproduced. Songs are keyed on phish.net song IDs, not display titles, which removes the era-alias problem at source: Taste That Surrounds and Taste are one song by construction.
What “Chance” means: the comparator is 20 blind picks from the identical eligible pool — pure arithmetic, not a model. Across the 394 shows that expects 1,626 hits; Forbin v10 and v11 got 2,573, Forbin v9 2,565, Forbin v7 2,538, Forbin v5 2,039.
The order survived the rebuild; the margins shrank. On the old 100-show sample the spread from Forbin v5 to Forbin v10 looked decisive. On the complete 394-case population the four frequency-based models sit within 1.4% of each other — 2,538 to 2,573 — and only Forbin v5's barbell is clearly worse. Forbin v10 and v11 are the same bar because v11 never changed the ranking; its depletion term only rescaled the odds, and that term is now withdrawn. The honest reading of six model versions is that plain theme-year frequency does nearly all the work, the rotation correction adds about 1%, and everything else has been noise or worse.
Before any of this existed, 20 picks per night were locked in from model knowledge alone — no API, no play counts, no burn list, no look at the setlist. That is the closest thing here to a raw “base Claude” baseline, and it is worth seeing next to Chance.
Ten hits out of forty, against 7.7 for darts. Claude working from memory alone is roughly 1.3× chance — statistically indistinguishable from guessing. Night two was the giveaway: it scored 2 of 20, because six of its picks had been played the night before and it had no burn list to know that.
Everything this project has gained came from the data pipeline, not from the model's knowledge of Phish. That is the honest read, and it is why the bar above is labelled Chance rather than anything implying a model — the untrained-Claude line is a different, much lower thing.
Every model roughly doubles its catch on an occasion night — 12.3 songs against 6.2 — while chance falls from 4.17 to 3.54. That is the single largest effect anywhere in this project, and it is not the models being clever: occasion shows play more songs from a pool the band has deliberately kept deep, so the head of the distribution gets played more completely. Forbin v9 is the one model that gets worse on these nights (11.96 against Forbin v7's 12.29), because its rotation term penalises recently-played songs and a curated night does not avoid them. Switching that term off is the whole of Forbin v10, and it recovers exactly the lost ground — no more.
Aggregate calibration is not the test that matters, and this chart is why. Forbin v11 scaled every probability by how depleted the pool was. At the anchor published on 28 July it over-calls by 31%; refit to this panel it lands at 0.999 — apparently perfect. But per-show the picture reverses: mean calibration error is 0.674 with the depletion term against 0.500 without it. Scaling by depletion fixes the average by adding error to every individual show. And the anchor itself moved 1.60 → 2.14 when the panel grew from 100 cases to 394 — a constant that swings 34% on a change of sample is absorbing noise, not measuring anything. Withdrawn.
Forbin v12 leads, and it is the only genuinely new result in this rebuild. Ranking the head of the slate is what the ladder measures, and v12 gains 163 points over Forbin v10 by suppressing songs that are long past due as well as those played too recently — see chart 7. It holds out: with the multiplier fitted on half the panel and scored on the other half, v12 beats v10 on the unseen half by 11 hits, 65 ladder points and 13 top-ten hits. It is queued behind tonight's show rather than released mid-run.
Under 40% for every model. Six in ten of even the most confident picks miss, and that number has barely moved across six versions — Forbin v5 to Forbin v12 spans 3.2 percentage points. The ceiling here is not a modelling failure so much as a fact about the band.
Forbin v5 looks steadiest only because it catches least — it cannot swing when it never reaches. Among the models that actually score, Forbin v12 is the calmest at 3.39, which is the same suppressor working: deep-vault picks are the high-variance ones.
Two findings here, and the second one Vlad called before the data did.
1. Being long past due is as bad as being played last night. Forbin v9 models rotation as a one-sided penalty — recently played, therefore unlikely. The full curve is an arch: songs at 0.606 lift when they are more than twice overdue, almost exactly the 0.613 penalty for being played too recently. A song that has not appeared in far longer than its rate predicts is not owed a slot; it has fallen out of rotation. This is what Forbin v12 adds, and it is worth 163 ladder points.
2. On curated nights the recency penalty disappears. 0.981 against 0.613 — on an occasion show, a song played three nights ago is no less likely than one played thirty shows ago. That is the correction Vlad made when he restored Prince Caspian to the #95 slate on a 3-show gap, against the model's objection. Caspian hit. The far end of the arch survives on those nights (0.609), so deep-vault material stays suppressed even when recency stops mattering.
The index cleanly finds short songs — After Midnight 31.9, Auld Lang Syne 28.4, Moby Dick 27.0 at the top; Jam 15.8 and Evolve 17.5 at the bottom. But as a ranking signal it fails. The highest-filler band underperforms in both populations (0.92 and 0.74). The short novelties appear in long shows because there is room, but stay individually rare — frequency already prices that in. The real effect is the mirror image: on special nights the lowest-filler band, the long-form modern jam vehicles, is penalised at 0.66. Adding either to the model produced 418 hits against Forbin v9's 417 and a worse ladder. Rejected. Measured on the superseded 100-show sample and not re-run on the rebuilt panel — the conclusion was decisive enough that re-testing a rejected idea was not the best use of the rebuild.
The model's edge collapses as the catalogue burns. With more than 160 songs still eligible it catches 3.24× what blind picks would; by the time the pool is down to 80–100 that is 1.34×, and below 80 it is barely better than drawing at random.
This is the exact opposite of what Forbin v11 assumed. v11 reasoned that a shrinking pool makes each survivor likelier, so it scaled every probability up. That is true of the individual song and false of the slate: the burn list removes precisely the high-frequency material a frequency model is good at calling, and what remains is a flat field of rare songs where ranking has almost nothing to work with. Depletion makes each pick likelier and the slate weaker at the same time, and only the first half was modelled.
The starred band is tonight. MSG #96 runs with 86 songs burned and 88 eligible. The 55 comparable shows in that band delivered a mean of 5.8 hits; the rebuilt slate for tonight independently sums to 5.9. Those two numbers were computed different ways and agree, which is the reason to believe them — and why the board now shows 5.9 expected hits rather than the 11.4 published yesterday.
This is a real blind test, not the published slates re-scored. Each model was re-run from the theme year's play counts minus the run's burn list, with rotation state measured against all Phish history to that night. None of these four nights is in the 394-case panel, so the rotation multiplier Forbin v12 uses is genuinely out of sample here.
Forbin v7, v9, v10 and v11.2 return identical slates on all four nights. Not similar — identical. On a themed night every candidate is a 1990s song with a 2026-sized gap, so the recency penalty that separates v9 from v10 never engages. The four versions are one model on this kind of show, which settles a question this page had been answering with theory.
Forbin v12 wins the hit count and loses the ladder — 36 against 31, but 95 against 113. It beat or matched the others on hits on every single night (3 wins, 1 tie) and lost the ladder on three of four. It finds more songs and ranks its top ten worse. Most of the ladder gap is #94, the 16-song night cut short by a front-of-house outage; excluding it the ladder is 70 to 74.
Forbin v5's 12 hits on #92 is the best single night any model has recorded, and it then scored 4, 6, 6. That is what a four-show sample looks like.
| Model | Hits | Mean | SD | Ladder | Top 10 | Calib. |
|---|---|---|---|---|---|---|
| Forbin v5mid-depth barbell | 2,039 | 5.18 | 2.64 | 8,796 | 1,440 | 1.01× |
| Forbin v7top 20 by frequency | 2,538 | 6.44 | 3.49 | 8,923 | 1,526 | 0.99× |
| Forbin v9frequency × rotation state | 2,565 | 6.51 | 3.41 | 9,037 | 1,546 | 0.93× |
| Forbin v10 / v11identical ranking | 2,573 | 6.53 | 3.44 | 9,034 | 1,547 | 0.95× |
| Forbin v12inverted-U rotation — candidate | 2,568 | 6.52 | 3.39 | 9,197 | 1,566 | 0.93× |
| Chance 20 blind picks, not a model | 1,626 | 4.13 | — | — | — | — |
Ceiling is 5,545, not 7,788. 2,243 of the songs played were never in the eligible pool — debuts, or a song's first appearance that year. Forbin v10/v11's 2,573 is 46% of what was reachable, and the four frequency models sit within 35 songs of each other across nearly eight thousand. Forbin v6 and Forbin v8 are absent: both are retired, and both need song metadata (cover status, Fish novelties, costume albums) that the rebuilt pipeline does not carry. Their old-panel results stand in the patch notes as history, not as comparable numbers.
1. The panel was rebuilt, and every number on this page changed with it. The original 100 shows were a sample whose selection could not be reconstructed. This panel is the complete population — all 394 shows from 1994 to 2026 that are night 3 or later of a consecutive run — so it is reproducible, extensible, and 4× the evidence. It reproduced the MSG run's burn list, pool size, all four night lengths and all twenty of tonight's play counts exactly, which is the check that it is measuring the right thing.
2. Forbin v11's depletion term is withdrawn, one day after release. It fixed aggregate calibration and made per-show calibration worse — 0.674 mean error against 0.500 without it — and its anchor constant moved 1.60 → 2.14 purely because the sample grew. It never changed a ranking, so nothing is lost but the odds, which are now the plain knot values.
3. The model's edge collapses as the catalogue burns — 3.24× chance on a full pool, 1.34× on a pool the size of tonight's. This is the opposite of v11's premise and it is the most useful thing the rebuild found. A depleted pool makes each survivor likelier and the slate weaker, because the burn list eats exactly the material a frequency model is good at.
4. The rotation signal is an arch, not a slope. Long-past-due songs are suppressed as hard as recently-played ones (0.61 either end). Forbin v9 only ever modelled one side. Adding the far side is Forbin v12, which gains 163 ladder points and holds up on a clean 50/50 holdout. Queued behind tonight rather than shipped mid-run.
5. On curated nights the recency penalty vanishes — lift 0.98 against 0.61 on ordinary shows. Vlad made that correction by hand on #95 when he restored Prince Caspian on a 3-show gap; it hit. The measurement now agrees with him. His corrections stand at 6 for 6; the model's called shots at 0 for 8.
6. Six versions in, plain frequency still does nearly all the work. Forbin v7 to v12 spans 35 songs out of 7,788 — 1.4%. Barbells, cover quotas, reserved slots, filler indices, song families and depletion scaling have all been measured and all but one have lost. The gains that survived are small, one-sided, and about honesty rather than cleverness.
Forbin v11.2 — but on weaker grounds than this page claimed yesterday, and the honest reason is that the two candidates cannot be told apart.
Three of the six models are not in the running. Forbin v5 is last on the panel by 500 songs and averaged 7 hits a night here against 7.75 for the frequency core. And Forbin v7, v9, v10 and v11.2 are the same model on a themed night — the blind test returns byte-identical slates for all four across all four nights, because every candidate is a 1990s song carrying a 2026-sized gap, so v9's recency penalty never engages. There is no choice to make among them. That leaves exactly one real question: v11.2 or v12.
And the two bodies of evidence answer it in opposite directions — on both metrics.
| Evidence | Hits | Ladder |
|---|---|---|
| 394-case panelall shows 1994–2026 | v11.2 by 5 | v12 by 163 |
| Blind test on this run#92–#95, four shows | v12 by 5 | v11.2 by 18 |
| Panel, occasion shows only24 shows | v11.2 by 5 | level |
Every cell contradicts the cell above or beside it. That is not a close call to be adjudicated — it is the signature of a difference too small to measure with the data that exists. When the larger sample and the more relevant sample disagree in both directions, the correct read is that the two models are indistinguishable, and the tiebreak goes to the one already published and pre-registered. Switching models hours before a show on a four-show sample is the precise error the last two releases were written to correct.
Two things this test overturned, both of which were on this page yesterday. First, the claim that v10's value on curated nights is switching the recency term off: the term was never engaging on these nights in the first place. Second, and more directly, the argument that v12 could not be trusted here because a 1996-themed show in 2026 makes every song look "overdue" against a 2026 baseline. The blind test is exactly that structure, four times, and v12 did not break — it caught more songs on every night. That objection is now disproved and has been withdrawn.
So tonight settles it, and it is pre-registered. The two slates agree on 17 of 20 songs and 9 of the top 10. They differ in three tail slots: v11.2 takes Talk, Uncle Pen and Amazing Grace — gaps of 129, 349 and 942 — where v12 takes I Didn't Know, The Sloth and Guyute, gaps of 17, 18 and 8. The mechanism is clean enough to score: v12 prefers theme-year songs that are also current in 2026, v11.2 is indifferent to that. Both trios are published on the board tonight. Whichever lands more is the model, and it will be one honest observation rather than an argument.
Which panel each entry was measured on. Entries dated 29 July 2026 use the rebuilt panel — the complete population of 394 shows. Everything dated 24–28 July was measured on a 100-show sample that has since been superseded and could not be reproduced. Those older figures are kept as the project's history; they are not comparable with the numbers in the charts above.
The rotation signal is an arch. Forbin v9 modelled only the near half.
Withdraws the depletion term released the day before, and rebuilds the panel that failed to catch it.
A calibration release. The ranking is Forbin v10's plus one data fix; everything else here is about the numbers printed next to the picks.
Occasion-aware. Two changes, both confined to special-occasion nights — NYE, Halloween, Baker's Dozen and this MSG run.
Added rotation state to the frequency core, and fixed the odds.
Forbin v7 plus three reserved slots: costume album, Fish novelty, most era-locked song.
Top 20 of the eligible pool by theme-year play count. The whole model.
Frequency head plus a cover quota and archetype buckets.
Mid-depth barbell: top 8 plus twelve from the 20–40% depth band.