← Setlist predictions

394-Simulation KPIs

Rules run blind against every qualifying Phish show from 1994 to 2026 — 394 of them, the complete population rather than a sample. Each test predicts a show that is night three or later of a consecutive run, using only what was knowable before it: that year's play counts up to the day before, minus every song already burned earlier in the run. Rebuilt 29 July 2026, after the earlier 100-show sample was lost and proved impossible to reconstruct.

Simulations
394
complete population, not a sample
Released
Forbin v11.2
2,573 caught, 1.58× chance
Edge on a full pool
3.24×
vs 1.34× on a burned one
Depletion term
withdrawn
made calibration worse
Pick for #96
v11.2
v12 tail pre-registered alongside
Span
1994–2026
26 years, 7,788 songs
What a test is, and why these are honest

Every model is a deterministic rule executed blind — no hand-picking, no knowledge of the setlist. Play counts come only from shows strictly before the target date, so nothing leaks backwards. The burn list is the actual set of songs played on earlier nights of that same run. The panel was rebuilt on 29 July 2026 and is now the complete population — every show from 1994 to 2026 that is night 3 or later of a consecutive-night run, 394 of them, rather than a 100-show sample that could not be reproduced. Songs are keyed on phish.net song IDs, not display titles, which removes the era-alias problem at source: Taste That Surrounds and Taste are one song by construction.

What “Chance” means: the comparator is 20 blind picks from the identical eligible pool — pure arithmetic, not a model. Across the 394 shows that expects 1,626 hits; Forbin v10 and v11 got 2,573, Forbin v9 2,565, Forbin v7 2,538, Forbin v5 2,039.

Chart 1Total songs caught across 394 simulations7,880 picks per model against 7,788 songs played · Chance = 20 blind draws, not a model
Forbin v5Forbin v7Forbin v9Forbin v10/v11Forbin v12
Forbin v5
2,039
Forbin v7
2,538
Forbin v9
2,565
Forbin v10/v11
2,573
Forbin v12
2,568
Chance
1,626

The order survived the rebuild; the margins shrank. On the old 100-show sample the spread from Forbin v5 to Forbin v10 looked decisive. On the complete 394-case population the four frequency-based models sit within 1.4% of each other — 2,538 to 2,573 — and only Forbin v5's barbell is clearly worse. Forbin v10 and v11 are the same bar because v11 never changed the ranking; its depletion term only rescaled the odds, and that term is now withdrawn. The honest reading of six model versions is that plain theme-year frequency does nearly all the work, the rotation correction adds about 1%, and everything else has been noise or worse.

SidebarWhat Claude scores with no data at allthe one genuinely “untrained” number this project has · 2 shows, not 100

Before any of this existed, 20 picks per night were locked in from model knowledge alone — no API, no play counts, no burn list, no look at the setlist. That is the closest thing here to a raw “base Claude” baseline, and it is worth seeing next to Chance.

Forbin v10
18
Base Claude
10
Chance
8

Ten hits out of forty, against 7.7 for darts. Claude working from memory alone is roughly 1.3× chance — statistically indistinguishable from guessing. Night two was the giveaway: it scored 2 of 20, because six of its picks had been played the night before and it had no burn list to know that.

Everything this project has gained came from the data pipeline, not from the model's knowledge of Phish. That is the honest read, and it is why the bar above is labelled Chance rather than anything implying a model — the untrained-Claude line is a different, much lower thing.

Chart 2Occasion nights are a different game entirelymean songs caught per show, 394 cases split by occasion type
special occasions — NYE, Halloween, runs of 8+ nights24 shows
Forbin v5
7.96
Forbin v7
12.29
Forbin v9
11.96
Forbin v10/v11
12.29
Chance
3.54
ordinary shows370 shows
Forbin v5
4.99
Forbin v7
6.06
Forbin v9
6.16
Forbin v10/v11
6.16
Chance
4.17

Every model roughly doubles its catch on an occasion night — 12.3 songs against 6.2 — while chance falls from 4.17 to 3.54. That is the single largest effect anywhere in this project, and it is not the models being clever: occasion shows play more songs from a pool the band has deliberately kept deep, so the head of the distribution gets played more completely. Forbin v9 is the one model that gets worse on these nights (11.96 against Forbin v7's 12.29), because its rotation term penalises recently-played songs and a curated night does not avoid them. Switching that term off is the whole of Forbin v10, and it recovers exactly the lost ground — no more.

Chart 3Calibration, and why the depletion term was withdrawnpredicted hits ÷ hits that landed, across 394 shows. 1.00 is honest.
Forbin v7
0.99
Forbin v9
0.93
Forbin v10
0.95
Forbin v11 (1.60)
1.31
Forbin v11 (refit 2.14)
1.00

Aggregate calibration is not the test that matters, and this chart is why. Forbin v11 scaled every probability by how depleted the pool was. At the anchor published on 28 July it over-calls by 31%; refit to this panel it lands at 0.999 — apparently perfect. But per-show the picture reverses: mean calibration error is 0.674 with the depletion term against 0.500 without it. Scaling by depletion fixes the average by adding error to every individual show. And the anchor itself moved 1.60 → 2.14 when the panel grew from 100 cases to 394 — a constant that swings 34% on a change of sample is absorbing noise, not measuring anything. Withdrawn.

Chart 4Ten-pick ladder points10 down to 1 per show · 21,670 available across 394 shows
Forbin v5
8,796
Forbin v7
8,923
Forbin v9
9,037
Forbin v10/v11
9,034
Forbin v12
9,197

Forbin v12 leads, and it is the only genuinely new result in this rebuild. Ranking the head of the slate is what the ladder measures, and v12 gains 163 points over Forbin v10 by suppressing songs that are long past due as well as those played too recently — see chart 7. It holds out: with the multiplier fitted on half the panel and scored on the other half, v12 beats v10 on the unseen half by 11 hits, 65 ladder points and 13 top-ten hits. It is queued behind tonight's show rather than released mid-run.

Chart 5Top-ten precisionhits among each model's ten highest-confidence picks · 3,940 available
Forbin v5
1,440
Forbin v7
1,526
Forbin v9
1,546
Forbin v10/v11
1,547
Forbin v12
1,566

Under 40% for every model. Six in ten of even the most confident picks miss, and that number has barely moved across six versions — Forbin v5 to Forbin v12 spans 3.2 percentage points. The ceiling here is not a modelling failure so much as a fact about the band.

Chart 6Consistencystandard deviation of hits per show, lower is better
Forbin v5
2.64
Forbin v7
3.49
Forbin v9
3.41
Forbin v10/v11
3.44
Forbin v12
3.39

Forbin v5 looks steadiest only because it catches least — it cannot swing when it never reaches. Among the models that actually score, Forbin v12 is the calmest at 3.39, which is the same suppressor working: deep-vault picks are the high-variance ones.

Chart 7The rotation signal is an arch, not a slope — and it changes shape on occasion nightshit rate by rotation state ÷ that frequency band's own average · 394 shows, 45,000 candidates
ordinary shows370 shows
played too recently
due < 0.4
0.61
0.4 – 0.8
1.08
0.8 – 1.2
1.19
due now
1.2 – 2
1.30
long past due
2 +
0.61
special occasions — NYE, Halloween, long runs24 shows
played too recently
due < 0.4
0.98
0.4 – 0.8
0.94
0.8 – 1.2
1.24
due now
1.2 – 2
0.89
long past due
2 +
0.61

Two findings here, and the second one Vlad called before the data did.

1. Being long past due is as bad as being played last night. Forbin v9 models rotation as a one-sided penalty — recently played, therefore unlikely. The full curve is an arch: songs at 0.606 lift when they are more than twice overdue, almost exactly the 0.613 penalty for being played too recently. A song that has not appeared in far longer than its rate predicts is not owed a slot; it has fallen out of rotation. This is what Forbin v12 adds, and it is worth 163 ladder points.

2. On curated nights the recency penalty disappears. 0.981 against 0.613 — on an occasion show, a song played three nights ago is no less likely than one played thirty shows ago. That is the correction Vlad made when he restored Prince Caspian to the #95 slate on a 3-show gap, against the model's objection. Caspian hit. The far end of the arch survives on those nights (0.609), so deep-vault material stays suppressed even when recency stops mattering.

Chart 8The idea that failed: “shorter songs, but more of them”filler index = mean song-count of the shows a song appears in
special-occasion (28)
fil <20
66%
20–22
109%
22+
92%
ordinary (72)
fil <20
100%
20–22
106%
22+
74%

The index cleanly finds short songs — After Midnight 31.9, Auld Lang Syne 28.4, Moby Dick 27.0 at the top; Jam 15.8 and Evolve 17.5 at the bottom. But as a ranking signal it fails. The highest-filler band underperforms in both populations (0.92 and 0.74). The short novelties appear in long shows because there is room, but stay individually rare — frequency already prices that in. The real effect is the mirror image: on special nights the lowest-filler band, the long-form modern jam vehicles, is penalised at 0.66. Adding either to the model produced 418 hits against Forbin v9's 417 and a worse ladder. Rejected. Measured on the superseded 100-show sample and not re-run on the rebuilt panel — the conclusion was decisive enough that re-testing a rejected idea was not the best use of the rebuild.

Chart 9The finding that replaced the depletion termmodel hits ÷ chance hits, 394 shows grouped by how much catalogue was still eligible
160+ songs left
3.24×
130 – 160
2.45×
100 – 130
2.05×
80 – 100 ★
1.34×
under 80
1.13×

The model's edge collapses as the catalogue burns. With more than 160 songs still eligible it catches 3.24× what blind picks would; by the time the pool is down to 80–100 that is 1.34×, and below 80 it is barely better than drawing at random.

This is the exact opposite of what Forbin v11 assumed. v11 reasoned that a shrinking pool makes each survivor likelier, so it scaled every probability up. That is true of the individual song and false of the slate: the burn list removes precisely the high-frequency material a frequency model is good at calling, and what remains is a flat field of rare songs where ranking has almost nothing to work with. Depletion makes each pick likelier and the slate weaker at the same time, and only the first half was modelled.

The starred band is tonight. MSG #96 runs with 86 songs burned and 88 eligible. The 55 comparable shows in that band delivered a mean of 5.8 hits; the rebuilt slate for tonight independently sums to 5.9. Those two numbers were computed different ways and agree, which is the reason to believe them — and why the board now shows 5.9 expected hits rather than the 11.4 published yesterday.

Chart 10Blind test: every model re-run mechanically on this run, #92–#95theme-year counts, burn list from the run's earlier nights, no knowledge of the setlist · 80 picks per model
Songs caught, 80 picks across four nightsof 86 played
Forbin v12
36
v7 = v9 = v10 = v11.2
31
Forbin v5
28
Chance
12.5
Ten-pick ladder, same four nights220 available
Forbin v5
115
v7 = v9 = v10 = v11.2
113
Forbin v12
95

This is a real blind test, not the published slates re-scored. Each model was re-run from the theme year's play counts minus the run's burn list, with rotation state measured against all Phish history to that night. None of these four nights is in the 394-case panel, so the rotation multiplier Forbin v12 uses is genuinely out of sample here.

Forbin v7, v9, v10 and v11.2 return identical slates on all four nights. Not similar — identical. On a themed night every candidate is a 1990s song with a 2026-sized gap, so the recency penalty that separates v9 from v10 never engages. The four versions are one model on this kind of show, which settles a question this page had been answering with theory.

Forbin v12 wins the hit count and loses the ladder — 36 against 31, but 95 against 113. It beat or matched the others on hits on every single night (3 wins, 1 tie) and lost the ladder on three of four. It finds more songs and ranks its top ten worse. Most of the ladder gap is #94, the 16-song night cut short by a front-of-house outage; excluding it the ladder is 70 to 74.

Forbin v5's 12 hits on #92 is the best single night any model has recorded, and it then scored 4, 6, 6. That is what a four-show sample looks like.

The dataAll models on the rebuilt panel, 394 simulations7,788 songs played · ceiling 5,545 · chance expects 1,626
ModelHitsMeanSDLadderTop 10Calib.
Forbin v5mid-depth barbell2,0395.182.648,7961,4401.01×
Forbin v7top 20 by frequency2,5386.443.498,9231,5260.99×
Forbin v9frequency × rotation state2,5656.513.419,0371,5460.93×
Forbin v10 / v11identical ranking2,5736.533.449,0341,5470.95×
Forbin v12inverted-U rotation — candidate2,5686.523.399,1971,5660.93×
Chance 20 blind picks, not a model1,6264.13

Ceiling is 5,545, not 7,788. 2,243 of the songs played were never in the eligible pool — debuts, or a song's first appearance that year. Forbin v10/v11's 2,573 is 46% of what was reachable, and the four frequency models sit within 35 songs of each other across nearly eight thousand. Forbin v6 and Forbin v8 are absent: both are retired, and both need song metadata (cover status, Fish novelties, costume albums) that the rebuilt pipeline does not carry. Their old-panel results stand in the patch notes as history, not as comparable numbers.

What this changes

1. The panel was rebuilt, and every number on this page changed with it. The original 100 shows were a sample whose selection could not be reconstructed. This panel is the complete population — all 394 shows from 1994 to 2026 that are night 3 or later of a consecutive run — so it is reproducible, extensible, and 4× the evidence. It reproduced the MSG run's burn list, pool size, all four night lengths and all twenty of tonight's play counts exactly, which is the check that it is measuring the right thing.

2. Forbin v11's depletion term is withdrawn, one day after release. It fixed aggregate calibration and made per-show calibration worse — 0.674 mean error against 0.500 without it — and its anchor constant moved 1.60 → 2.14 purely because the sample grew. It never changed a ranking, so nothing is lost but the odds, which are now the plain knot values.

3. The model's edge collapses as the catalogue burns — 3.24× chance on a full pool, 1.34× on a pool the size of tonight's. This is the opposite of v11's premise and it is the most useful thing the rebuild found. A depleted pool makes each survivor likelier and the slate weaker, because the burn list eats exactly the material a frequency model is good at.

4. The rotation signal is an arch, not a slope. Long-past-due songs are suppressed as hard as recently-played ones (0.61 either end). Forbin v9 only ever modelled one side. Adding the far side is Forbin v12, which gains 163 ladder points and holds up on a clean 50/50 holdout. Queued behind tonight rather than shipped mid-run.

5. On curated nights the recency penalty vanishes — lift 0.98 against 0.61 on ordinary shows. Vlad made that correction by hand on #95 when he restored Prince Caspian on a 3-show gap; it hit. The measurement now agrees with him. His corrections stand at 6 for 6; the model's called shots at 0 for 8.

6. Six versions in, plain frequency still does nearly all the work. Forbin v7 to v12 spans 35 songs out of 7,788 — 1.4%. Barbells, cover quotas, reserved slots, filler indices, song families and depletion scaling have all been measured and all but one have lost. The gains that survived are small, one-sided, and about honesty rather than cleverness.

The pick for #96

Forbin v11.2 — but on weaker grounds than this page claimed yesterday, and the honest reason is that the two candidates cannot be told apart.

Three of the six models are not in the running. Forbin v5 is last on the panel by 500 songs and averaged 7 hits a night here against 7.75 for the frequency core. And Forbin v7, v9, v10 and v11.2 are the same model on a themed night — the blind test returns byte-identical slates for all four across all four nights, because every candidate is a 1990s song carrying a 2026-sized gap, so v9's recency penalty never engages. There is no choice to make among them. That leaves exactly one real question: v11.2 or v12.

And the two bodies of evidence answer it in opposite directions — on both metrics.

EvidenceHitsLadder
394-case panelall shows 1994–2026v11.2 by 5v12 by 163
Blind test on this run#92–#95, four showsv12 by 5v11.2 by 18
Panel, occasion shows only24 showsv11.2 by 5level

Every cell contradicts the cell above or beside it. That is not a close call to be adjudicated — it is the signature of a difference too small to measure with the data that exists. When the larger sample and the more relevant sample disagree in both directions, the correct read is that the two models are indistinguishable, and the tiebreak goes to the one already published and pre-registered. Switching models hours before a show on a four-show sample is the precise error the last two releases were written to correct.

Two things this test overturned, both of which were on this page yesterday. First, the claim that v10's value on curated nights is switching the recency term off: the term was never engaging on these nights in the first place. Second, and more directly, the argument that v12 could not be trusted here because a 1996-themed show in 2026 makes every song look "overdue" against a 2026 baseline. The blind test is exactly that structure, four times, and v12 did not break — it caught more songs on every night. That objection is now disproved and has been withdrawn.

So tonight settles it, and it is pre-registered. The two slates agree on 17 of 20 songs and 9 of the top 10. They differ in three tail slots: v11.2 takes Talk, Uncle Pen and Amazing Grace — gaps of 129, 349 and 942 — where v12 takes I Didn't Know, The Sloth and Guyute, gaps of 17, 18 and 8. The mechanism is clean enough to score: v12 prefers theme-year songs that are also current in 2026, v11.2 is indifferent to that. Both trios are published on the board tonight. Whichever lands more is the model, and it will be one honest observation rather than an argument.

Patch notes

Which panel each entry was measured on. Entries dated 29 July 2026 use the rebuilt panel — the complete population of 394 shows. Everything dated 24–28 July was measured on a 100-show sample that has since been superseded and could not be reproduced. Those older figures are kept as the project's history; they are not comparable with the numbers in the charts above.

Forbin v12validated, queued29 Jul 2026

The rotation signal is an arch. Forbin v9 modelled only the near half.

Forbin v11.2released29 Jul 2026

Withdraws the depletion term released the day before, and rebuilds the panel that failed to catch it.

Forbin v11 / v11.1superseded — measured on the old 100-show panel29 Jul 2026

A calibration release. The ranking is Forbin v10's plus one data fix; everything else here is about the numbers printed next to the picks.

Forbin v10superseded27 Jul 2026

Occasion-aware. Two changes, both confined to special-occasion nights — NYE, Halloween, Baker's Dozen and this MSG run.

Forbin v9superseded27 Jul 2026

Added rotation state to the frequency core, and fixed the odds.

Forbin v8retired — old panel25 Jul 2026

Forbin v7 plus three reserved slots: costume album, Fish novelty, most era-locked song.

Forbin v7the baseline that would not die25 Jul 2026

Top 20 of the eligible pool by theme-year play count. The whole model.

Forbin v6retired — old panel25 Jul 2026

Frequency head plus a cover quota and archetype buckets.

Forbin v5retired24 Jul 2026

Mid-depth barbell: top 8 plus twelve from the 20–40% depth band.