New to these reports? Start here
- Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
- "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
- Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
- A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
- If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.
Winning a 30-Person Confidence Pool
Date: 2026-07-20 · Harness: pool_strategy_sim.py →
pool_strategy_sim_results.json · 6,000 simulated seasons per regime · Seed: 20260720
Pool format (confirmed): straight-up winners, all games, confidence N down to 1, 30 entrants. 2025 structure: 272 games, 18 weeks, season max 2,200 points.
The objective actually changed
FINDINGS_STRATEGY_2026.md maximized expected points and concluded "pick every
favorite" (~1,593 of 2,200). A pool pays only the maximum, so the real
objective is P(finish first). Those come apart: if every rival also picks
favorites, everyone scores the same number plus noise and P(win) collapses toward
1/30 no matter how correct you are. Being right in the same way as everyone else
wins nothing.
So the question is whether deviating — paying expected points to decorrelate from the field — buys win probability. That depends entirely on rival skill, which is unknown, so rival skill is swept rather than assumed.
Rivals pick the favorite with probability sigmoid(K·logit(p_fav) + public + idio)
and order confidence by a noisy p_fav. K=0 is a coin flip, K=1 is
probability-matching, K=3–5 is a realistic pool player, K=∞ is mechanical
always-favorite. The public term is a shared per-game lean, so rivals are
correlated with each other the way real pools are — omitting it would biasbiasWhether the misses lean consistently one way. A projection can have a good average error size but still be biased if it is almost always too high. Bias is often the more fixable problem. the
answer toward conformity. Outcomes are drawn from the devigged closing moneyline,
which is justified rather than assumed: over 2018–2025 mean p_fav = 0.6645
against an actual favorite win rate of 66.4%.
Result: differentiation is a trap — except in one regime
| Rival skill | Field mean | Best rival | baseline P(win) | flip-1 P(win) | Δ |
|---|---|---|---|---|---|
| coin flip | 1,111 | 1,263 | 100.0% | 100.0% | 0.0 |
| K=1 | 1,345 | 1,474 | 99.1% | 99.1% | 0.0 |
| K=2 | 1,464 | 1,565 | 80.8% | 79.3% | −1.5 |
| K=3 | 1,517 | 1,598 | 51.2% | 49.9% | −1.3 |
| K=5 | 1,552 | 1,613 | 23.3% | 22.2% | −1.1 |
| K=8 | 1,564 | 1,615 | 15.7% | 15.2% | −0.5 |
| always-favorite (noisy order) | 1,572 | 1,611 | 24.5% | 24.5% | 0.0 |
| CLONES (exact baseline, perfect order) | 1,597 | 1,597 | 3.3% | 39.8% | +36.5 |
Across every human regime, baseline maximizes P(win) and every contrarian variant costs win probability — aggressive ones catastrophically (flipping 2 games to top confidence drops P(win) to 0.7–7% at realistic skill). The intuition that a big pool rewards variance is wrong here, and the reason is worth stating: your rivals are not clones. They waffle on picks and order confidence badly, so they are already scattered below you. You do not need manufactured variance; you are differentiated by being correct.
The CLONES row is the exception and it is decisive. If all 29 rivals played exact baseline with perfect market ordering, playing baseline yourself guarantees a 30-way tie — P(win) is exactly 1/30 = 3.3%. Flipping a single game takes it to 39.8%.
Recommendation: baseline, plus one flipped tossup
Pick every Vegas favorite. Order confidence by devigged moneyline probability. Then flip your single closest-to-tossup game to the underdog and leave it at the bottom of the confidence ladder.
That one flip is close to free insurance:
- Cost: ~1 expected point out of ~1,597, and at most −1.5pp of win probability in any human regime.
- Gain: +36.5pp in the clone regime.
Maximum downside 1.5pp against maximum upside 36.5pp. You do not need to know which regime your pool is in to take that trade — which is the point of taking it.
On the reported ~1,100 winning score
That figure cannot be a winning season total under this format. A pool of 29 literal coin-flippers produces a best-of-29 winner of 1,263, and that is the floor — every less-random rival model produces a higher winner (1,474 → 1,615). There is no rival skill level anywhere on the sweep that yields an 1,100 winner.
What 1,100 does match, almost exactly, is the field mean of a coin-flipping pool (1,111) — an average score, not a winning one. The likely explanations, in order:
- A different denominator — the pool ran fewer weeks, or you joined late. A season max near 1,520 would put 1,100 at 72.4%, i.e. exactly the always-favorite baseline, which would mean the winner played the baseline and nothing more.
- Misremembered (it was given as "something like 1100").
- The pool genuinely is that weak — in which case picking every favorite wins it by roughly 400 points and nothing else in this document matters.
Worth checking the actual leaderboard, because branches 1 and 3 imply very different amounts of effort.
On finishing 13th and winning zero weeks
Zero weekly wins is expected, not a failure. With 30 entrants and 18 weeks, a perfectly average player wins zero weeks with probability (29/30)^18 = 54%. It is the single most likely outcome. Worse, the simulation shows baseline play wins ≥1 week only 8.7–37% of the time at realistic rival skill: with 29 rivals, someone spikes a lucky week almost every week, and the mathematically correct slate is rarely the week's high score. Weekly prizes are a variance game and the season prize is an accuracy game — they are in direct conflict, and chasing weekly wins costs ~37pp of season equity to buy ~22pp of weekly equity.
Finishing 13th is the real signal. Baseline play lands top-3 with 49–97% probability at realistic rival skill. Finishing 13th of 30 is very hard to reconcile with baseline play — it is what happens when you make your own picks and add noise to the market. The gap between 13th and 1st is almost certainly not a strategy subtlety; it is the cost of deviating from favorites at all.
Calibrate to the real leaderboard before trusting any P(win)
The P(win) figures above are only as good as the rival model, which is unvalidated. Two facts should temper them:
- P(win)=80% at K=2 is not credible. No 30-person pool has a strategy that wins four years in five; winners rotate. That number is a symptom of an implausibly weak rival model, not a real edge. Believable pools sit in the K=8-to-always-fav range where baseline P(win) is 15–24% — you should expect to lose most years even playing perfectly, because someone else always spikes a lucky season.
- The sim is partly circular. Outcomes are drawn from the market's
p_favand baseline picks byp_fav, so baseline is the true generating distribution and rivals can only add noise to it, never read a matchup better than the close. That is defensible (this platform has never found a sport where humans beat closing lines) but it is an assumption doing real work, and it inflates baseline's edge against any genuinely sharp opponent.
The discriminating diagnostic — a self-consistent pool where all 30 players share skill K — makes the regime checkable against a real leaderboard:
| Rivals | Winner | 5th | 13th | median | 1st–13th gap | P(anyone ≥ 1590) |
|---|---|---|---|---|---|---|
| coin flip | 1263 | 1188 | 1127 | 1115 | 136 | 0% |
| K=1 | 1475 | 1411 | 1359 | 1348 | 116 | 1% |
| K=2 | 1567 | 1516 | 1476 | 1468 | 91 | 33% |
| K=3 | 1598 | 1558 | 1525 | 1519 | 73 | 56% |
| K=5 | 1615 | 1584 | 1559 | 1554 | 56 | 67% |
| K=8 | 1615 | 1590 | 1569 | 1565 | 46 | 67% |
| always-fav | 1610 | 1591 | 1576 | 1573 | 35 | 64% |
How to read it. Always-favorite scores ~1593. So a pool whose winner tops out below ~1590 is a pool where nobody picks pure favorites — a soft pool (K≈1–2), in which baseline genuinely does dominate (the "too good to be true" result is true there). A pool where the winner is 1600+ and several players break 1590 is sharp: the top third is bunched within ~40 points, outcomes are luck-dominated, and baseline's edge shrinks to the modest 15–24% range. A winner near 1500 and a sharp field are mutually exclusive. The single fact that settles which world a pool is in: does anyone ever break ~1590 (≈72% of max)?
The recommendation (baseline + one tossup flip) is unchanged by this, because it is chosen to be robust across the whole regime span rather than optimized for one point on it. What changes is the honest advertised win probability: modest, not dominant, unless the pool is demonstrably soft.
Honest bounds
- The rival model is assumed, not measured. No data on actual pool-mate behaviour exists. This is why K is swept across the full range from coin flip to clone rather than fitted — the recommendation is chosen to be robust across the entire sweep, not optimized for one point on it.
- The clone regime is a bound, not a forecast. Real pools contain nobody who plays a perfectly market-ordered slate. It is included because it is the only regime where differentiation pays, and the recommendation hedges it cheaply.
- Ties are credited as 1/(co-leaders), appropriate for a split pot. If your pool breaks ties outright by the highest/lowest-scoring-team tiebreaker, the clone-regime baseline number is 1/30 either way.
- One season's schedule structure (2025) is used as the template.
- Not modelled: in-season adaptation. Late in a season, a player trailing the leader by a large margin should get more contrarian and a leader should mirror the field. That is a genuinely different (and better) strategy than any fixed rule here, and it is the natural next build if the pool matters.
UPDATE 2026-08-12 — the ~1,100 anomaly, resolved
The section above treats 1,100 as the winner's score and concludes no rival skill level produces it. That holds. But it assumed the wrong owner.
2025 season max is 2,200, and a coin-flipper scores exactly half of any max
— E[sum c_i X_i] = 0.5 * sum c_i, independent of confidence ordering. So:
| score | |
|---|---|
| coin flip (any ordering) | 1,100 |
| coin-flip pool, 13th of 30 | 1,116 |
| coin-flip pool, winner | 1,255 |
| always-favorite, 2025 realized | 1,578 |
1,100 is the coin-flip mean, and 13th of 30 in a coin-flip pool scores 1,116. The user finished 13th and reported ~1,100. Two independent matches, so the reading that fits every reported fact — the score, the 13th-place finish, and zero weekly wins — is that ~1,100 was the user's own score, not the winner's.
That is branch 3 of the three listed above ("the pool genuinely is that weak"), and it makes the rest of this document mostly moot: always-favorite scores 1,578 and beats a coin-flip pool's expected winner by 323 points. The user's 2025 picks were statistically indistinguishable from random, giving up ~478 points against simply taking every favorite.
SUPERSEDED 2026-08-12 by FINDINGS_REAL_POOL_2026.md. Real Yahoo exports for weeks 1-12 show the field picks favourites 68.5-92.1% of the time and averages 69.3% of max — nothing like coin flips. The reading above (1,100 = the user's score in a weak field) is WRONG; ~1,100 is the leader's running total around week 13. The 1,100-equals-coin-flip-mean match was a coincidence that happened to fit two facts at once. Weeks 13-18 exports do not exist, so the final standings and true winning score can no longer be established.
Method: calibration_test.py::load for 2025 outcomes, 20k-draw simulation of a
30-entrant pool at p_skill in {0.50, 0.60, 1.0}.