New to these reports? Start here
- Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
- "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
- Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
- A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
- If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.
Calibration Hygiene Campaign — Item 6 Synthesis
Date: 2026-07-12 | Scope: 5 tracks, adversarially checked | House framing: calibrationcalibrationA deliberate sanity check on the method itself: run it on something already known to be true. If it fails to detect the known thing, the method is broken and its other results mean nothing. is hygiene, not alpha — a change ships only if it clears a market-relative gate, never on ECE/log-loss improvement alone.
Headline
Of five tracks (six sub-results, since Track 3 splits MLB/NBA), the campaign produces one certified PASS (MLB EloElo ratingA rating system, originally from chess, that moves a team up or down based on results and the strength of the opponent. v2 beta calibration, shipped as hygiene with an explicitly disclosed marginal/noise-level effect size), two clean FAILs (NHL v3, MLB props), two BLOCKED (NBA Elo — no usable odds data to test; tennis — corrupted alert store + polling disabled + no underlying edge to protect), and one FLIPPED to needs-more-work by adversarial check (NFL win-head beta: passed its point-estimate gates but does not survive a bootstrapbootstrapRe-running a calculation on thousands of resampled versions of the data to see how much the answer wobbles. The spread of those answers becomes the confidence interval. the house rules require for a decisive production swap). This is exactly the outcome the pre-registered-gate discipline is designed to produce: four of five tracks close honestly negative/inconclusive, and the one adoption ships as hygiene, explicitly not touted as an accuracy win.
| Track | Experimenter verdict | Checker verdict | Final decision |
|---|---|---|---|
| NHL v3 (Platt/Beta) | FAIL | sound, confirmed | reject |
| NFL win-head (Platt vs Beta) | PASS | flawed — gate margins don't survive bootstrap | needs-more-work |
| MLB Elo v2 (Platt vs Beta) | PASS | sound, confirmed | adopt (hygiene) |
| NBA moneyline (bolt-on Beta) | FAIL | sound, confirmed | blocked (no usable odds data) |
| MLB props: K + hits (Beta) | FAIL | sound, confirmed | reject |
| Tennis Elo v1 | BLOCKED | sound, confirmed | blocked |
Track 1 — NHL v3 walk-forward Platt/Beta recalibration — REJECT
Gate: (a) ROIreturn on investmentProfit as a percentage of the money wagered. +2% means $2 profit per $100 bet. matches/beats raw, (b) volume ≥70% of raw, (c) Brier vs market does not worsen. Result: FAIL on (a) and (c) for both calibrators.
| Variant | Bets | ROI | Game-clustered 95% CI95% confidence intervalThe range the true value is plausibly in. If this range includes zero, we cannot rule out that the real effect is nothing at all. | Model Brier | Market Brier |
|---|---|---|---|---|---|
| raw (production) | 9,949 | +4.275% | [-1.98%, +10.48%] | 0.23633 | 0.23643 |
| Platt | 18,303 (1.84×) | -0.556% | [-6.90%, +5.74%] | 0.23870 | 0.23643 |
| Beta | 18,508 (1.86×) | -0.305% | [-6.11%, +5.59%] | 0.23718 | 0.23643 |
Raw model was already essentially at market Brier — no latent miscalibration to fix. Both calibrators shrink probabilities toward the population mean (Platt slopes 0.09-0.78 on logit(p)), which against a fixed absolute 3% EV threshold roughly doubles bet volume by pulling marginal/noise quotes over the line rather than reducing volume. Checker: full independent rerun bit-for-bit identical; leakage hunt (OOF pool break at k≥ts, per-game dedup, consistent clipping) found nothing; one cosmetic dead-comment issue only. Action: leave calibrate_model_prob() as identity.
Track 2 — NFL win-head Platt vs Beta — NEEDS MORE WORK (checker override)
Gate: (a) beta ECE improvement ≥1.0pp vs Platt on incumbent walk-forwardwalk-forwardEvaluating week by week using only what was knowable before each week, mimicking how the model would actually have been used at the time. OOS; (b) beta Brier/log-loss ≤ Platt's on 2025 holdoutholdoutData deliberately set aside and never looked at while developing an idea, then used once at the end as a fair test. Peeking at it first would defeat the purpose. vs devigged close. Experimenter: both PASS.
| Leg | Arm | n | ECE (pp) | Brier | Log-loss |
|---|---|---|---|---|---|
| 1 — walk-forward OOS 2018/19 | raw | 502 | 7.73 | 0.2163 | 0.6336 |
| Platt | 502 | 4.11 | 0.2146 | 0.6197 | |
| Beta | 502 | 2.88 | 0.2149 | 0.6220 | |
| 2 — 2025 holdout vs market | raw | 162 | — | 0.2218 | 0.6397 |
| Platt | 123 | — | 0.2183 | 0.6255 | |
| Beta | 102 | — | 0.2177 | 0.6236 | |
| market (devigged close) | 285 | — | 0.2108 | 0.6060 |
Beta ECE-improvement point estimate +1.23pp clears the 1.0pp bar; beta Brier/log-loss both ≤ Platt's. Checker's bootstrap (game-clustered, 3,000 resamples) using the experiment's own frozen probabilities: ECE-delta 95% CI [-1.93pp, +2.61pp] spans zero AND the 1.0pp gate; beta loses to Platt on Brier in 29% of resamples and on log-loss in 26%. Neither gate was ever bootstrapped by the experimenter despite the script defining an unused RNG_SEED "for any bootstrap CI." This is the same point-estimate trap already logged for NHL v3. Do not promote on this evidence; mechanical code changes documented in the promotion_spec are reusable once a properly bootstrapped re-run clears both gates.
Track 3 — MLB Elo v2 + NBA moneyline: global logistic vs beta — SPLIT
3a. MLB Elo v2 — ADOPT (hygiene only)
Gate: beats incumbent test log losslog lossA score for probability forecasts that punishes confident wrong answers harshly. Lower is better.; mid-bucket gap improves; market-relative conclusion unchanged. All 3 PASS.
| Arm | Test log-loss | Brier | ECE | Mid-bucket (40-50%) pred/act gap | Accuracy vs market (57.0%) | ROI flat @ close |
|---|---|---|---|---|---|---|
| raw | 0.68158 | 0.24437 | 0.02302 | 46.0% / 49.4% (+3.45pp) | 51.8% | -9.89% |
| Platt (incumbent) | 0.68149 | 0.24432 | 0.02151 | 46.0% / 49.4% (+3.36pp) | 51.8% | -9.89% |
| Beta | 0.68134 | 0.24429 | 0.02029 | 46.0% / 49.0% (+3.01pp) | 52.1% | -9.49% |
Bootstrap (2000 reps, cluster-by-game): beta-vs-Platt log-loss delta -0.00016, 95% CI [-0.00091, +0.00060] spans zero. Disclosed prominently by the experimenter and explicitly shipped as hygiene, not an accuracy claim — the market-relative gate (which house rules make binding) genuinely passes and the calibrator can only help (contains the identity map). Checker: bit-for-bit incumbent Platt reproduction (a=0.9731912514406493, b=-0.0034158760632627917), full leakage hunt clean, market-relative leg genuinely computed. Confirmed sound.
3b. NBA moneyline — BLOCKED (no usable data)
NBA's current trainer has no standalone calibration step — identity by construction — so beta was bolted on against a self-constructed causal split (val 2019-2023 / test 2024-2025, disclosed).
| Arm | Test log-loss | Odds-joined n | Model acc vs market | ROI @ close |
|---|---|---|---|---|
| identity (=incumbent) | 0.61784 | 62 | 66.1% vs 59.7% | +22.00% |
| Beta bolt-on | 0.61770 | 62 | 66.1% vs 59.7% | +22.00% |
The n=62 test-window odds sample (coverage gap: sbr_archive ends 2023-01-17, unified_odds.db starts 2026) inverts NBA's real full-sample conclusion (64.5% vs market 66.1%, ROI -2.44%, n=4,294) identically in both arms — a pre-existing small-sample artifact, not caused by beta. Gate fails on market-relative criterion as specified. Checker confirmed: genuine data-coverage gap, not a bug. No code change; NBA calibration cannot be honestly judged until real odds coverage exists for a valid test window.
Track 4 — MLB props (pitcher strikeouts, batter hits) Beta calibration — REJECT
Gate: market-relative log loss must materially close the gap to devigged close (≥25% of raw→close gap) with ROI not worsening; kill gate = beats raw on 2023-2025 anchor-line holdout.
| Market | Kill gate (holdout LL) | Market-window raw LL | Market-window beta LL | Market (devigged) LL | Gap closed | ROI raw → beta |
|---|---|---|---|---|---|---|
| Pitcher Ks (n=647) | 0.6818→0.6795 PASS | 0.7146 | 0.7269 (worse) | 0.6784 | -34.2% (widened) | -6.85% → -4.81% |
| Batter hits (n=7,614) | 0.6411→0.6393 PASS | 0.6703 | 0.6726 (worse) | 0.6691 | -183.5% (widened) | -7.76% → -8.64% |
Both markets pass the non-binding kill gate but FAIL the binding market-relative gate — beta makes real-market log loss worse, not just fails to improve it. Game-clustered bootstraps (2000 reps) confirm both degradations are real: K raw-minus-beta LL 95% CI [-0.0193, -0.0051] (excludes zero, beta confidently worse); hits CI [-0.0034, -0.0012] (excludes zero). This is the exact kill-gate-passes/market-fails trap the house rule exists to catch — a script working correctly, not a bug. Checker: bit-for-bit reproduction, independent incumbent re-derivation, leakage hunt clean, market leg genuinely computed. No promotion.
Track 5 — Tennis Elo v1 — BLOCKED
Not a calibration result — a data-availability failure. Live alert store ([redacted]) shows all 159 currently-settled tennis_elo_v1 rows with result='void' (zero usable outcomes), a corruption regression post-dating the 2026-07-02 dedup (pre-dedup backup shows the same row IDs correctly graded win/loss, 529W/1123L, n=1,652, matching the historical -12.3% ROI figure in memory). Both tours' odds polling has been enabled=0 since 2026-06-16 (WTA) / 2026-06-30 (ATP) — 26 days with zero new alerts through grass season. Even ignoring both problems, backtests/tennis_eval/FINDINGS.md already found the underlying edge unsupported (market beats Elo on log-loss in every slice; every ROI point estimate's bootstrap CI crosses zero) — there is no alpha here worth calibrating. Checker confirmed every DB-level figure and the FINDINGS.md quotes verbatim. Correctly, no calibration experiment was attempted. Two follow-ups flagged for the owner (not fixed, out of scope): locate/fix the void-corruption source; decide whether to re-enable tennis odds polling.
Net effect on item 6
- calibrate_model_prob() (NHL): remains identity. No change.
- MLB Elo v2 calibrator: recommended swap Platt→Beta, shippable as hygiene per the promotion spec above (owner action required — no production files were touched by any track).
- NFL win-head calibrator: stays Platt. Beta needs a properly bootstrapped re-run before reconsideration.
- NBA moneyline: no calibration layer added; blocked on odds-data coverage, not on the calibration method.
- MLB props (K, hits): stays uncalibrated (raw, params-tuned only). Beta rejected on real-market evidence.
- Tennis Elo v1: blocked on data hygiene (void corruption, disabled polling) independent of and prior to any calibration question.
Five of six sub-results close as honest negatives/blocks and one closes as a certified, explicitly-marginal hygiene adoption — no adoption was manufactured to make the campaign look more productive than the evidence supports.