New to these reports? Start here
  • Dotted-underlined words have a plain-English definition — hover or tap them. Every term is also on the glossary page.
  • "Null" means we found nothing, not that something broke. Most reports here are negative results, on purpose — knowing an idea doesn't work is the point.
  • Two questions get asked separately. First, is the effect real? Second, is it already priced into the betting odds? An effect can be completely real and still useless to bet on.
  • A "calibration" row is a self-check. It runs the same method on something already known to be true. If that fails, the whole report is unreliable — so it's reported alongside the findings.
  • If a confidence interval includes zero, the real effect might be nothing at all, so no claim gets made.

Four NFL Context Signals: Are They Real, and Are They Already Priced?

One-off study, not a model, not production. Four hypotheses the user raised as interesting: contract-year motivation, referee crew tendencies, injury-report signal, and backup-QB drop-off. Each is tested twice — Stage AStage AThe first question: is the effect real at all? Measured against what actually happened, ignoring betting markets.: is the effect real? and Stage BStage BThe second question: is the effect already reflected in the betting odds? An effect can be completely real and still worthless to bet, because the price already includes it.: is it already in the closing lineclosing lineThe final odds right before a game starts. It reflects everything the betting market knows, which makes it the hardest benchmark to beat.? — and every effect is converted into the units books actually price in (spread points, total points, win-probability, prop units).

Headline: one real and useful finding (injury practice status), two clean nulls (contract year, referees), one solid quantification with no edge (backup QB) — and one methodological catch that manufactured a fake 1.85-point edge before it was caught.

TL;DR

Hypothesis Market n Stage A — is it real? Stage B — mispricing [95% CI95% confidence intervalThe range the true value is plausibly in. If this range includes zero, we cannot rule out that the real effect is nothing at all.] Tradeable floortradeable floorThe minimum edge needed to overcome the bookmaker's cut. Anything smaller is real but unprofitable. Verdict
Contract year Props (PPRpoints per receptionA fantasy scoring format that awards a point for every catch, which raises the value of high-volume receivers./g) 375 players −0.095 SD [−0.199, +0.008] no historical prop lines — untestable 2.6 rec yds NULLnull resultA test that found nothing. "Null" is the starting assumption that there is no real effect; a "null result" means the data gave us no reason to abandon that assumption. It does not mean the data was missing or the test failed to run. (point estimate negative)
Referee crews Total 2,892 games, 42 refs split-halfsplit-half reliabilityA repeatability test: split the data in two (here, odd vs even seasons) and check whether the same units score similarly in both halves. If a trait does not repeat, it was noise, not a trait. r = −0.09 permutationpermutation testRandomly shuffling the labels thousands of times to see what results pure chance produces, then checking whether the real result stands out from that. p = 0.83 0.79 pts null resultnull resultA test that found nothing real. The starting assumption is "there is no effect here", and a null result means nothing in the data argued against it. we are confident in, because the same method demonstrably CAN find a real effect when one exists (see calibrationcalibrationA deliberate sanity check on the method itself: run it on something already known to be true. If it fails to detect the known thing, the method is broken and its other results mean nothing.). Distinct from simply not having enough data.">DECISIVE NULLdecisive nullA null result we are confident in, because the same method demonstrably CAN find a real effect when one exists (see calibration). Distinct from simply not having enough data.
Referee crews Spread same split-half r = +0.05 permutation p = 0.61 0.76 pts DECISIVE NULL
Injury: practice status Props / usage 4,653 Questionable 31.4 pp play-rate spread no historical prop lines REAL & LARGE
Injury: team burden Spread 2,895 games −0.46 [−0.84, −0.08] 0.76 pts Detectable, below floor
Injury: team burden Total 2,895 games +0.31 [−0.09, +0.70] 0.79 pts NULL
Backup QB Spread 1,997 games −3.41 pts lost vs −3.64 line moves +0.23 [−1.39, +1.91] 0.76 pts Correctly priced
Backup QB Total 1,997 games −4.24 lost vs −2.31 line −1.93 [−3.07, −0.53] 0.79 pts Failed holdoutholdoutData deliberately set aside and never looked at while developing an idea, then used once at the end as a fair test. Peeking at it first would defeat the purpose. → dead
Backup QB Moneyline 1,996 games +1.34 pp [−4.97, +8.10] 2.38 pp NULL
(calibration) result ~ spread_line slope 2,895 1.043 (expect ~1) ✅ signs OK
(calibration) moneyline devigdevigRemoving the bookmaker's cut from odds to recover the market's actual implied probability. vs actual 2,893 0.554 vs 0.544 ✅ devig OK
(calibration) prior-season → this-season z(PPG) 963 r = 0.489 ✅ H1 pipeline OK
(calibration) home-team split-half, raw totals 33 teams r = +0.380 ✅ H2 harness OK
(calibration) report_status='Out' → play rate 3,491 0.0% ✅ H3 join OK
(calibration) result ~ backup start 1,997 −3.41 pts ✅ H4 join OK

Every hypothesis carries a calibration row proving the pipeline finds signal when signal exists. That is the difference between "we found nothing" and "our code was broken," and this project has been burned by the latter before.


⚠ The most important result is a methodological one

H4 initially produced a −1.85 point spread edge — more than double the tradeable floor. It was entirely an artifact. Worth reading even if you skip everything else.

The starting QB was identified, reasonably enough, as the QB with the most pass attempts that game. That silently admits two situations the market could not possibly have priced:

  1. Garbage time — in a blowout the backup out-throws the starter (confirmed real: 2023 wk18 PHI shows Mariota 20/36 despite Hurts starting).
  2. Early in-game injury — the starter is knocked out in Q1, the backup throws more, and the team predictably underperforms a line set for a healthy starter.

Both correlate mechanically with losing while being unknowable pre-kickoff. Requiring the identified starter to have thrown a dominant share of team attempts in both the current and prior game removes them:

Starter definition Points actually lost Points the line moves Apparent mispricing
Most attempts (contaminated) −5.18 −3.33 −1.85 [−2.65, −1.15]
≥70% of team attempts −4.26 −3.64 −0.62 [−1.70, +0.52]
≥85% of team attempts (clean) −3.41 −3.64 +0.23 [−1.39, +1.91]

The "edge" shrinks monotonically as contamination is removed and reverses sign once clean. Note the actual effect also shrinks (−5.18 → −3.41), because the contaminated version was partly measuring "games where the QB got hurt," not "games a backup started."

This is the same family of error as the NHL model's de-snooped +4.32% and the FRED biasbiasWhether the misses lean consistently one way. A projection can have a good average error size but still be biased if it is almost always too high. Bias is often the more fixable problem." data-def="Accidentally using information that was not available at the time. It makes predictions look brilliant and is the single most common way a backtest fools you.">look-aheadlook-ahead biasAccidentally using information that was not available at the time. It makes predictions look brilliant and is the single most common way a backtest fools you.: a signal that is only knowable after the fact will look like alpha. It cost about ten minutes to catch and would have been the study's headline finding if it hadn't been.


Setup: what counts as a tradeable effect

Computed from this data's own residual SDs (common_market.py), not assumed. At −110/−110 (4.55% hold) you must win 52.38%, i.e. beat coin-flip by 2.38 pp:

Market Residual σ Win rate per point Minimum effect to break even
Spread 12.73 pts 3.13 pp/pt 0.76 spread points
Total 13.19 pts 3.03 pp/pt 0.79 total points
Moneyline 2.38 pp of fair win probability

Bets needed to resolve an edge (80% power) — the 2% row independently reproduces the ~17,500 figure already in this project's memory:

True ROIreturn on investmentProfit as a percentage of the money wagered. +2% means $2 profit per $100 bet. edge Bets to detect
1% 76,927
2% 19,232
3% 8,547
5% 3,077

Data: 2,895 regular-season games, 2015–2025, with closing spread/total/moneyline ~100% complete from nflreadpy.load_schedules(). Sign-convention validator passes all six checks (result ~ spread_line slope 1.043, R² 0.197; moneyline devig 0.5539 predicted vs 0.5439 actual, buckets on the diagonal).


H1 — Contract year: NULL

Design. Within-player fixed effectsfixed effectsA design where each subject is compared against themselves, which cancels out their permanent differences. Here it means comparing a player to his own other seasons rather than to other players. on z-scored PPR points per game. This is the whole point: the dominant confound is selection — good players get extended early and so never appear in the contract-year pool — meaning any cross-sectional comparison measures player quality, not motivation. Each player is his own control.

Population: veteran multi-year deals only (yrs ≥ 2, excluding Drafted/UDFA/Practice and any deal signed in the player's draft year). Rookie-scale deals are perfectly confounded with the development curve — a rookie deal's final year is also the age-24 peak-improvement season — so they are excluded rather than "controlled."

Deduplication was mandatory, not hygiene: contract records are heavily duplicated (319,038 raw → 39,732 unique, 87.5% dropped). Trades emit one row per team. Correlated duplicate rows are exactly what inflated the NHL study's ROI.

⚠ 1,993 player-seasons excluded as ambiguous: OTC records year_signed as a year, not a date, so an extension signed in March is indistinguishable from one signed in November.

Results (n = 2,449 player-seasons, 579 contract years, 375 players contributing both a contract and non-contract season):

  • Primary: −0.095 SD, 95% CI [−0.199, +0.008] — CI spans zero, and the point estimate is negative (slightly worse in contract years).
  • Dose-response is non-monotone in the wrong direction. Mean z(PPG) by years remaining: 0 yrs +0.058, 1 yr +0.104, 2 yrs +0.151, 3 yrs +0.300. Players with the most time left perform best — that's the just-signed-a-big-deal selection effect showing up exactly where the FE design predicts it would, and it's why the cross-sectional curve can't be read as motivation.
  • Games played: −0.19 games [−0.45, +0.09] — no evidence contract-year players play through injury more.
  • Calibration: prior-season → this-season z(PPG) r = 0.489 (attenuated vs the preseason study's 0.773 by range restriction and within-position z-scoring, but clearly positive — pipeline finds real signal).

Exploratory, do not act on: WR contract years show −1.66 PPR pts/game [−2.76, −0.55]. That is 1 of 4 positions tested with n=57 players and no multiplicity correction; the split_miner harness found 96.6% of naive p<0.05 findings evaporate out-of-sampleout-of-sampleTested on data that was not used to build or tune the idea. This is the honest test; results on the data you built with are almost always flattering.. Flagged, not claimed.

Verdict: NULL. All three kill criteria met (CI spans zero, |estimate| < 0.10 SD, dose non-monotone). Powered to ~0.10 SD; published contract-year effects in other sports run 0.0–0.2 SD and mostly fail to replicate, so this rules out anything tradeable but is only marginally powered for a small psychological effect.


H2 — Referee crews: DECISIVE NULL

Design. Crew = the referee (white hat) only — the 7-man crew isn't a stable entity across seasons, and the referee is the unit bettors actually discuss. 42 referees with ≥30 games, 2,892 games joined.

⚠ Join gotcha: load_officials keys on the legacy YYYYMMDDXX game ID, not the modern 2023_01_DET_KC form. old_game_id in load_schedules is the only bridge.

Two tests, both structured to avoid multiple comparisonsmultiple comparisonsTesting many ideas at once. Test twenty things and one will usually look significant by luck alone, so results need a stricter bar. entirely — split-half reliability across odd/even seasons, and a within-season label permutation (one global test rather than 42 per-referee t-tests).

Split-half (odd vs even seasons) r
Referee → raw total points −0.093
Referee → market-adjusted total error −0.101
Referee → market-adjusted spread error +0.050
Home team → raw total points (calibration) +0.380
Within-season permutation Observed between-ref SD Null mean p
Raw total points 1.495 pts 1.793 0.945
Market-adjusted total error 1.516 pts 1.690 0.830
Market-adjusted spread error 1.606 pts 1.662 0.608

Referee tendencies do not persist across seasons at all — and observed between-referee variance is below the permutation null in every case. The same harness finds clear persistence for home-team identity (r = +0.380), so it detects real effects when they exist.

Notably this is more null than expected: the anticipated result was real variance in raw totals from non-random assignment (senior crews get primetime/dome games), washing out after market adjustment. Instead there's nothing even in raw totals.

Verdict: DECISIVE NULL. Don't model referee identity for totals or spreads.


H3 — Injury report: REAL, LARGE, and the one worth using

The only hypothesis that survives, and it isn't a betting edge — it's a usage and projection input.

⚠ Probe result: "Probable" was abolished after 2015 (2,553 rows in 2015, zero from 2016 on), exactly the vocabulary discontinuity to watch for. Kept as its own level, never merged into Questionable. Only 4 of 58,445 player-weeks have multiple rows, so the report is effectively one row per player-week.

Calibration (must reproduce the trivially-true): Out → 0.0% play rate (n=3,491), Doubtful → 0.7%, no designation → 84.5%. The join works.

Primary — within "Questionable," practice participation massively stratifies whether a player suits up:

Practice status Play rate n
DNP 38.8% 704
Limited 61.2% 2,994
Full 70.2% 906

31.4 percentage-point spread, monotonemonotoneConsistently moving one direction as the input increases. A real dose-response effect should be monotone; if more of the cause does not mean more of the effect, the pattern is suspect., against a 10 pp kill threshold. "Questionable" is not one status — it's three very different ones, and the practice line tells you which.

And it grades severity even among those who do play — production vs. each player's own trailing-4-game baseline:

Practice status Δ PPR pts vs own baseline 95% CI n
DNP −2.12 [−2.97, −1.27] 267
Limited −1.29 [−1.63, −0.96] 1,771
Full −0.78 [−1.32, −0.23] 610

All three CIs exclude zero and the ordering is monotone. A Questionable player who practiced fully still loses ~0.8 PPR pts; one who didn't practice loses ~2.1 — on top of being far less likely to play at all.

Team-level market test (stretch): spread error ~ injury-burden differential = −0.46 pts/player [−0.84, −0.08]. Statistically detectable but below the 0.76 tradeable floor, and computed with naive SEs (no bootstrapbootstrapRe-running a calculation on thousands of resampled versions of the data to see how much the answer wobbles. The spread of those answers becomes the confidence interval." data-def="A bootstrap that resamples whole SEASONS rather than individual games, because games within a season are related. Ignoring that makes results look more certain than they are.">season-block bootstrapseason-block bootstrapA bootstrap that resamples whole SEASONS rather than individual games, because games within a season are related. Ignoring that makes results look more certain than they are.), so the CI is optimistic. Totals: +0.31 [−0.09, +0.70], null.

Verdict: REAL for usage/props, not tradeable for team markets. This is directly usable by the fantasy projection engine and the draft/lineup tooling — see Bottom Line.


H4 — Backup QB: correctly priced, but the quantification is the deliverable

Framed from the start as primarily quantification, secondarily mispricing — with QB status known by Friday, "priced" was the expected answer, but "how many points is a backup worth" is a number this project's models can use regardless.

Using the clean starter definition (≥85% attempt share both games): 322 backup starts across 1,997 games.

Market Points actually lost Points the line moves Mispricing [95% CI]
Spread −3.41 [−4.88, −1.96] −3.64 [−4.77, −2.76] +0.23 [−1.39, +1.91]
Total −4.24 [−5.46, −2.84] −2.31 [−2.89, −1.54] −1.93 [−3.07, −0.53]
Moneyline +1.34 pp [−4.97, +8.10]

Spread: the market has this almost exactly right. Teams lose 3.41 points; the line moves 3.64. Residual +0.23, comfortably inside the floor.

Totals looked like a survivor and then failed the pre-registered holdout. The in-samplein-sampleMeasured on the same data used to build or tune the idea. Nearly always looks better than reality. −1.93 was the only Stage-B result to clear a floor, so per protocol it got exactly one holdout test:

Split Actual Line moves Mispricing [95% CI]
Develop 2015–2021 −4.58 −2.16 −2.42 [−3.58, −0.77]
Holdout 2022–2025 −3.65 −2.35 −1.30 [−2.90, +1.96]

Point estimate roughly halved and the CI now spans zero. Directionally consistent, statistically unconfirmed. Dead per protocol — no refitting on develop. Honest statement: consistent with either a modest real effect or noise; n=764 holdout games cannot separate them.

Exploratory: injury-driven backup starts (QB listed Out/Doubtful) cost −4.11 pts with mispricing −0.31 [−1.97, +1.03] — cleanly priced, as expected when the market has a full week's notice.

Power caveat: ~322 clean backup starts gives a spread CI half-width of ~1.65 points. We can rule out mispricing larger than ~1.7 points; we cannot resolve a 0.5-point mispricing. UnderpoweredunderpoweredNot enough data to detect an effect even if it is really there. An underpowered null means "we could not tell", which is very different from "there is nothing there". ≠ null.


What we could have detected

Two of four hypotheses are underpowered to resolve a tradeable-sized effect even if one existed. The NFL supplies only ~285 regular-season games a year; 11 seasons is 2,895 games, and that ceiling is the binding constraint on this entire class of study.

Hypothesis Effective n CI half-width Can it resolve the tradeable floor?
H1 contract year 375 players ±0.10 SD Yes for props
H2 referees 42 refs × ~49 games τ ≈ 0.8 pts Marginal — right at the floor
H3 (a) usage 4,653 Questionable ±2 pp Yes, comfortably
H3 (b) market 2,895 games ±0.38 pts Yes
H4 backup QB 322 backup starts ±1.65 pts No — can only rule out >1.7 pts

"Already priced in," defined

Formal null for every Stage B test: E[market_error | signal] = 0, where market_error is result − spread_line, total − total_line, or home_win − p_home_fair (multiplicative devig).

These are closing lines only. nflverse carries no opening lines for 2015–2025, so every Stage B result here tests priced at close, never priced at open. Nothing in this study can detect a stale-opening-line or line-movement edge. Since injury reports and crew assignments are published days before kickoff, closing-line tests are the appropriate — and hardest — bar for them.

Methodology guardrails applied

  • Season-block bootstrap (1,000 draws, resampling whole seasons) for H1/H4 CIs — handles within-season correlationcorrelationHow closely two things move together, from -1 (opposite) through 0 (unrelated) to +1 (in lockstep). It does not by itself mean one causes the other. and league-era drift, the direct antidote to the correlated-rows failure that inflated the NHL estimate. H3(b) uses naive SEs and is flagged as optimistic.
  • One pre-registered primary test per hypothesis; everything else labeled exploratory. No subgroup mining after a null primary.
  • No composites. Signals were never combined. The OSINT study showed k=50 known-zero signals lift in-sample SharpeSharpe ratioReturn relative to how much it bounced around. Higher means smoother returns for the same profit. 0.34→1.80 with OOS 0.00.
  • No threshold optimization. Cutoffs fixed a priori (grid-argmax has no skill here).
  • Develop 2015–2021 / holdout 2022–2025, holdout touched exactly once, only for the single hypothesis that survived Stage B in-sample (H4 totals — which then failed).

Caveats

  • No historical prop closing lines exist for 2015–2025 anywhere in nflverse, and this project's own [redacted] only starts Aug 2026. The prop "is it priced?" test is therefore impossible historically for H1 and H3 — Stage A effect sizes are reported against the vig-beating threshold instead, which is sufficient when the effect is smaller than the threshold.
  • OTC year_signed is year-granularity; 1,993 ambiguous player-seasons dropped.
  • Contract rows are 87.5% duplicated before dedup.
  • "Probable" exists only in 2015.
  • [redacted] is active-contracts-only (2,237 rows) and would introduce fatal survivorship biassurvivorship biasStudying only the things that stuck around. Looking at current players ignores everyone who washed out, which flatters the results.. This study uses nflreadpy.load_contracts() (51,952 rows incl. 48,962 expired). Don't use the local DB for historical work.
  • H4's clean starter definition costs sample (581 → 322 backup starts). That's the correct trade — the discarded games are precisely the unknowable ones.

Bottom line — effect on our modeling

Hypothesis ML Spread Total Props Action
Contract year none none none none (−0.10 SD, wrong sign) Don't model it.
Referee crews none none none none Don't model it.
Injury practice status none −0.46 pts (below floor) none large Use it — projections, not betting.
Backup QB priced priced (−3.4 lost / −3.6 priced) failed holdout useful Use the magnitude, don't bet it.

Two things are worth actually doing:

  1. Feed practice status into fantasy projections. A Questionable player is 38.8% / 61.2% / 70.2% to play depending on DNP / Limited / Full, and if he plays he's −2.12 / −1.29 / −0.78 PPR pts below his own baseline. Combining those gives a materially better expected valueexpected valueThe average result if you could repeat a bet or decision endlessly. Positive expected value means it pays off on average, even though any single instance can lose. than treating "Questionable" as one bucket — which is what the current engine effectively does. Directly relevant to lineup and start/sit decisions.
  2. Use −3.4 points as the backup-QB prior in any game-level model, and expect the market to already know it. For fantasy, the backup's offense scores ~4.2 fewer points, which should shade skill-player projections down.

Nothing here is a betting edge. The only Stage-B result that cleared a tradeable floor failed its holdout, and the one large real effect (practice status) lives in a market with no historical line data — and is almost certainly priced by anyone running props seriously. That's the expected outcome, and the study is worth the hour for the projection inputs and for the contamination lesson in H4.

On this page

Terms in this report

Source

backtests/nfl_context_signals/FINDINGS.md
updated 2026-08-18 21:46