← All findingssource · memory/project_wnba_week_counterfactual.md
Counterfactual replay run 2026-07-08 (scripts + data under session scratchpad wnba_cf_backtest/, RESULTS.md there). All replicas validated first: disciplined replica matched prod line_edges 1,547/1,547 cells and reproduced prod's real post-fix pick (Johannes u9.5, edge to 7 decimals); scout replay matched today's served heuristic board 95.8%.
Findings (one week — directional only):
- Scout board: actual (stale ML) 50.1% ±3.6 vs fixed-system replay 48.1% ±3.5 on 761 matched decided props — indistinguishable coin flips. Fixes were harm removal (side mix 59% vs 75% UNDER against 53% reality; honest ladder), NOT accuracy. The walk-forward backtest's 53.6%/63.7%-HIGH claims did not transfer (HIGH tier 48.4% n=128 this week). Treat all scout backtest credentials skeptically; 7.9% of board picks were DNP players (injuries stubbed in batch — still unfixed).
- Disciplined family: would have emitted exactly ONE pick all week (lost). Root cause is structural: at the 18:00Z fanout Pinnacle has both-sided mains on only ~6-17 props/game (quotes WNBA near tip) and the both_agree+sanity gates cut the rest. Fix = later/T-60 fanout slot for WNBA, not more model. pinn_devig replay 13-5 (+45.9% ROI) n=18 — noise-level.
- system_lineedges PRICED subset: 206 decided, 64.1% [57.3, 70.3], +17.1% ROI at capture prices — CI lower bound clears −110 break-even; the strongest signal in the entire investigation. The formula's headline 83.9% over 11,816 line-picks is a construction artifact (unpriced/synthetic lines) — only the priced subset is quotable. Proposed: a
priced_lineedgesformula variant (emit only book-priced exact lines) to test forward under CLV. - Pick gate counterfactual: would have dropped 71/93 published picks; drops settled BETTER than survivors (51.5% vs 35.0%; the no_edge-dropped subset hit 61.5%). The gate is an integrity/credibility fix, not an accuracy fix — expect no hit-rate lift from it. (All CIs overlap, n small.)
- [internal detail removed]
How to apply: don't cite pre-7/7 WNBA CLV rankings without the contamination caveat; volume fixes (fanout timing, priced-lineedges) rank above model work; demand forward CLV on the 64.1% before any promotion (multiple-comparisons discipline per feedback_model_edge_calibration). Related: project_wnba_audit_2026_07, feedback_devig_phantom_edge.