FINDINGS LIBRARY · 81 notes · measured, not remembered
What we measured.
The analyst may only quote what is written here: validated signals, refuted ideas, and the rules that came out of the mistakes. A finding that changes is a diff in this library, not a mystery.
Measured findings
What shipped, what was measured, and what the numbers said — including the refutations.
- autopsy-premortem-pairParlay autopsy (post-game) + slip pre-mortem (pre-game) endpoints in grading-service, plus the WNBA vs_opponent hydrator fix — DEPLOYED + E2E-verified 2026-07-30 (autopsy on slip 48; premortem live; 23/26 players carry vs_opponent).
- claude-combined-frozenclaude_combined freeze (7/18→8/02): BOTH halves fixed and deployed — advanced-check latch (f1a8509, 7/29) + retry slot pre-tip 23:45Z→21:30Z (33d623b, 8/02). First live test = the 8/02 21:30Z pass; verify last_pick_at moved.
- feature-experiments-measuredprojections-engine features that were built, measured via controlled A/B, and REVERTED for no lift — don't re-propose without new evidence
- project_adjusted_projection_auditDH adjusted_projection SCORED for the first time (2026-08-27): 13.9% worse than the market line, worse than a plain L10 mean on 8/8 stats, +0.01pp directional info over 'always bet the under'. H2H term is a UNITS-MISMATCH BUG — delete it (12 lines, halves the gap to Pinnacle).
- project_arm_promotion_bar_correctedThe 'positive units + positive CLV over 200 picks' promotion bar is WRONG — it promotes noise. Corrected bar + measured refutation of form_wnba and the deepseek arm (2026-08-28).
- project_brier_board_trackerscripts/brier_board.py is the forward accuracy tracker for WNBA props — Brier + MAE per slate with a write gate; first increment (Platt calibration) shipped, four others killed.
- project_brier_vs_vegas_closedBeating Vegas on WNBA prop Brier is closed — well-powered negative across six paths; market's total edge over base rate is only 1.31%; the minutes oracle was a post-hoc artifact.
- project_cv_rebound_tracking_spec8-agent research+design workflow (2026-07-23) on a proposed broadcast-video CV player-tracking system — biggest finding is that NBA already has real public rebound-tracking data, no CV needed there.
- project_daily_props_tooldaily_props.py — reusable one-command disciplined prop screen (hydrator + odds de-vig); built 2026-06-01
- project_debate_persona_lockdownMulti-persona debate (Koerner/Levitan/Raybon + judge) now reasons about FIX 1-5 lockdown flags + line_edges per pick. Closes the picker-facing gap end-to-end. Shipped 2026-05-21 via sport-agents e7be3cf + debate-v2 d7525b2.
- project_disciplined_slate_card_prototypeDisciplined "Top Plays — survived" prop-card prototype across the NBA+WNBA slate; destined for nightly-picks (NBA first). Built ad-hoc 2026-05-30.
- project_dvp_blowout_gate_refutedDO-NOT-BUILD: DvP gate + blowout gate both REFUTED (2026-08-27). DvP is ~85% noise (split-half r=-0.033); blowout gating DESTROYS 28u — high-spread games are the money formula's BEST bucket (+18.4% ROI on overs).
- project_grade_v2_systemgrade_v2 = calibrated P(hit) (user's definition: 'most likely to hit'), fitted on the settled ledger, monotone out-of-time; separate +EV value_flag; served-picks odds floor 1.70 decimal, NO upper cap. Live end-to-end incl. nightlypicks.com display, 2026-07-08.
- project_kelly_stakinggrading-service Kelly-derived stake replaced the hardcoded core/list formula-tier table — shipped end to end, default ON
- project_market_anchored_simulatorMulti-session build of nightlyhoops_build_plan.md §5/§6 (joint game simulator + allocation-compare) — Phase 1 shipped, phases 2-4 remaining
- project_matchup_page_perfectionVerified audit + prioritized plan to fix the nightlyhoops game details (matchup) page; plan doc lives in nightly-hoops/docs/matchup-page-perfection-plan.md
- project_methodology_lockdown4 methodology discipline checks hard-coded into data-hydrator PropDiagnostic (commit 8354c43, live in prod 2026-05-20)
- project_minutes_lever_measuredMinutes/redistribution/news-latency all MEASURED 2026-08-06: minutes projection is genuinely better, but does NOT transfer to prop accuracy. Three do-not-builds with numbers.
- project_mlb_pricing_fix_and_prereg_armsMLB priced=0 root cause (odds-service ODDS_ACTIVE_LEAGUES lacked mlb) FIXED 2026-08-28 by env; first PRE-REGISTERED hypothesis/control pair (mlb_always_under + mlb_coinflip) shipped in grading afdbd4a with kill rule in docs/mlb-prereg-2026-08-28.md.
- project_mlb_under_is_base_rateMLB's 93% UNDER is NOT the NBA missing-feature bug — it's a genuine 75% UNDER base rate on 0.5 lines. The 76% 'accuracy' is base rate, not skill.
- project_monstergpt_writeup_planPlan to produce MonsterGPT-style WNBA prop projections+narrative — verified data inventory across SDS/hydrator/scout/odds/agents (2026-06-16).
- project_multi_formula_clv_harnessMulti-formula CLV harness — runs N pick-generation formulas in parallel, tags each pick with formula_id, measures CLV per formula. 5-7 day build. Design doc at grading-service/docs/multi-formula-clv-harness.md.
- project_nba_season_readinessNBA-season readiness for sports-scout-service: the training leg was crash-looping since 2026-08-02; fixed on branch. What still must happen before opening night.
- project_news_speed_first_lightFirst real news-speed replay run (2026-07-28, 503 events): WNBA prop lines mostly don't exist when injury news breaks — the §10 play is beating late-posted OPENERS, not front-running line moves.
- project_odds_floor_validatedThe 1.70 decimal odds floor is MEASURED, not just reasoned: same formula, same ~62% hit rate both sides, but served picks (−143 or better) = +17.1% ROI vs floor-dropped = −1.7%. Price, not signal, is what's monetizable.
- project_parlay_pipelineParlay pipeline (2/3-leg slips, dual leg source, measured nightly) shipped to grading-service main 2026-06-11; included the odds game-scoping fix that ends the void avalanche
- project_programmatic_seo_livenightly-hoops programmatic-SEO prop-discovery engine is LIVE in prod (2026-06-20); free-first, $5/mo A+/A gate deferred with seam ready
- project_projections_engine_state[…] is BUILT + DEPLOYED end-to-end (NBA+WNBA crons writing nightly slates) but UNVALIDATED and NOT the source of truth. Blockers: (1) bet settlement n_settled=0, (2) ~0 model lift over L5, (3) not wired into live picks. Gap-audit 2026-05-29.
- project_public_board_kelly_gate_empties_cardRESOLVED 2026-08-28 (nightly-picks 4f85412): Board now posts the flat-1u CORE population (Ledger's) with Kelly as an annotation. Before: Board hid kelly<=0 rows since 8/09 → 0-3 picks/night all August while the Ledger showed +64.8u on the same feed.
- project_rank_bucket_endpointGET /clv/by-rank (grading b080939, 2026-08-04): self-computing top-1/3/5/10 bucket performance with z-tests. First read: clean gradient (+40%/+28%/+22% vs +6.8% tail) that is NOT significant (z≈1.0, n=170).
- project_scout_model_under_biasScout pipeline runs fine, but the NBA prop MODEL is weak + systematically UNDER-biased; precise root cause.
- project_sds_is_the_data_layerSDS is the single source of truth for stats/metadata/schedule/lineups/injuries/history/scout-data across all leagues. Only live/pbp/odds/projections/ml-scout are legitimately independent. Stop name-dropping legacy services (players/lineup/games/matchup/injury) as if they're authoritative.
- project_six_layer_gap_auditCode-verified audit (2026-07-23) of the stack against the "six-layer basketball intelligence" write-up + ranked upgrade roadmap. Corrects several stale memory notes.
- project_surprise_wnbaSurprise Desk is league-aware (NBA + WNBA) + hardened — shipped nightly-hoops eafdd44 (2026-06-20)
- project_two_month_retro_2026_08THE 2-month WNBA retro (13-agent audit, 2026-08-26): +71.4u verified but ROI-validated/CLV-neutral; ranked gap list (demotion gate #1, stale-line entry #2, playoff blindness #3); Kelly board refused 57u. Read before any 'how did we do' or NBA-prep work.
- project_wnba_active_gate_blog_generatorsport-agents blog_generator (WNBA/multi-sport) now drops Out/Doubtful/Questionable players at source — fd1ad23, 2026-06-22
- project_wnba_audit_2026_072026-07-07 full-chain WNBA prediction audit (9 agents, prod-probed): 5 confirmed live bugs + tiered roadmap. Verdict: plumbing/measurement problem, not modeling. Artifact: https://claude.ai/code/artifact/[…]
- project_wnba_backlogThe open WNBA work queue after the 2026-07-07/08 audit-and-fix arc — ordered, with evidence links. Check off / update as items land.
- project_wnba_batch_misses_afternoon_gamesThe 21:00Z WNBA prop batch runs AFTER tip for ~21% of games (weekend afternoons). Verified 2026-08-06. Fix is a cron slot, but it's coupled to a risky CDK deploy.
- project_wnba_debate_chain_blindWNBA debate/blog chain refused ALL picks because player agents got 0 odds_lines — extract_player_bundles read odds from the wrong shape AND joined on null player_id (+ a 5000-char prompt truncation cut prop_diagnostic). FIXED, MERGED TO MAIN + PUSHED (sport-agents 5029278) 2026-06-03; added B5 odds-aware selection. Data-layer verified. Remaining: deploy cron image + full LLM chain run.
- project_wnba_eb_shrinkageWNBA Gamma-Poisson EB shrinkage on the season-average blend input — shipped, default ON, not backtested (leakage), forward-shadow validation pending
- project_wnba_injury_joinWNBA injury join (3f31812): MERGED+PUSHED to main, but NOT DEPLOYED — stranded in the 7/07 batch image. Measured 2.6% catchable, not the stale 7.9%.
- project_wnba_lineups_gapRotoWire lineups chain SHIPPED end-to-end (verified 8/26): creds set, /wnba/lineups/date live with role+projected_minutes, PE consumes it. Still UNCONSUMED by the money formulas — the demotion gate (retro Gap 1) is the open use.
- project_wnba_lockdown_liveWNBA PropDiagnostic (lockdown FIX 1-5) end-to-end live in prod 2026-05-21. Verified against LV@PHO 2025 final — 18 players, 8 stats each, flags populated.
- project_wnba_model_auditWNBA prop-model audit (2026-06-15) — confirmed bug cluster FIXED in working tree (not committed/deployed); 25 WNBA tests + 622 suite green.
- project_wnba_projections_playbookWNBA projections-engine port — DONE 2026-05-31. The "4-6 day Phase 11" estimate below was WRONG; engine side was ~75-80% built (models trained+served), the only gap was one data-hydrator attach call.
- project_wnba_sds_emptyCLOSED 2026-05-21 by sport-data-service commit b2dba69 — WNBA snapshot reachability was a wnba_ prefix-strip bug in lookup queries, NOT an ingest gap. Rows existed; SDS just couldn't find them because data-hydrator/games-service emit wnba_X but SDS stores raw X.
- project_wnba_season_endgameWNBA regular season ends 2026-08-19 (~13 days from 8/06). What is and is NOT measurable in that budget — decides what's worth shipping now vs deferring to NBA.
- project_wnba_snapshot_depth_gapWNBA data-hydrator snapshot is missing 9 of 14 per-player depth blocks (vs_opponent, prop_metrics, advanced, projection, rotation_pattern, clutch_stats, series_history, memory, props) AND the entire snapshot.stats team-level block. Caught 2026-06-08 via new MCP get_player_depth tool returning legitimate Nones.
- project_wnba_snapshot_players_regressionROOT-CAUSED 2026-07-07 — never a regression. Bare /api/v1/snapshot/{id} serves a warmer-cached players:null WNBA snapshot; league route is healthy (37 players). Real victim: grading fanout.go:390 fetches the bare route → disciplined family = 0 WNBA picks all season.
- project_wnba_sport_agents_playbookConcrete work breakdown to build the WNBA sport-agents adapter. Three touchpoints, each with NBA regression risk; documented 2026-05-21 after attempting the build and hitting the complexity wall.
- project_wnba_ui_liveWNBA free public picks UI is live on nightlyhoops.com; two latent backend gaps remain (game_id prefix, grading WNBA history)
- project_wnba_week_counterfactualLast-week (6/30-7/6) replay of the fixed WNBA system, validated replicas: scout fixes = harm removal not edge (48% vs 50%); disciplined volume ~1 pick/wk at 18:00Z (Pinnacle quotes late); system_lineedges PRICED subset 64.1% n=206 +17% ROI = strongest signal; pick gate = integrity not accuracy; CLV was 32% contaminated with directional bias.
- project_worldcup_pipelineFIFA World Cup arc — ESPN data, game lines, Poisson props, LLM analyst, post-game recap + re-grade — shipped end-to-end in grading layer; frontend surfacing pending.
- project-analytics-engine/ask historical-question surface — Phases 1-2 live 2026-08-23; spec in nightly-picks docs/analytics-engine-spec.md.
- project-ask-availability-blockTonight-availability context block on /ask answer pages — shipped 2026-08-26 (nightly-picks e744230+3eeb523)
- project-board-warnings/verdict/card + board warning surface — shipped 2026-08-25, closes the "engine knows but never says" gap
- project-clv-settle-strandingCLV settle only retries today+yesterday, so any skipped row strands forever; DNP identity gap stranded 12.4k rows until fixed 2026-08-22.
- project-conditional-minutes-refutedDO NOT BUILD a conditional-on-minutes estimator for promoted players — measured null vs the model, 2026-08-25
- project-decision-layer-audit2026-07-24 full-stack audit — price is the only validated signal; live bugs found; do-not-build list
- project-slip-buildernightlypicks.com /slip SHIPPED (verified 2026-08-26: merged, /verdict/* routes live, /slip returns 200). Remaining: game-line legs, guest gating, ledger logging.
- project-window-provenanceLayoff + cross-season window flags in DH, suppressor + verdict rungs in grading — LIVE in prod (verified 2026-08-26 in WNBA snapshots)
- sportlib is the source of truth for prop types and league configFor any league/prop/odds-market work in sports-scout-service, check sportlib ([…]) before hand-coding tables.
- wnba-launch-stateWNBA projections launch — what's verified working vs remaining, as of 2026-05-16
Working rules
Mistakes the operator caught and the discipline that came out of them.
- do not let scope creep turn a small task into a deep refactorStop and check in when an "add X" task starts cascading into multi-file refactors. The user has rejected this twice in one session.
- feedback_active_tonight_gateFIX 5 revised 2026-06-09 (commit bf97c5f) — active_status now distinguishes "active"/"unknown"/{out,doubtful,questionable,probable,gtd}. Original injury-only join silently dropped healthy starters (Brunson/Hart Finals G3 miss).
- feedback_cache_ttl_bugCLOSED 2026-05-21 by commit dc380de — force-refresh now propagates into singleflight leader, ?refresh=true alias accepted, DELETE endpoint was always there (operator-discoverability gap, not capability).
- feedback_check_line_edges_before_quotingBEFORE quoting any prop pick (especially from Pinnacle de-vig), ALWAYS check the snapshot's prop_diagnostic.stats.{stat}.line_edges.quality + adjusted_projection. They override the de-vig math when they disagree.
- feedback_devig_phantom_edgeTwo ways a disciplined prop screen manufactures FAKE +EV: (1) single-book/stale-line de-vig, (2) trusting WNBA adjusted_projection. Caught my own screen doing both, 2026-06-22.
- feedback_feature_change_trapsFour traps that made my own scout feature/training changes net-harmful, caught only by adversarial review (2026-08-06). Check these before ANY feature-schema or train-stage edit, in any league.
- feedback_pipeline_exit_status_false_greenPushed SDS on a false green: `go test … | tail` returns tail's exit status, not go test's, so `&& ok=1` fired on failing tests. Also `git add -A <dir>` swept an untracked stray test into the commit.
- feedback_role_anchor_flag_gapThe role-anchor guard does NOT fire on bench->starter promotions: L10 pools two roles and manufactures a false mispriced_under. Caught on Rebecca Allen 2026-08-01.
- feedback_verify_before_extendStop stacking conclusions on unverified premises. 5 mistakes in one session (2026-05-21) all had the same root cause: skipped verification on "95% sure" answers. The data was always there.
- feedback-betting-disciplineHow to give the user sports-betting advice — what held vs what misled across a long live session
- model-edge-calibrationDon't quote model EV as real edge; +25-40% in a mature props market means the model is miscalibrated, not that books left money on the table. CLV over many bets is the only real evidence.
- pinnacle-defaultPinnacle is the default reference book for CLV / sharp-line analytics. Don't ship code that defaults to FanDuel/DraftKings as "primary" without first checking that Pinnacle is in the data — it usually is, and it's the right reference.
References
Where things live and how to read them.
- adding-a-league playbookAuthoritative phase-by-phase recipe for taking a new league end-to-end. Read first when adding wnba/ncaab/nfl/nhl/soccer.
- reference_nba_summer_league_statsstats.nba.com serves NBA Summer League by stint via distinct numeric LeagueIDs (13 California, 16 Salt Lake, 15 Vegas); pulled through sport-data-service's existing NBA proxy+headers. The gotcha is the required x-nba-stats-origin/token headers, NOT the LeagueID.
- reference_odds_book_qualityPer-book quality rules for odds-service props: BetRivers is one-way-only (never de-viggable), Pinnacle posts once and goes stale by tip. Verified 2026-07-31.
- reference_odds_service_wnba_routeodds-service ingests WNBA at /api/v1/odds/today?league=wnba but only 4 soft books (FD/DK/Caesars/BR) — NO PINNACLE. Sharp de-vig leg unavailable for WNBA until odds-service adds Pinnacle ingestion.
- reference-hypothesis-harnessscripts/test_hypothesis.py — permutation-test any "player X always does Y" hunch before believing it