Better than the consensus, a step shy of the league
Out of sample, Draftanomics's board predicts real NFL career value at 0.55 — ahead of the analyst consensus (0.45) and just behind NFL war rooms (0.58). Here's how it gets there, and where it still falls short.
Written by Draftanomics’s models from the live board and validated before publish — every number traces to a served value. How our writing works
There’s an honest way to grade a draft board and a flattering one. The flattering one picks a generous yardstick after the fact and reports a number that makes you look like a genius. We used to do that. This is the honest version.
The question that actually matters isn’t “did the model match the draft order” — front offices reach and slide for reasons that have nothing to do with talent. The question is: did the players the model ranked highly go on to have good NFL careers? So that’s what we measure, against the one number nobody can game after the fact.
What we’re measuring
For every prospect in a backtested class, we line up Draftanomics’s pre-draft ranking against the player’s eventual NFL career value — nflverse weighted Approximate Value, accumulated over their whole career, which is never a model input. Then we take the rank correlation: 0 is a coin flip, 1 is perfect foresight. We do it leave-one-year-out, so each class is scored by a model that never trained on it, and pool the result across 2017–2024.
Three boards, same test:
- Draftanomics’s board: 0.55. The best a public-information model gets.
- The analyst consensus: 0.45. The pooled public big-board field.
- NFL war rooms: 0.58. The bar — and they clear it every single year.
Read that honestly. Draftanomics orders prospects by eventual career value better than the entire public analyst field, and it does not beat the league. NFL teams lead in all eight backtested years. No public model we know of gets closer to them than this one — but “closer” is the claim, not “ahead.”
Why the model beats the consensus
Draftanomics trains on the same public data analysts use — consensus boards, combine drills, college production — but ranks prospects without the post-hoc rationalization that creeps into every public take. It doesn’t “fall in love” with a workout. It doesn’t down-weight a small-school prospect because nobody else has tape on him. It doesn’t promote a teammate of this year’s QB1 to inflate fit.
Three places the model has a structural edge over the field:
- Consensus signal weighting: every analyst board, weighted by years of hindsight accuracy. PFF, The Athletic, Daniel Jeremiah, Dane Brugler, Bleacher Report, Pro Football Network, and 22 others. The depth of that pooled signal — not any one scouting primitive — is where the edge over the field actually lives.
- Bust calibration per position group: the bust model is fit cohort-by-cohort, not pooled. A 4.4 forty means one thing for WR, another thing for OL. The model knows that. The field argues about it every year.
- Position-aware AAV: every prospect gets a comp-anchored second-deal AAV projection. Rookie scale is the floor; the ceiling comes from the regression on every veteran second deal in the last decade.
Why it doesn’t beat the NFL
NFL teams have private medicals, character interviews, and position-coach video sessions. Draftanomics doesn’t. If a top-15 prospect has a flagged medical that all 32 teams know about and analysts only whisper about, Draftanomics keeps him high while the room slides him. That information gap is most of the 0.58-vs-0.55 difference, and no amount of public-data modeling closes it.
Where the model gets it wrong
Three failure modes we publish:
- Trench positions (OT, IOL, IDL): the bust label for offensive and interior defensive linemen leaned on snap-share until this cycle. Pad-level tape doesn’t show up in box scores. Folding in PFF blocking grades moved the offensive-line bust model from anti-predictive (~0.30 AUC) to ~0.76 — the kind of swing that tells you how blind the old label was.
- Quarterbacks: small cohort, high variance, big stakes. The QB bust model is below random when forced to fit on its own, so it falls back to the overall model at ~0.66 AUC — better than a coin flip but nowhere near the position-specific accuracy of the WR or skill models.
- Pre-combine classes: the live class runs on stubbed combine data until real measurables land in late February. Until then the top tier compresses — every consensus #1 looks the same to the model — which is why the homepage’s “Top P(elite) right now” cards sort by canonical Draftanomics rank instead of raw p_elite.
The cleanest test
If you want to test the model yourself, open the track record and click any year. The page shows the full board against real NFL careers — every hit, every miss, every save, no selective citation, no hindsight editing.
That’s the bar: the best board public data can build, with the failures named and the league still out in front.