What does out-of-sample mean, and why does it matter?
The model is graded on draft classes it never trained on — the only test that says anything about future performance.
A model fitted on data will fit that data. Show it the same players it learned from and it will look extraordinary, and that number will tell you nothing about the next class.
Out-of-sample testing removes the trick. Each draft class is held out of training entirely, the model is fitted on the others, and then it is asked to rank the class it has never seen. Repeat for every year. What you get is an estimate of how the board performs on players nobody had outcomes for yet — which is the only situation that ever actually occurs.
It is a less flattering number, always. That is the point of publishing it. Any accuracy figure that does not say which data the model was trained on is unfalsifiable, and unfalsifiable numbers are marketing.
The same discipline applies to the boards themselves. Ours are frozen and published before the draft, then graded afterward against what happened. A board that can be quietly edited after the fact cannot be graded at all.