The honest test of a hand-tuned heuristic is whether it separates real outcomes. We run the engine’s collapse signal over 4,273 historical qualifying player-seasons (1996-97 – 2022-23) and check whether it predicted a real scoring decline in the next 3 seasons. The numbers below are reported as they fall — flattering or not.
| collapseProb ≥ | Precision | Recall | F1 |
|---|---|---|---|
| 20 | 20.1% | 56.8% | 29.7% |
| 25 | 22.8% | 48.2% | 30.9% |
| 30 | 25.0% | 41.4% | 31.2% |
| 35 | 31.2% | 19.4% | 23.9% |
| 40 | 32.9% | 8.7% | 13.7% |
| 45 | 31.4% | 4.7% | 8.1% |
Precision beats the 13.5% base rate at every threshold — so the signal is real, not noise (roughly a 1.9× lift at threshold 30). But it is a modest edge, and it does not clearly beat a trivial “flag everyone over 31” rule (precision 27.2%, recall 30.7%). That is the honest ceiling of hand-tuned weights fed a games-played injury proxy.
clamp(round((1 - gp/72) * 100), 0, 100)) because the stats snapshot has no injury data. Contract and salary inputs are held neutral. So this tests the age + availability mechanism, not the full heuristic.