Model Health
The technical area for checking whether BanksPitch’s live scores behave sensibly over time.
Calibration by score band
Does a higher evidence score actually correspond with a higher real hit rate? Small calibration gaps are better.
Overall health
Technical measurements are kept here instead of cluttering the match pages.
Accuracy by market
A market should eventually be judged on its own settled sample, not just the overall record.
What do these terms mean?
Calibration gap
Actual hit rate minus average model score. A gap near zero means the score and observed result rate are closer together.
Brier score
A forecast-error measure. Lower is better. It is not very informative with a tiny sample.
Sample size
Small samples move sharply after each result. BanksPitch should not make major model changes from a handful of matches.
Historical 90-day backtest
Separate from the live record. This replays the model on past fixtures for research and calibration checks.
Performance by model version
Old predictions stay attached to the model version that created them, so improvements can be measured without rewriting history.