⌂ Home

Model Health

The technical area for checking whether BanksPitch’s live scores behave sensibly over time.

This page does not change the model. It measures forecasts that were already stored before kickoff. No historical result is used to rewrite the original prediction.
—
Last 7 days
—
Last 30 days
—
Live sample

Calibration by score band

Does a higher evidence score actually correspond with a higher real hit rate? Small calibration gaps are better.

Loading calibration

Overall health

Technical measurements are kept here instead of cluttering the match pages.

Loading metrics

Accuracy by market

A market should eventually be judged on its own settled sample, not just the overall record.

Loading markets

What do these terms mean?

Calibration gap

Actual hit rate minus average model score. A gap near zero means the score and observed result rate are closer together.

Brier score

A forecast-error measure. Lower is better. It is not very informative with a tiny sample.

Sample size

Small samples move sharply after each result. BanksPitch should not make major model changes from a handful of matches.

Historical 90-day backtest

Separate from the live record. This replays the model on past fixtures for research and calibration checks.

Keep the distinction clear: backtesting is historical research. The live prediction record above is the stronger evidence for how forecasts actually behave when stored before kickoff.
Backtest has not been run in this session.

Performance by model version

Old predictions stay attached to the model version that created them, so improvements can be measured without rewriting history.

Loading model versions
Loading data-integrity status…