Calibration: Do a Model's Probabilities Mean What They Say?
How calibration is checked
Group every pick by the probability it was given, then count how often each group won. A well-calibrated model's 70% group wins about 70% of the time. Being right often is not the same as being calibrated: a model that gives every pick 90% and wins 60% is confidently wrong.
Each bucket's actual rate comes with a range, because a bucket of 100 picks is a small sample (see hit rate). A bucket is off only when what it said falls outside that range.
Scoring one probability: log loss
Log loss scores a probability by the chance it gave to what actually happened, as -ln(p). A confident miss costs much more than a confident hit saves.
- It wins
- -ln(0.70) = 0.357
- It loses
- -ln(0.30) = 1.204
- A 50/50 guess, either way
- -ln(0.50) = 0.693
Lower is better. A model beats a coin flip only if its average log loss stays under 0.693. Averaged over many games it rewards both being right and being honest about uncertainty.
Where Probetrics shows it
Every sport with a model shows its calibration bucket by bucket on its Results and model pages, and the live buckets fill in as picks are graded. A 70% pick still loses about three times in ten; that is the model working as designed. Calibration is about winner probabilities, which are the record. Probetrics does not publish an over or under record.
Questions and answers
- What does it mean for a model to be calibrated?
- Its stated probabilities match outcomes: of everything it called about 70%, about 70% happened. It is a different test from how often the model is right.
- Why does a calibration table show a range?
- Each bucket holds a limited number of picks, so its actual rate is uncertain. The range shows how far the true rate could plausibly be from the observed one.
- What is log loss?
- A score for a probability forecast equal to minus the natural log of the chance given to what happened. Lower is better, and a coin-flip forecast scores 0.693.
Read next
- Walk-forward backtestingWalk-forward backtesting scores a model on games it has not seen by pricing each game only from what was known before it, moving forward through time.
- Market-anchored probabilityA market-anchored probability starts from the sportsbook's no-vig price for a game and lets a model change it only where testing has shown the model adds information.
- Hit rateA hit rate is the share of a player's games that finished on one side of a prop line, and over a small sample it is a wide range rather than a single number, which a Wilson interval makes visible.
- No-vig oddsNo-vig odds, also called fair odds, are a two-sided price with the sportsbook's margin (the vig, or juice) taken out, so the two sides' chances add up to 100%.
Results and calibration →How the models are tested →All terms →
This page explains arithmetic and vocabulary. It does not recommend a bet, a side, an amount or a sportsbook, and the examples are illustrations. For information only, not betting advice. 21+ only. Gambling problem? Call or text 1-800-GAMBLER.