Skip to main content
ProbetricsProbetrics

Walk-Forward Backtesting: Testing a Model Honestly

Walk-forward backtesting scores a model on games it has not seen by pricing each game only from what was known before it, moving forward through time.
Updated · Education, not advice

The problem it solves

A model fitted to a season and then scored on that same season always looks good, because it has seen the answers. The same happens if a stat quietly includes information from after the game, such as the starting quarterback of that day when only last week's starter was knowable. That is look-ahead bias, and it is why most backtests look better than the live results.

How a walk-forward test runs

  • Go through the games in date order. Each game is priced using only games played before it.
  • Tune settings on one stretch of seasons, the tuning window.
  • Freeze the settings. Then score a later stretch, the test window, that was not used for tuning. Look at it once.
  • Keep a change only if it beats the baseline on the tuning window and still does on the test window. Report a null result as a null result.
An illustration of the windows (example years)
StepSeasonsWhat happens
Tune2015 to 2019Try settings and pick the best, scoring each game on prior games only
Freeze-The settings are locked; no further changes
Test2020 to 2022Score once on seasons the tuning never touched
LiveThis seasonPicks are locked before the start and graded in public

Example years only. The actual windows differ by sport.

What it does not prove

A good walk-forward result shows a method held up on games it had not seen. It does not guarantee the next season will look the same. That is why live picks are locked before the start and graded in public: the live record, with its calibration, is the final check. Probetrics does not publish the exact inputs, weights or settings, or what was tried and rejected.

Questions and answers

What is walk-forward testing?
A backtest that moves through time, pricing each game only from earlier games, tuning on one window and scoring a later, untouched window.
What is look-ahead bias?
Using information in a test that would not have been known at the time, such as a final box score or a starter announced after the pick. It makes results look better than they could be live.
Why freeze the settings before testing?
Because adjusting them after seeing the test result quietly turns the test window into a tuning window, and the result no longer measures anything the model has not seen.

How the models are tested →Results →All terms →

This page explains arithmetic and vocabulary. It does not recommend a bet, a side, an amount or a sportsbook, and the examples are illustrations. For information only, not betting advice. 21+ only. Gambling problem? Call or text 1-800-GAMBLER.