NHL Prediction Accuracy | Model Performance Explained
Prospective Record — Saved Before Puck Drop
Regular-season raw-model forecasts recorded before scheduled puck drop. Each game uses the latest saved forecast before both its archived and current start time. Results include overtime and shootouts. Historical date-only records are excluded.
No completed games have a qualifying timestamped forecast yet. This record starts with new producer captures; historical predictions are not backfilled.
This page explains how to interpret our model's prediction accuracy metrics and provides historical evaluation context. Visit Model Performance for detailed results.
Historical Fit — Includes Training Data
All four cards use the same historical evaluation summary. These results include training data and are not an estimate of performance on unseen games. The export does not identify a model version or evaluation date range. Monthly report coverage: 2023-10-01 to 2026-04-01. Brier = probabilistic score (lower is better, 0.25 = 50% assigned to every game). MAE Total = mean absolute error on predicted total goals (the model's point estimate is the median, which MAE — not RMSE — is the consistent error metric for).
What the Metrics Mean
Accuracy
The fraction of games where the model correctly predicted the winner (the team with win probability >50%). A 50/50 baseline is useful for comparison, but an evaluation must identify its model version, date range, and whether forecasts preceded the games.
Brier Score
The Brier score measures probabilistic accuracy: it is the mean squared difference between the predicted probability and the binary outcome (1 = home win, 0 = home loss). A random model predicting 50% every game scores 0.25. Lower Brier scores are better. Compare models on the same held-out games and report the sample size.
Calibration
Calibration measures whether predicted probabilities match observed frequencies. If the model says 65% in 100 games, those teams should win about 65 of them. Production win probabilities use the raw model blend, without post-processing calibration. Calibration quality must be measured on unseen games. See the Performance page for calibration curves.
MAE (Totals)
Mean absolute error on predicted total goals (over/under). The model's point estimate is the median of its simulated total, and the median minimizes MAE (the mean minimizes RMSE), so MAE is the error metric consistent with how the point estimate is chosen. A perfect model would score 0. The cross-validation table below reports RMSE instead, because those folds score the mean (expected goals), where RMSE is the consistent metric.
Cross-Validation Results (3 folds)
Walk-forward cross-validation. Feature and upstream-data timing require separate audits. Avg Brier: 0.2549 | Avg Log-loss: 0.7033 | Avg RMSE (Totals): 2.393
| Fold | Brier | Log-loss | RMSE Total | Train N | Val N |
|---|---|---|---|---|---|
| 1 | 0.2561 | 0.7059 | 2.431 | 703 | 2,089 |
| 2 | 0.2550 | 0.7036 | 2.369 | 1,396 | 1,396 |
| 3 | 0.2534 | 0.7003 | 2.380 | 2,094 | 698 |
Monthly Accuracy Trend
Win/loss prediction accuracy by calendar month. Larger samples = more stable estimates.
Monthly Breakdown
| Month | Games | Accuracy | Brier Score |
|---|---|---|---|
| 2023-10 | 140 | 57.1% | 0.2383 |
| 2023-11 | 213 | 56.8% | 0.2422 |
| 2023-12 | 219 | 61.6% | 0.2394 |
| 2024-01 | 208 | 54.3% | 0.2358 |
| 2024-02 | 172 | 61.0% | 0.2369 |
| 2024-03 | 228 | 62.3% | 0.2267 |
| 2024-04 | 132 | 53.0% | 0.2495 |
| 2024-10 | 166 | 67.5% | 0.2128 |
| 2024-11 | 220 | 63.6% | 0.2213 |
| 2024-12 | 214 | 65.9% | 0.2097 |
| 2025-01 | 224 | 63.4% | 0.2346 |
| 2025-02 | 122 | 56.6% | 0.2259 |
| 2025-03 | 234 | 65.4% | 0.2169 |
| 2025-04 | 132 | 65.9% | 0.2308 |
| 2025-10 | 180 | 66.7% | 0.2098 |
| 2025-11 | 225 | 63.1% | 0.2210 |
| 2025-12 | 226 | 65.5% | 0.2255 |
| 2026-01 | 240 | 60.8% | 0.2284 |
| 2026-02 | 74 | 67.6% | 0.2255 |
| 2026-03 | 242 | 52.9% | 0.2505 |
| 2026-04 | 125 | 61.6% | 0.2236 |
Evaluating Unseen Games
A forward test trains only on earlier games and freezes each forecast before its outcome is known. Historical fits and cross-validation folds are separate evaluations and should retain their own model versions, date ranges, and sample sizes.
No current-model forward accuracy estimate is supplied by this page's summary export. See Model Performance for the available dated reports.
Why Predictions Remain Uncertain
A probability describes uncertainty in an outcome. Even a strong favorite can lose. Sources of unpredictability include:
- Goalie performance variance (hot/cold streaks)
- Injuries and lineup uncertainty announced close to game time
- Puck luck (post hits, lucky bounces)
- Small sample sizes — 82-game seasons with nightly scheduling
We evaluate whether probabilities reflect uncertainty, even when no single prediction is guaranteed.
See also: Live Model Performance | Full Methodology | Today's Predictions