Skip to content

Contact us

Send us a message

Questions, bugs or ideas: it goes straight to the Scoutics team, and we reply by email.

What is it about?

Prefer a full page? Open Support

EPL Premier League
Your matchups

Pick a fixture, run the simulator, and explore the pitch in detail.

Model report Internal historical evaluation

Model performance.

See how our match probabilities performed on later games kept out of training, calibration and model selection.

Chronological holdouts Deployed model artifacts
Forward record   Read the methodology
Four sports. One transparent standard.

The current scorecard

Select a sport to explore its results

AUC measures ranking, not the percentage of correct picks: 0.50 is chance ranking; 1.00 is perfect ranking. EPL draws and NFL ties are excluded from AUC only.

Selected probability model

NBA holdout results

Built 04 Oct 2026

Logistic regression · Selected on validation log loss before the holdout was scored.

Winner ROC-AUC0.716Higher is better · non-tied games
Outcome accuracy67.7%Correct most-probable outcome
Log loss0.617Lower is better · all outcomes
Brier score0.426Lower is better · summed over 3 outcomes
Test window to
Holdout population1,458 games · 1,458 for AUC
Outcome definitionFull-game winner · including overtime
Compare all models and data splits 4 evaluated

Every evaluated model. The same later games.

All candidates use the same data splits and outcome definition. The deployed model is selected by the lowest log loss on the earlier selection block; holdout scores do not decide the winner.

NBA · 1,458 holdout games · 1,458 for AUC
ModelSelection blockHoldout test
Log loss ↓Winner AUC ↑Accuracy ↑Log loss ↓Brier ↓
Logistic regressionDeployed 0.6194 0.7161 67.7% 0.6172 0.4263
Histogram gradient boosting 0.6286 0.6991 65.2% 0.6231 0.4337
Extra trees 0.6231 0.7155 67.0% 0.6150 0.4258
Training-frequency baselineBenchmark 0.6894 0.5000 55.4% 0.6874 0.4942

The training-frequency baseline assigns identical probabilities to every game. It is a statistical benchmark, not bookmaker odds. AUC evaluates non-tied games; accuracy, log loss and summed three-outcome Brier use the full holdout.

Earlier data builds the model. Later data tests it.

  1. 1 · Training6,613 games
  2. 2 · Calibration1,080 games
  3. 3 · Selection1,690 games
  4. 4 · Holdout test1,458 games

30 pregame features · Version independent-score-form-v3 · Metrics loaded from the deployed artifact

Selected probability model

NFL holdout results

Built 04 Oct 2026

Logistic regression · Selected on validation log loss before the holdout was scored.

Winner ROC-AUC0.691Higher is better · non-tied games
Outcome accuracy63.7%Correct most-probable outcome
Log loss0.656Lower is better · all outcomes
Brier score0.449Lower is better · summed over 3 outcomes
Test window to
Holdout population947 games · 945 for AUC
Outcome definitionFull game · home / tie / away

2 ties are excluded from winner AUC. Accuracy, log loss and Brier include the full 947-game holdout.

Compare all models and data splits 4 evaluated

Every evaluated model. The same later games.

All candidates use the same data splits and outcome definition. The deployed model is selected by the lowest log loss on the earlier selection block; holdout scores do not decide the winner.

NFL · 947 holdout games · 945 for AUC
ModelSelection blockHoldout test
Log loss ↓Winner AUC ↑Accuracy ↑Log loss ↓Brier ↓
Logistic regressionDeployed 0.6764 0.6910 63.7% 0.6562 0.4491
Histogram gradient boosting 0.6788 0.6707 61.4% 0.6579 0.4539
Random forest 0.7044 0.6904 63.0% 0.6805 0.4480
Training-frequency baselineBenchmark 0.7212 0.5000 54.7% 0.7035 0.4987

The training-frequency baseline assigns identical probabilities to every game. It is a statistical benchmark, not bookmaker odds. AUC evaluates non-tied games; accuracy, log loss and summed three-outcome Brier use the full holdout.

Earlier data builds the model. Later data tests it.

  1. 1 · Training4,676 games
  2. 2 · Calibration673 games
  3. 3 · Selection1,013 games
  4. 4 · Holdout test947 games

30 pregame features · Version independent-score-form-v3 · Metrics loaded from the deployed artifact

Selected probability model

NHL holdout results

Built 04 Oct 2026

Logistic regression · Selected on validation log loss before the holdout was scored.

Winner ROC-AUC0.547Higher is better · non-tied games
Outcome accuracy53.3%Correct most-probable outcome
Log loss0.690Lower is better · all outcomes
Brier score0.497Lower is better · summed over 3 outcomes
Test window to
Holdout population1,596 games · 1,596 for AUC
Outcome definitionFull-game winner · including overtime / shootout

NHL's current winner ranking is close to chance. Treat this as limited historical separation, not evidence of a dependable betting edge.

Compare all models and data splits 4 evaluated

Every evaluated model. The same later games.

All candidates use the same data splits and outcome definition. The deployed model is selected by the lowest log loss on the earlier selection block; holdout scores do not decide the winner.

NHL · 1,596 holdout games · 1,596 for AUC
ModelSelection blockHoldout test
Log loss ↓Winner AUC ↑Accuracy ↑Log loss ↓Brier ↓
Logistic regressionDeployed 0.6696 0.5467 53.3% 0.6901 0.4968
Histogram gradient boosting 0.6742 0.5388 53.1% 0.6943 0.5009
Extra trees 0.6720 0.5467 53.4% 0.6910 0.4977
Training-frequency baselineBenchmark 0.6873 0.5000 52.5% 0.6920 0.4989

The training-frequency baseline assigns identical probabilities to every game. It is a statistical benchmark, not bookmaker odds. AUC evaluates non-tied games; accuracy, log loss and summed three-outcome Brier use the full holdout.

Earlier data builds the model. Later data tests it.

  1. 1 · Training6,734 games
  2. 2 · Calibration1,188 games
  3. 3 · Selection1,747 games
  4. 4 · Holdout test1,596 games

30 pregame features · Version independent-score-form-v3 · Metrics loaded from the deployed artifact

Selected probability model

EPL holdout results

Built 04 Oct 2026

Logistic regression · Selected on validation log loss before the holdout was scored.

Winner ROC-AUC0.765Higher is better · non-tied games
Outcome accuracy52.9%Correct most-probable outcome
Log loss0.981Lower is better · all outcomes
Brier score0.585Lower is better · summed over 3 outcomes
Test window to
Holdout population1,342 games · 1,017 for AUC
Outcome definitionRegulation time · home / draw / away

325 draws are excluded from winner AUC. Accuracy, log loss and Brier include the full 1,342-game holdout.

Compare all models and data splits 4 evaluated

Every evaluated model. The same later games.

All candidates use the same data splits and outcome definition. The deployed model is selected by the lowest log loss on the earlier selection block; holdout scores do not decide the winner.

EPL · 1,342 holdout games · 1,017 for AUC
ModelSelection blockHoldout test
Log loss ↓Winner AUC ↑Accuracy ↑Log loss ↓Brier ↓
Logistic regressionDeployed 0.9995 0.7647 52.9% 0.9808 0.5849
Histogram gradient boosting 1.0189 0.7397 49.9% 1.0133 0.6058
Extra trees 1.0067 0.7515 52.6% 0.9998 0.5968
Training-frequency baselineBenchmark 1.0780 0.5000 44.3% 1.0708 0.6477

The training-frequency baseline assigns identical probabilities to every game. It is a statistical benchmark, not bookmaker odds. AUC evaluates non-tied games; accuracy, log loss and summed three-outcome Brier use the full holdout.

Earlier data builds the model. Later data tests it.

  1. 1 · Training6,247 games
  2. 2 · Calibration952 games
  3. 3 · Selection1,259 games
  4. 4 · Holdout test1,342 games

30 pregame features · Version independent-score-form-v3 · Metrics loaded from the deployed artifact

Read the numbers with context

Ranking is one part of the story.

01

Does the model rank well?

Winner AUC checks whether higher home-win scores tend to correspond to home wins, conditional on a non-tied result. It does not measure calibration or betting returns.

02

How good are the probabilities?

Log loss and Brier score evaluate the probability assigned to every outcome. Confident wrong forecasts cost more. Our Brier is a sum across three outcomes, not binary Brier.

03

Is there a market advantage?

That requires comparison with odds available at the same time, after bookmaker margin and costs. These historical tests do not establish profitable betting returns.

Evaluation details

How we test. What this covers.

Chronology comes first

Features use earlier completed game scores, including recent form, scoring history, rest and Elo-style strength. Results from the same date are withheld as a group, so they cannot enter another game's inputs that day.

Models fit the training block. Temperature scaling uses the calibration block; model family selection uses validation log loss. The later holdout is scored after selection, without refitting on it.

Scope and next evidence

This is an internal evaluation of match-outcome probability models. It does not validate lineup-impact estimates, the scenario or replay engine, FPL points, or specialty markets such as totals and corners.

Compare model revisions only when their features, test games and metric definitions match. A score from an earlier research experiment is not directly comparable with a different deployed evaluation.

US source histories include preseason and postseason games. Sports use different periods and outcome definitions, so this is not a controlled ranking of one sport against another.

Prospective, timestamped forecasts, same-time market benchmarks and calibration analysis are the next evidence needed to assess live decision quality.

From evidence to exploration

Compare the next matchup.

Explore model probabilities alongside available market prices in Betting Lab.

Explore upcoming games