Live Tennis Prediction Accuracy

Every ATP & WTA match prediction The Drop Shot makes, scored openly against the betting closing line. Log loss, Brier score, accuracy, and AUC — sliced by tour, surface, and round, and updated as matches finish.

How we score predictions

Each prediction is a pre-match win probability for both players, generated from tennis signal alone — surface, ranking, head-to-head history, recent form, and serve and return performance. After the match finishes we compare that probability to the market's closing line, the consensus price just before play begins. The closing line is a deliberately hard benchmark: it already folds in everything the public knows, so beating it consistently is difficult.

Headline metrics — Drop Shot vs the closing line

The dashboard reports the same four metrics professional forecasters use, on the overlapping set of matches where both a prediction and a closing line exist:

  • Log loss — punishes confident wrong calls hardest; lower is better (a coin flip scores 0.69).
  • Brier score — mean squared error of the probabilities; 0 is perfect, 0.25 is a random guess.
  • Accuracy — the share of matches where the rated favorite won.
  • AUC — how cleanly the probabilities rank winners above losers; 0.50 is chance, 0.75–0.80 is bookmaker-grade.

Why publish prediction accuracy?

Most tennis prediction products show you a number and never tell you whether it was right. We do the opposite: every scored match is in the sample, the metrics move as results come in, and the slices where the closing line still wins are shown next to the ones where the model does. Predictions are only worth trusting if the track record is in the open — so it is.

Limitations & caveats

Small slices are noisy: a surface-and-round cell with a handful of matches can swing on a single upset, so weight the all-tour, all-surface numbers most and treat narrow cells as directional. The closing line is not a fixed target either — it's a strong, adaptive benchmark, and a period where the model trails it is information, not a bug. Use the since-date filter to judge recent form rather than the entire history.

Frequently asked questions

How does The Drop Shot measure prediction accuracy?

Every ATP and WTA match we predict is scored after it finishes against the betting closing line — the market's final consensus probability just before play begins, the toughest public benchmark there is. We report four standard metrics on the overlapping set of matches: log loss and Brier score (lower is better, they reward well-calibrated probabilities), accuracy (share of matches where the favorite won), and AUC (how well the probabilities rank winners above losers). The dashboard slices every metric by tour, surface, and round, with a since-date filter.

What do log loss, Brier score, and AUC mean?

Log loss penalizes confident wrong calls heavily — a model that says 90% and is wrong is punished far more than one that says 55%. A blind coin-flip scores 0.69; lower is better. Brier score is the mean squared error of the probabilities (0 is perfect, 0.25 is a random guess). AUC is the probability that a randomly chosen winner was rated higher than a randomly chosen loser — 0.50 is coin-flip, 0.75–0.80 is bookmaker-grade. Together they describe both how calibrated and how discriminating the model is.

Are betting closing lines used to train the model?

No. The closing line is the benchmark we score against, never an input to the model. The predictions are generated from tennis signal alone — surface, ranking, head-to-head history, recent form, and serve/return performance — and only afterward compared to where the market closed. That separation is what makes the comparison honest: the model never sees the answer it's being graded on.

Does the model beat the betting lines?

On some slices yes, on others no — and we publish both. The point of the dashboard is accountability, not a marketing claim: every scored match is in the sample, the metrics update as results come in, and the slices where the closing line still wins are shown alongside the ones where the model does. The current head-to-head, by tour and surface, is on the dashboard above.

How often is the accuracy dashboard updated?

It re-scores continuously as matches finish — new results enter the sample within hours, and the headline metrics recompute against the latest closing lines. Use the since-date filter to window the dashboard to a recent period (for example, since the last model retrain) rather than the all-time sample.