Every ATP & WTA match prediction The Drop Shot makes, scored openly against the betting closing line. Log loss, Brier score, accuracy, and AUC — sliced by tour, surface, and round, and updated as matches finish.
Each prediction is a pre-match win probability for both players, generated from tennis signal alone — surface, ranking, head-to-head history, recent form, and serve and return performance. After the match finishes we compare that probability to the market's closing line, the consensus price just before play begins. The closing line is a deliberately hard benchmark: it already folds in everything the public knows, so beating it consistently is difficult.
The dashboard reports the same four metrics professional forecasters use, on the overlapping set of matches where both a prediction and a closing line exist:
Most tennis prediction products show you a number and never tell you whether it was right. We do the opposite: every scored match is in the sample, the metrics move as results come in, and the slices where the closing line still wins are shown next to the ones where the model does. Predictions are only worth trusting if the track record is in the open — so it is.
Small slices are noisy: a surface-and-round cell with a handful of matches can swing on a single upset, so weight the all-tour, all-surface numbers most and treat narrow cells as directional. The closing line is not a fixed target either — it's a strong, adaptive benchmark, and a period where the model trails it is information, not a bug. Use the since-date filter to judge recent form rather than the entire history.
Every ATP and WTA match we predict is scored after it finishes against the betting closing line — the market's final consensus probability just before play begins, the toughest public benchmark there is. We report four standard metrics on the overlapping set of matches: log loss and Brier score (lower is better, they reward well-calibrated probabilities), accuracy (share of matches where the favorite won), and AUC (how well the probabilities rank winners above losers). The dashboard slices every metric by tour, surface, and round, with a since-date filter.
Log loss penalizes confident wrong calls heavily — a model that says 90% and is wrong is punished far more than one that says 55%. A blind coin-flip scores 0.69; lower is better. Brier score is the mean squared error of the probabilities (0 is perfect, 0.25 is a random guess). AUC is the probability that a randomly chosen winner was rated higher than a randomly chosen loser — 0.50 is coin-flip, 0.75–0.80 is bookmaker-grade. Together they describe both how calibrated and how discriminating the model is.
No. The closing line is the benchmark we score against, never an input to the model. The predictions are generated from tennis signal alone — surface, ranking, head-to-head history, recent form, and serve/return performance — and only afterward compared to where the market closed. That separation is what makes the comparison honest: the model never sees the answer it's being graded on.
On some slices yes, on others no — and we publish both. The point of the dashboard is accountability, not a marketing claim: every scored match is in the sample, the metrics update as results come in, and the slices where the closing line still wins are shown alongside the ones where the model does. The current head-to-head, by tour and surface, is on the dashboard above.
It re-scores continuously as matches finish — new results enter the sample within hours, and the headline metrics recompute against the latest closing lines. Use the since-date filter to window the dashboard to a recent period (for example, since the last model retrain) rather than the all-time sample.