Scored on held-out ATP & WTA tour main-draw matches it never trained on: the sample size, the Brier against 0.250 for a coin flip and the straight-up accuracy are all read live from the calibration feed, so this page never quotes a measurement it has not just loaded. And the probabilities are honest β when we say 60%, it happens about 60% of the time.
We sort every held-out match above into 10 buckets by the win probability we gave the favorite, then check what fraction actually won. A perfectly honest model sits on the diagonal: the dots should hug the dashed line.
| Predicted | Won (actual) | Miss | Matches |
|---|
Brier score is the average squared error between the probability we gave and what happened (lower is better; 0.25 = a coin flip, 0 = perfect). Straight-up accuracy is how often the side we favored actually won. Calibration is the most important: across the whole range, our stated probabilities match real-world frequencies.
These are measured on the same surface-aware power-rating + serve/return blend we use to price every match on the site β not a curated subset.