🤖 AI benchmark: hit-rate of 7 models

Prematch Live (in-play)

Seven external AI models (Hermes contour) independently analyze the same matches — predicting the outcome (1X2), total (Over/Under), both teams to score (BTTS) and the exact score. Here we honestly compare their predictions against the real result after the final whistle and combine everything into a single accuracy rating. An informational and analytical snapshot, not betting advice.

⚠️ Data is still accumulating — counting starts from 09.07.2026, so all models are compared on the same events (early test predictions are excluded). The sample is still small and not representative. Right now the snapshot holds 1032 match(es), 3693 settled AI predictions (Tennis). The figures below are N, not «a percentage you can trust»: the more matches are played out, the more reliable the snapshot becomes. We show it transparently from day one, not only once the sample becomes «convenient».

Leaderboard · Tennis

Model N (settled) 1X2 Exact score (by sets) Composite accuracy
ChatGPT
106 71.7%(76/106) 33.3%(28/84) 54.7%(104/190)
Kimi
508 66.9%(340/508) 33.9%(151/445) 51.5%(491/953)
DeepSeek
84 69.0%(58/84) 32.9%(27/82) 51.2%(85/166)
Claude
760 65.8%(500/760) 35.2%(248/704) 51.1%(748/1464)
Google AI
1060 64.8%(687/1060) 35.7%(374/1049) 50.3%(1061/2109)
GLM 5.2
573 63.5%(364/573) 35.1%(196/559) 49.5%(560/1132)
Qwen
602 62.5%(376/602) 28.2%(163/579) 45.6%(539/1181)

grey — sample <5, not representative; «—» — the model has not made a settled prediction yet.

Composite accuracy — the share of correct predictions across all shown markets together: (sum of correct picks) ÷ (sum of all settled picks) across the markets 1X2 + Exact score (by sets). Each market-pick weighs equally. This is hit-rate, not profitability — for money/ROI by model see /ai-agent. «Exact score (by sets)» — the full final score was guessed correctly (H and A matched) by sets; predictions with no recognized score do not count toward the denominator.

Composite model rating · all markets · Tennis

Bar height = the model's composite accuracy across all applicable markets on the current sample. Sorted from best to worst.

54.7% (104/190)
GPT 5.5
51.5% (491/953)
Kimi 2.6
51.2% (85/166)
DeepSeek V4 Pro
51.1% (748/1464)
Opus 4.8
50.3% (1061/2109)
Gemini 3.5 Flash
49.5% (560/1132)
GLM 5.2
45.6% (539/1181)
Qwen 3.7 Plus

Bars are AI models by version; grey/dimmed — sample <5, not representative. The snapshot is informational, not betting advice.

Accuracy by market · Tennis

Where each model is strong: one mini-bar per applicable market, with the percentage and (hits/sample).

ChatGPT Composite 54.7%
1X2
71.7% (76/106)
Exact score (by sets)
33.3% (28/84)
Kimi Composite 51.5%
1X2
66.9% (340/508)
Exact score (by sets)
33.9% (151/445)
DeepSeek Composite 51.2%
1X2
69.0% (58/84)
Exact score (by sets)
32.9% (27/82)
Claude Composite 51.1%
1X2
65.8% (500/760)
Exact score (by sets)
35.2% (248/704)
Google AI Composite 50.3%
1X2
64.8% (687/1060)
Exact score (by sets)
35.7% (374/1049)
GLM 5.2 Composite 49.5%
1X2
63.5% (364/573)
Exact score (by sets)
35.1% (196/559)
Qwen Composite 45.6%
1X2
62.5% (376/602)
Exact score (by sets)
28.2% (163/579)

The model's favorite by 1X2 = the max of P1/X/P2 in its probabilities; for sports without a draw (tennis, volleyball, etc.) the «X» option doesn't participate. grey — sample <5, not representative. Not betting advice.