🤖 AI benchmark: hit-rate of 7 models

Prematch Live (in-play)

Seven external AI models (Hermes contour) independently analyze the same matches — predicting the outcome (1X2), total (Over/Under), both teams to score (BTTS) and the exact score. Here we honestly compare their predictions against the real result after the final whistle and combine everything into a single accuracy rating. An informational and analytical snapshot, not betting advice.

⚠️ Data is still accumulating — counting starts from 09.07.2026, so all models are compared on the same events (early test predictions are excluded). The sample is still small and not representative. Right now the snapshot holds 2679 match(es), 9041 settled AI predictions (Football, Tennis, Basketball, Volleyball, Hockey, Table tennis). The figures below are N, not «a percentage you can trust»: the more matches are played out, the more reliable the snapshot becomes. We show it transparently from day one, not only once the sample becomes «convenient».

Leaderboard

Model N (settled) 1X2 Double chance (1X) Exact score Composite accuracy
ChatGPT
263 65.0%(171/263) 74.6%(47/63) 17.2%(35/203) 44.2%(206/466)
Claude
2166 61.3%(1328/2166) 77.6%(612/789) 21.8%(412/1887) 42.9%(1740/4053)
Kimi
1109 60.8%(674/1109) 73.8%(253/343) 20.3%(196/964) 42.0%(870/2073)
Google AI
2806 60.3%(1691/2804) 74.7%(735/984) 20.0%(552/2764) 40.3%(2243/5568)
GLM 5.2
1285 59.5%(765/1285) 74.4%(302/406) 19.5%(246/1261) 39.7%(1011/2546)
DeepSeek
175 60.6%(106/175) 56.1%(23/41) 18.5%(32/173) 39.7%(138/348)
Qwen
1237 59.3%(733/1237) 75.0%(267/356) 16.8%(196/1166) 38.7%(929/2403)

grey — sample <5, not representative; «—» — the model has not made a settled prediction yet.

Double chance (1X) — the same pick counts as a win if the chosen side won or the match drew. Of the 1X2 losses in football/hockey: 680 draws, 743 underdog (total settled 1X2 picks in these sports: 2982, double chance combined 75.1% (2239/2982)). The models almost always take the favorite and don't bet on a draw — double chance shows how many bets are eaten specifically by draws.

Composite accuracy — the share of correct predictions across all shown markets together: (sum of correct picks) ÷ (sum of all settled picks) across the markets 1X2 + Exact score. Each market-pick weighs equally. This is hit-rate, not profitability — for money/ROI by model see /ai-agent. «Exact score» — the full final score was guessed correctly (H and A matched); for tennis, sets are compared; predictions with no recognized score do not count toward the denominator. Double chance (1X): a pick counts as a win if the chosen side won OR it was a draw — it accounts for frequent draws that «eat» bets on the favorite. This metric is informational and is not included in composite accuracy. On «All sports» we don't show total and BTTS — they're tied to the sport and make no sense in a mixed pool.

Composite model rating · all markets

Bar height = the model's composite accuracy across all applicable markets on the current sample. Sorted from best to worst.

44.2% (206/466)
GPT 5.5
42.9% (1740/4053)
Opus 4.8
42.0% (870/2073)
Kimi 2.6
40.3% (2243/5568)
Gemini 3.5 Flash
39.7% (1011/2546)
GLM 5.2
39.7% (138/348)
DeepSeek V4 Pro
38.7% (929/2403)
Qwen 3.7 Plus

Bars are AI models by version; grey/dimmed — sample <5, not representative. The snapshot is informational, not betting advice.

Accuracy by market

Where each model is strong: one mini-bar per applicable market, with the percentage and (hits/sample).

ChatGPT Composite 44.2%
1X2
65.0% (171/263)
Exact score
17.2% (35/203)
Claude Composite 42.9%
1X2
61.3% (1328/2166)
Exact score
21.8% (412/1887)
Kimi Composite 42.0%
1X2
60.8% (674/1109)
Exact score
20.3% (196/964)
Google AI Composite 40.3%
1X2
60.3% (1691/2804)
Exact score
20.0% (552/2764)
GLM 5.2 Composite 39.7%
1X2
59.5% (765/1285)
Exact score
19.5% (246/1261)
DeepSeek Composite 39.7%
1X2
60.6% (106/175)
Exact score
18.5% (32/173)
Qwen Composite 38.7%
1X2
59.3% (733/1237)
Exact score
16.8% (196/1166)

The model's favorite by 1X2 = the max of P1/X/P2 in its probabilities; for sports without a draw (tennis, volleyball, etc.) the «X» option doesn't participate. grey — sample <5, not representative. Not betting advice.