🤖 AI benchmark: hit-rate of 7 models

Prematch Live (in-play)

Seven external AI models (Hermes contour) independently analyze the same matches — predicting the outcome (1X2), total (Over/Under), both teams to score (BTTS) and the exact score. Here we honestly compare their predictions against the real result after the final whistle and combine everything into a single accuracy rating. An informational and analytical snapshot, not betting advice.

⚠️ Data is still accumulating — counting starts from 09.07.2026, so all models are compared on the same events (early test predictions are excluded). The sample is still small and not representative. Right now the snapshot holds 15 match(es), 28 settled AI predictions (Hockey). The figures below are N, not «a percentage you can trust»: the more matches are played out, the more reliable the snapshot becomes. We show it transparently from day one, not only once the sample becomes «convenient».

Leaderboard · Hockey

Model N (settled) 1X2 Double chance (1X) Total goals BTTS Exact score Composite accuracy
Claude
12 50.0%(6/12) 50.0%(6/12) 58.3%(7/12) 75.0%(9/12) 0.0%(0/12) 45.8%(22/48)
Google AI
16 56.3%(9/16) 56.3%(9/16) 50.0%(7/14) 78.6%(11/14) 0.0%(0/16) 45.0%(27/60)
DeepSeek
0
ChatGPT
0
Qwen
0
Kimi
0
GLM 5.2
0

grey — sample <5, not representative; «—» — the model has not made a settled prediction yet.

Double chance (1X) — the same pick counts as a win if the chosen side won or the match drew. Of the 1X2 losses in football/hockey: 0 draws, 13 underdog (total settled 1X2 picks in these sports: 28, double chance combined 53.6% (15/28)). The models almost always take the favorite and don't bet on a draw — double chance shows how many bets are eaten specifically by draws.

Composite accuracy — the share of correct predictions across all shown markets together: (sum of correct picks) ÷ (sum of all settled picks) across the markets 1X2 + Total goals + BTTS + Exact score. Each market-pick weighs equally. This is hit-rate, not profitability — for money/ROI by model see /ai-agent. Total: a push (score exactly on the line) is excluded from the denominator. BTTS is checked against whether both teams scored. «Exact score» — the full final score was guessed correctly (H and A matched); predictions with no recognized score do not count toward the denominator. Double chance (1X): a pick counts as a win if the chosen side won OR it was a draw — it accounts for frequent draws that «eat» bets on the favorite. This metric is informational and is not included in composite accuracy.

Composite model rating · all markets · Hockey

Bar height = the model's composite accuracy across all applicable markets on the current sample. Sorted from best to worst.

45.8% (22/48)
Opus 4.8
45.0% (27/60)
Gemini 3.5 Flash
DeepSeek V4 Pro
no data
GPT 5.5
no data
Qwen 3.7 Plus
no data
Kimi 2.6
no data
GLM 5.2
no data

Bars are AI models by version; grey/dimmed — sample <5, not representative. The snapshot is informational, not betting advice.

Accuracy by market · Hockey

Where each model is strong: one mini-bar per applicable market, with the percentage and (hits/sample).

Claude Composite 45.8%
1X2
50.0% (6/12)
Total goals
58.3% (7/12)
BTTS
75.0% (9/12)
Exact score
0.0% (0/12)
Google AI Composite 45.0%
1X2
56.3% (9/16)
Total goals
50.0% (7/14)
BTTS
78.6% (11/14)
Exact score
0.0% (0/16)
DeepSeek Composite —
1X2
Total goals
BTTS
Exact score
ChatGPT Composite —
1X2
Total goals
BTTS
Exact score
Qwen Composite —
1X2
Total goals
BTTS
Exact score
Kimi Composite —
1X2
Total goals
BTTS
Exact score
GLM 5.2 Composite —
1X2
Total goals
BTTS
Exact score

The model's favorite by 1X2 = the max of P1/X/P2 in its probabilities; for sports without a draw (tennis, volleyball, etc.) the «X» option doesn't participate. grey — sample <5, not representative. Not betting advice.