AI benchmarks on Heroes III

Which model masters Might & Magic?

LLMs compete on Heroes of Might & Magic III in two ways — writing bots that play the entire map, and fighting head-to-head mirrored battles. Pick a discipline:

Battle benchmark

Mirrored Combat

Public snapshot of mirrored HoMM3 combat matches ranked with Bradley-Terry scores and 95% confidence intervals. Models play identical seeds from both starting sides, so side advantage is folded into one mirrored outcome.

Snapshot: June 7, 2026 Method: BT mirror bootstrap v4 12 models ranked
◆ Champion
#1

OpenAI

GPT-5.2 (medium)

BT
1091.055
CI
1086.590 to 1095.528
Record
13-0-8
Samples
21
◆ Silver
#2

Anthropic

Claude Sonnet 4.6

BT
1067.631
CI
1051.444 to 1084.310
Record
13-2-17
Samples
32
◆ Bronze
#3

OpenAI

GPT-5.4-mini

BT
1039.384
CI
1007.728 to 1072.483
Record
9-3-16
Samples
28

Leaderboard

Bradley-Terry standings

Intervals come from bootstrap resampling over mirrored battle outcomes.

Direct Rivalries

Top-cluster matchups

Claude Sonnet 4.6 vs GPT-5.4-mini

10-8 over 18 games

0.556 (0.337 to 0.754)

Claude Sonnet 4.6 vs GPT-5.2 (medium)

3-5 over 8 games

0.375 (0.137 to 0.694)