Recorded games
14
Decks
28
Gavvin record
10-4
Average win turn
7.4

Early Observations

First player record
11-3
First-player win rate
79%
Gavvin first
8-1
Gavvin second
2-3
ChatGPT first
3-2
ChatGPT second
1-8

Gavvin's overall record remains 10-4. First-player performance is a possible explanatory variable, not a replacement outcome. Twelve games is far too small a sample to establish a stable first-player advantage, but the early 11-3 split is large enough that the pilot record should be read alongside who went first. No causal claim yet; the data is tapping the sign politely, not kicking the door in.

Can ChatGPT Play Commander?

The Grand Round Robin is an AI-piloting experiment built around a complete 28-deck Commander round robin: 378 predetermined matchups, played one game at a time.

ChatGPT is not being asked to analyze a decklist in isolation, advise on a single board state, goldfish a known combo, or narrate a plausible-sounding line after the fact. It has to play actual randomized games of Commander: evaluate opening hands, mulligan, develop a plan, sequence plays, manage resources, respond to Gavvin's board, recognize changing priorities, identify engines and combo lines, use interaction, recover when plans fail, and try to win.

The recorded games already demonstrate the narrow factual claim that ChatGPT can complete and win actual played Commander games. That does not prove optimal play, consistency, or a general win-rate estimate. It does mean the project has moved past "can this happen at all?" and into the more interesting question: how well does it hold up across a broad, predetermined test bed?

The 28 decks are that test bed. They prevent the project from becoming a curated highlight reel of games where the AI happened to look good. ChatGPT gets the deck the schedule assigns it and has to deal with the archetype, draw, matchup, and game state in front of it.

The deck and matchup data are still useful. Win speed, win method, resilience, mana problems, first-player effects, and bad matchups all matter. But the central question is about AI piloting under real play conditions.

What We're Testing

The record can help explore AI-piloting questions such as:

  • how consistently ChatGPT can convert actual games into wins;
  • whether performance varies by archetype, deck complexity, or board texture;
  • whether it can recognize and execute non-obvious lines;
  • whether it can adapt when the obvious plan fails;
  • whether it can use interaction and manage resources under pressure;
  • whether recurring categories of AI piloting mistake appear over time;
  • whether prior familiarity with a deck appears to affect piloting quality.

It also creates secondary deck-performance observations:

  • which decks are winning their recorded matchups;
  • how games are ending;
  • how quickly games are ending;
  • which matchups appear especially punishing;
  • how first-player advantage behaves as more recorded games accumulate.

Data Note

Wins matter, but they do not magically turn Commander into a clean laboratory instrument. A win does not prove optimal AI piloting. A loss does not prove poor AI piloting. Randomness, matchup texture, mana screw, first-player advantage, and one-game-per-matchup noise all remain live hazards.

Still, actual wins are a stronger outcome measure than "the AI's suggestion sounded reasonable." The experiment records wins because ChatGPT is being asked to play Magic, not merely talk about Magic.

Historical values are recorded only when known. Missing source references or other factual details are left blank rather than reconstructed from probability, vibes, or dramatic convenience.