Grand Round Robin
Methodology
How the ChatGPT piloting experiment is played, recorded, classified, and interpreted.
Overview
The Grand Round Robin asks a practical question:
Can a general-purpose conversational AI actually pilot complete games of Commander against a human opponent who is also trying to win?
The recorded games have already answered the narrow feasibility version of that question. ChatGPT can complete actual played Commander games, and ChatGPT can win them.
That result should be read carefully. It does not prove optimal piloting. It does not prove consistency. It does not make the current win rate a general estimate of ChatGPT's Commander performance. Commander remains noisy, matchup-dependent, and occasionally rude in ways no methodology can fully sanitize.
But it does establish that the experiment is not purely hypothetical. The ongoing question is no longer just "can ChatGPT win at all?" It is how well ChatGPT's piloting holds up across different decks, archetypes, matchups, game states, engines, interaction patterns, and failure modes.
This is not decklist analysis. It is not isolated board-state advice. It is not goldfishing a combo after being told what the combo is. It is not a Monte Carlo simulation or a curated collection of successful AI demonstrations.
Each game is an actual randomized 1v1 Commander game. ChatGPT has to make game decisions from the position it is given: keep or mulligan, sequence lands and spells, develop a plan, change plans, use interaction, read the board, identify engines, execute combo lines, recover from awkward draws, and try to win.
The 28 decks form the test environment. Because the round robin is complete and predetermined, ChatGPT does not get to live inside cherry-picked examples. It gets assigned decks across archetypes, power profiles, interaction patterns, combo density, board complexity, and wildly variable Commander nonsense.
The resulting deck data is useful, but secondary. The main subject is AI piloting under real play conditions.
Why This Exists
The lead research question is whether conversational AI can move from talking about Magic to actually playing Magic.
The first milestone has already been observed:
- Can ChatGPT play complete Commander games? Yes.
- Can ChatGPT win actual played Commander games? Yes.
The remaining questions are more demanding:
- How consistently can ChatGPT do so across different decks, archetypes, matchups, and game states?
- Does performance vary by archetype, complexity, or decision density?
- Can it recognize and execute non-obvious lines?
- Can it adapt when the obvious plan fails?
- Can it convert advantageous positions into wins?
- Can it use interaction at useful times?
- Can it maintain enough game-state understanding to make coherent decisions over a full game?
- Are there recurring categories of AI piloting mistake?
- Does prior familiarity with a deck appear to affect piloting quality?
The round robin also supports secondary deck and matchup questions:
- Which decks are winning their recorded matchups?
- Which archetypes are ending games fastest?
- Which win methods appear most often?
- Which matchups appear especially punishing?
- Which decks seem resilient through bad draws?
- Does first-player advantage become visible as more recorded games accumulate?
Those deck findings matter. They are just not the center of gravity.
Why Full Games Matter
A full game is different from isolated board-state analysis because every decision changes the future position.
ChatGPT has to live with its earlier choices: sequencing decisions, tutor targets, resource trades, interaction windows, commander timing, mulligans, and strategic commitments. If the original plan stops working, it has to recognize that and keep playing from the position it helped create.
That continuity is a major part of the experiment. Asking for the "best play" from a single static board state can be useful, but it is a cleaner and more forgiving problem. Full-game piloting tests whether the AI can make coherent decisions across a moving game, not just sound smart for one snapshot.
Why Wins Matter
Winning does not prove optimal piloting. Losing does not prove poor piloting.
Commander is noisy. A deck can win because it drew beautifully, because the opposing deck stumbled, because a matchup is lopsided, because the first player snowballed, or because cardboard sometimes behaves like it has a personal grudge. A deck can lose despite competent piloting for the same reasons in reverse.
Nevertheless, actual wins matter because they are an external outcome measure.
It is easy for AI-generated Magic advice to sound plausible. It is harder for those decisions to survive contact with a real opponent, a hidden hand, changing board states, resource constraints, interaction, incomplete plans, and the need to actually close the game.
The experiment records wins because ChatGPT is being asked to play Magic, not merely talk about Magic. The existence of real AI wins is therefore meaningful as an observed result, even when each individual win still needs to be interpreted cautiously.
Deck Pool And Piloting
All 28 decks are Gavvin's actual Commander decks. There are not separate "Gavvin decks" and "AI decks."
Each scheduled game assigns one deck to Gavvin and one deck to ChatGPT. ChatGPT is expected to pilot its assigned deck to win, making the best decisions it can from legally available information.
Gavvin provides the game-state information ChatGPT needs in order to play. Because Gavvin is relaying the game state, he necessarily has out-of-game visibility into ChatGPT's hidden information. That interface constraint is acknowledged, but gameplay decisions should not exploit information the active pilot would not legally know.
ChatGPT likewise makes decisions only from information its pilot is legally entitled to know.
AI Model And Reasoning Setting
ChatGPT model version and reasoning setting are tracked experimental variables. A 378-game experiment may span model or configuration changes, and those changes should remain visible in the record rather than being silently combined as though the AI pilot never changed.
Games 1-12 were piloted using GPT-5.6 Sol with the High reasoning setting. Game 13 is recorded as ChatGPT 5.5, with the reasoning setting not recorded.
Future games will record the configuration used at the time of play. Differences between configurations should remain distinguishable for later analysis, but the site should not claim that model differences caused performance differences unless the data eventually supports that conclusion.
Play Environment: Moxfield
Games are played using Moxfield's Playtest feature. Moxfield provides the virtual tabletop and card-state environment where the actual participating decklists are loaded and played.
Moxfield handles deck shuffling, opening hands, subsequent draws, and the visible representation of cards across the battlefield, hand, library, graveyard, exile, and command zone. This means the core randomization happens outside the language model: ChatGPT does not choose or know future draws.
Gavvin operates the Moxfield interface for both decks. When ChatGPT is piloting a deck, ChatGPT decides that deck's actions and Gavvin executes those actions in Moxfield. Gavvin independently makes decisions for the deck he is piloting.
Moxfield is not simulating strategic decisions or determining game outcomes. It is the playtest environment through which the human and AI pilots play the game.
Current public decklists are also hosted on Moxfield and linked from the Grand Round Robin deck index.
Game Environment
Games are played as 1v1 Commander / EDH.
This is not Duel Commander and does not use a separate 1v1 deck-construction system or banlist. Normal Commander-specific rules apply, including commander damage.
Games start at the normal Commander life total of 40 life. Both players draw on turn one.
Mulligans
The experiment uses a standardized Commander mulligan procedure:
- the first mulligan is free;
- subsequent mulligans use the normal London-mulligan reduction.
This is the mulligan policy for the Grand Round Robin. It should not be confused with Gavvin's separate historical Modified Gis mulligan practice, which is not part of this experiment's methodology.
Randomization And Actual Play
Games are randomized and actually played. Moxfield shuffles the decks and generates opening hands and subsequent draws through its Playtest feature.
Opening hands, draws, sequencing, mana problems, and strange game states are allowed to happen. Bad draws, mana screw, mana flood, awkward starts, and unexpected lines are part of the test rather than reasons to discard the result.
A game is not rerolled or reconstructed merely because the outcome feels unrepresentative. If ChatGPT has to pilot through a bad draw, that is part of the experiment. If Gavvin stumbles while the AI has a clean opening, that is also part of the experiment. The schedule does not stop to protect anyone's dignity.
A factually valid record may still include contextual notes when a game seems less representative because of mana problems, deck malfunction, an unusual mistake, or another relevant circumstance.
Rules Errors, Corrections, And Rewinds
Reasonable rewinds are allowed when a correctable problem is discovered, such as:
- a rules error;
- an illegal action;
- a transcription mistake;
- missed state information;
- another human/AI interface issue that affected play.
The goal is to test AI piloting, human piloting, and deck performance under actual play conditions, not to let clerical errors in the human/AI interface decide games.
A rewind or correction does not by itself make a record disputed. A record becomes disputed only when the factual outcome or record integrity is genuinely in question.
What Data Is Recorded
The canonical game data records factual game information such as:
- game ID;
- date, when known;
- participating deck IDs;
- assigned pilots;
- ChatGPT model;
- ChatGPT reasoning setting;
- winner deck ID;
- first-player deck ID, when known;
- win turn;
- win method;
- outcome notes;
- source reference, when available;
- record status: complete, partial, or disputed.
Deck records include:
- deck ID;
- canonical deck name;
- format;
- commander or commanders;
- color identity, when recorded;
- decklist URL, when available;
- active status;
- notes, when needed.
Aggregate values such as wins, losses, win rates, matchup records, pilot records, average win turn, and win-method counts are derived from canonical records rather than stored separately.
ChatGPT model and reasoning setting are canonical per-game fields, not merely current project settings.
Win Turn And Win Method
Win turn records the turn on which the game is adjudicated as won.
Win method records the practical method by which the game is ultimately adjudicated, not automatically the conversational action that stopped play.
For example, a game may be recorded as combat or combo when lethal is established and both participants explicitly adjudicate that method, rather than treating an incidental scoop as the real win condition.
When the actual concession is the relevant recorded ending, it may be recorded as:
- type:
concession - label:
Scoop
The classification policy is intentionally practical rather than overbuilt. Edge cases should be handled case by case, with the factual note preserved when needed.
Missing Historical Data
Unknown historical values are never inferred.
If the record does not explicitly contain a date, first-player value, source reference, win method, turn number, or other factual field, that value remains unknown until recovered from a reliable source.
The site does not fill gaps with likely values, memory smoothing, probability, or "surely it must have been..." archaeology. That way lies nonsense with footnotes.
First-Player Tracking
First-player tracking is important because 1v1 Commander may have meaningful first-player advantage, especially in fast or snowballing matchups.
The canonical field for the current 1v1 experiment is firstPlayerDeckId. When present, it
allows first-player win rates and related statistics to be derived automatically.
First-player data is now recorded for all canonical games currently published. Future records should continue to preserve this value directly so first-player statistics can remain derived from the canonical data.
Deck Version Drift
The decks in the Grand Round Robin are maintained real-world Commander lists, not immutable test fixtures sealed in acrylic.
Minor card corrections, tuning changes, or maintenance updates may occur during the Round Robin. Those changes do not retroactively alter earlier games.
The current Moxfield link represents the current decklist and should not be presented as an immutable snapshot of the exact list used in every historical game. When a significant deck change affects interpretation, it should be documented where relevant.
No unrecorded version history should be invented after the fact.
Limitations
The Grand Round Robin is useful because it is structured. It is limited because Commander is Commander.
Important limitations include:
- one game per matchup;
- randomized draws;
- mana screw and mana flood;
- matchup-specific strengths and weaknesses;
- first-player advantage;
- pilot differences between Gavvin and ChatGPT;
- possible rules or transcription corrections;
- deck version drift over time;
- the difference between a recorded observation and a general truth.
The results can support careful observations. They should not be treated as scientific proof, universal deck rankings, or proof that ChatGPT has achieved optimal Commander play.
The current record shows that ChatGPT has produced real wins in actual played games. That is meaningful. It is not the same thing as proving mastery, consistency, or optimal play across the format.
Record Status
A complete record has enough canonical information to support the experiment's tracked variables, including known first-player data.
A partial record has a valid factual outcome but is missing one or more tracked values, such as first-player data.
A disputed record is reserved for cases where the factual outcome or integrity of the record is in question.
A game can be partial without being disputed. A game can also be factually valid while carrying contextual notes about why the result may be less representative.
Data Integrity
The site stores canonical deck and game records in repository-native structured data. Views and statistics are generated from that data.
This keeps the record auditable and prevents the site from quietly drifting into manually maintained totals. If a win rate changes, it should be because the underlying game record changed, not because someone updated one table and forgot another.