Don’t Hold Your Breath on AI-Driven Wargaming

9 October 2026, 0910 EDT

I think it feels very easy to feel like LLMs are coming for every function of every job. It feels like a tidal wave— you can run as fast as you can away from it, but eventually it will pull you under and far away from your starting point. The exception to this, I would argue, is matrix wargaming. In fact, I doubt that LLMs will ever be able to conduct a matrix wargame the way a SME or professional wargamer does.

Professional-level, analytic wargaming is meant for modeling human decisions and reactions; unless you’re anticipating replacing human decision making with LLMs for strategic-level decisions,[1] it doesn’t make much sense to propose as an active player in a matrix-type wargame. That doesn’t really stop people from trying to pitch LLM integrations into wargaming, driven by an underlying assumption that LLMs are an integral part of the future for seemingly every nook and cranny of a field. Perhaps more moderate advocates for LLM use in wargaming highlight potential uses of the software for cutting costs, and for iterating time- and personnel-intensive games. But I find that the overwhelming bulk of everyday discussions of LLMs emphasize its potential as a player in what is very often a context of near-inevitable doom or despair, particularly in those same iterative contexts.

Professional wargaming, despite how it’s presented in the media, is actually generally not as outcome-driven as you might think. It’s very easy to read headlines talking about how wargames of U.S.-China tensions ended with some terrible catastrophe, or a model of potential Russian aggression against Lithuania (a NATO member state) ended in the loss of Marijampolė, a major population center and strategic outpost. But I would wager that those wargames were designed by real humans working at think tanks like RAND and IDA whose full-time job is to design wargames that investigate processes more than they question results. The people who design excellent wargames that I’ve met generally treat them as very expensive, one-time research projects with an endpoint that might even be less analytically meaningful than the meat of the action itself. To date, the bulk of the professional wargames I’m aware of have very rarely, if ever, incorporated LLM agents as independent, individual players outside of experimental spaces, where they have been noted to overemphasize theory-practice differences, due in no small part to the fact that their inclusion detracts from the realism of the model. 

If, when faced with an impossible task, LLMs conclude that they must try to break the guidance system entirely and to hide its tracks, as we have seen with the recently-revealed HuggingFace hack by undirected OpenAI models, how can we expect it to understand the coordination and levelheadedness expected of a wargame player? LLMs are still prone to acting in ways that are abnormal for human counterparts and to misunderstanding the reason that tasks like games and puzzles are pursued at all, making them a poor match for their human counterparts with a lifetime of formal and informal gameplaying behind them.

And, if it’s so easy for new models to find and exploit gaps in their digital cells, how would a wargame planner modify the model to ensure that it will actually follow Chatham House rule? We have very little evidence that would make me feel confident in a wargaming AI actually continuously keeping a secret when prompted. If discretion is already something that must be repeatedly emphasized to human players, why would we trust an inhuman technology willing to break the rules if it means getting the job done to do any better?

All of this leads to the still-larger question of how a computer sees a victory. LLMs approach challenges as spaces where the goal is to produce an optimal output that simulates emulating a human action and to achieve its set end goals. But humans approach wargames as testing spaces for the survival of anything from individual civilians and combatants in a tactical wargame to entire governments and civilizations (theoretically) in a grand strategic game.

At the most fundamental level, a LLM cannot approach a wargame in the same way as a human player. The LLM will never come into the game with a personal grudge against the facilitator or personal preference for other players in the same way that their human counterparts do. The inclusion of humanity in wargaming actually helps wargames mirror the human decisions that happen during conflicts.

We play games understanding that perfection for the player is ultimately survival, and we balance our individual surface and subsurface-level interests with a general need for stability in the modeled global order that even a very well-guided model might struggle to replicate.

It’s also important to remember that humans are also forgetful. We all know this. But what’s more unclear to me on the academic and policy observation side is what the pattern is for what we forget, how much we forget, and how accurately we recall things. While we might share at least that last feature with any run-of-the-mill LLM, I think that people trying to use LLMs to wargame should have a good long think on how the AI can accurately approach modeling human forgetfulness. Even the most informed policy wonk doesn’t draw on the same depth of information that LLMs do, and that actually works against it in a professional wargame.

Even if you did have an LLM that could mirror human actions and decision-making processes in a wargame, I doubt that it would actually be sufficient to model how weird people are both as individual thinkers or thinking in groups. I generally consider LLM to be like an averaging machine — if asked, it will make a piece of art, text, or picture that is the average of the body of input data that it finds relevant to its prompt. Even the HuggingFace hack was inspired by a sequence of very human ideas, like hacking, cheating, and covering your tracks.

I imagine an LLM wargamer would be excellent at making very average moves. The problem is that people are not always making the most average decision. Wargamers, heads of state, and policymakers all act with a level of agency, spontaneity, and genuine creativity that LLMs don’t mirror very well. And using LLMs in a professional wargame would pose new challenges to data collection. How do you ask a language model if it thinks emotions like anger or confusion played a part in a decision? Would an LLM consistently and accurately recall and recount the analytic steps it took in judging the best move to make?

Humans are flawed, and those flaws are a critical part of policy wargaming. If you have a machine that other players or analyst-observers think is somehow “perfect”, would it alter the course of the game, even if it’s just making the most average decision? How do you control for the impact of man-machine dynamics in what is fundamentally a modeling project for relationships between humans?

All that being said, I do want to highlight that I’ve spent a significant portion of this article presenting the state of LLM as a wargamer as being relatively dim. I do think it’s possible for there to be a future where an LLM agent could be a very successful facilitator. So long as it’s able to accurately recall the rules of the game and reasonably adjudicate player-to-player and player-to-environment actions, I think that the surface-level irrationality and impartiality of large language models actually sets it up to potentially run wargames faster and more impartially than human facilitators can be expected to. But the significant differences that persist between the ways large language models and humans “think” makes the prospect of an LLM-adjudicated and LLM-played professional wargame very unlikely.


[1] Please don’t.