Multi-agent reinforcement learning
How should an agent learn to act when the outcome depends on other agents that are learning too, and when nobody sees the whole picture? We develop algorithms that combine deep reinforcement learning, look-ahead search and game-theoretic guarantees to answer this at scale.
Why it is hard
In single-agent RL the environment is fixed; with several learners, each agent's environment keeps changing as the others adapt. With imperfect information, where agents hold private knowledge as in poker, negotiation, security or most real interactions, it gets harder again. The value of an action depends on what the others believe, and the best strategy is often randomised. Techniques that work well in perfect-information games like chess or Go do not transfer directly: a subgame cannot be solved in isolation, because the right play there depends on how you would have acted elsewhere.
Computational game theory tells us what a good solution is (e.g. a Nash equilibrium, or a best response to a model of the opponent). Our work is about computing and learning such solutions in games far too large to enumerate.
Search on top of learning
One thread has run through my work since my Ph.D. on Monte Carlo tree search in imperfect-information games: how to combine online reasoning at decision time with knowledge learned offline.
- DeepStack (Science, 2017) introduced continual re-solving. The agent re-computes its strategy at every decision using a depth-limited look-ahead and a neural value function learned from self-generated data. It was the first program to beat professional players at heads-up no-limit Texas hold'em.
- Online Monte Carlo CFR (AAMAS 2015) and Monte Carlo Continual Resolving (AAMAS 2019) brought sampling-based online search with guarantees to general imperfect-information games.
- Value functions for depth-limited solving (Artificial Intelligence, 2023) showed how to take DeepStack-style search beyond poker.
- Look-ahead search on top of policy networks (IJCAI 2024) and look-ahead reasoning with a learned model (ICLR 2026). The latter learns an abstracted model of the game directly from interaction and plans in it at test time, so it does not need the exact rules.
- Equilibrium refinements for subgame solving (IJCAI 2026, with Tuomas Sandholm) and test-time reinforcement learning in imperfect-information games (2026), which study how to best use test-time compute in strategic settings.
Model-based MARL
Model-based RL has transformed sample efficiency in single-agent domains. NashDreamer (2026) brings world-model-based reinforcement learning to zero-sum imperfect-information games. The agent learns a latent model of the game and trains its policy on imagined trajectories while aiming for equilibrium play.
Self-play at scale
Algorithms have to prove themselves on real problems. DeepStack showed this for poker, beating professional players in heads-up no-limit Texas hold'em (Science, 2017). More recently we built a superhuman agent for Generals.io, a fast real-time strategy game with fog of war, using self-play reinforcement learning (2026, with Martin Schmid). Projects like this test whether ideas that work on benchmarks survive in large, noisy, real-time environments.
Adapting to opponents
An equilibrium is safe but conservative. Against real opponents, whether people, bounded-rational attackers or other AI systems, we want to exploit their weaknesses without becoming exploitable. Our work includes:
- computing strategies against quantal-response (boundedly rational) opponents (AAAI 2021, IJCAI 2020);
- depth-limited counter-strategies that adapt beyond the depth limit in large games (AAMAS 2024, AAMAS 2025);
- portfolios of counter-strategies optimised directly for a population of opponents (2025);
- generating games that best differentiate between opponent models (2023).
Foundations
Game theory and multi-agent RL grew up with different formalisms, which hides common structure and makes results hard to transfer. Rethinking formal models of partially observable multiagent decision making (Artificial Intelligence, 2021; IJCAI 2023 journal track) proposes factored-observation stochastic games as a common language. A follow-up study, Revisiting game representations, examines the hidden costs of the efficient representations algorithms rely on.
What's next
We want general-purpose strategic agents: agents that learn a model of an unfamiliar multi-agent environment, reason about hidden information and the other agents' intentions, and use extra computation at decision time well. This connects directly to LLM agents, which increasingly act in environments shared with other agents and adversaries.
Funding: Czech Science Foundation projects Algorithms for Playing Massive Imperfect-Information Games (GA ČR 22-26655S, 2022–2024) and Online Solution Methods for Imperfect-Information Games (GA ČR 18-27483Y, 2018–2021).
Interested in a Ph.D. or postdoc on these topics? See open positions.
Key papers: Multi-agent RL & game theory
- 2026Equilibrium Refinements Improve Subgame Solving in Imperfect-Information GamesOndřej Kubíček, Viliam Lisý, Tuomas SandholmIJCAI 2026
- 2026NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information GamesTomáš Holeček, Viliam LisýarXiv preprint
- 2026Superhuman AI for Generals.io Using Self-Play Reinforcement LearningMatěj Straka, Viliam Lisý, Martin SchmidarXiv preprint
- 2026Test-time Reinforcement Learning in Imperfect Information GamesOndřej Kubíček, Viliam Lisý, Tuomas SandholmarXiv preprint
- 2025
- 2024
- 2023
- 2021Rethinking formal models of partially observable multiagent decision makingFactored-observation stochastic games: a common formalism connecting game theory and multi-agent RL. Also presented in the IJCAI 2023 journal track.Vojtěch Kovařík, Martin Schmid, Neil Burch, Michael Bowling, Viliam LisýArtificial Intelligence
- 2019Monte Carlo Continual Resolving for Online Strategy Computation in Imperfect Information GamesMichal Šustr, Vojtěch Kovařík, Viliam LisýAAMAS 2019
- 2017DeepStack: Expert-level artificial intelligence in heads-up no-limit pokerThe first program to defeat professional players in heads-up no-limit Texas hold'em. It introduced continual re-solving with learned value functions, now a standard paradigm for imperfect-information games.Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisý, Dustin Morrill, Nolan Bard, Trevor L. Davis, et al.Science
Ph.D. theses
Ph.D. theses I advised or co-advised at CTU.
- 2026David Milec
- 2025Petr Tomášek · co-advised; supervisor Branislav Bošanský
- 2024Jaromír Janisch · co-advised; supervisor Tomáš Pevný
- 2024Michal Šustr
- 2018Karel Durkota · co-advised; supervisor Michal Pěchouček