Multi-agent reinforcement learning
Scalable algorithms for agents that learn to act among other learning agents, especially when each sees only part of the world. We combine deep RL with look-ahead search and game-theoretic guarantees.
- Search on top of learning: from DeepStack and continual resolving to look-ahead search over learned abstract models and test-time RL.
- Model-based MARL: world models for zero-sum imperfect-information games (NashDreamer).
- Self-play at scale: from DeepStack, which beat professional poker players, to a superhuman agent for the real-time strategy game Generals.io.
- Adaptation: robust counter-strategies that exploit opponents without becoming exploitable themselves.
- Foundations: formal models of partially observable multi-agent decision-making.
Open question: how can agents exploit test-time compute and learned models in strategic settings as effectively as they do in single-agent ones?