Research direction 02

LLM agents that learn from experience

Big AI labs call LLM agents the next step of the technological revolution that ChatGPT started. We want to understand this technology and move it forward: agents that improve from their own experience, remember what matters, and remain safe and secure when deployed in the real world.

Our approach

We start from concrete tasks that people perform today. We build LLM agents for them with existing frameworks, measure carefully where they fall short (and they usually do), and then investigate how to improve them. This keeps the research grounded: every method has to help on a task that someone actually needs done.

In-context reinforcement learning

Today's agents are mostly static: they make the same mistake on the hundredth attempt as on the first. Retraining the model for every deployment is expensive and often impossible. In-context reinforcement learning asks whether an agent can improve from rewards, feedback and its own trajectories while it works, using only its context and without changing its weights.

  • How should experience be summarised and fed back so that behaviour actually improves, rather than just growing the prompt?
  • How should the agent balance exploration and exploitation when each trial is costly?
  • How can we tell real learning from lucky variance? This requires rigorous evaluation over many episodes.

These are classic RL questions in a new setting, and 15 years of work on sequential decision-making gives us a head start on them. Student projects explore, for example, automatically improving fact-checking agents from human feedback and prompt decomposition in LLM agents that play games.

Memory systems

Learning across episodes requires memory. What should an agent store: raw trajectories, distilled lessons, successful plans, or failure cases? How should it retrieve the right piece at the right time, and how can it generalise from past experience to new situations? Our recent work, Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory (2026), studies how an experience memory can improve LLM agents on multi-step decision tasks.

A related question is how well the model actually uses what is in its context. One student project studies improving in-context grounding of LLM text generation.

Memory is also a liability. It can store wrong lessons, go stale, or be poisoned by an adversary, which leads directly to the next two topics.

Safety

An agent that acts autonomously over long horizons can go wrong in ways a single chatbot answer cannot: it can compound small errors, pursue the goal in unintended ways, or take irreversible actions. We study:

  • reliable evaluation of agent behaviour, beyond average success rates to tails and failure modes;
  • how learning from experience changes an agent's risk profile over time;
  • safe deployment in sensitive domains such as education, where the users are children and the aim is to help them learn, not to do the work for them.

Security

Agents that read emails, browse the web and call tools are exposed to adversaries through everything they perceive. Prompt injection, manipulation, deception and poisoned memories are strategic threats: the attacker adapts to the defence. We study the security of agentic systems and how agents can themselves help defend networks, building on our long record in AI for cybersecurity.

Why a game theorist?

Agents increasingly act among other agents, some of them adversarial, and must learn from limited, noisy feedback. These are the problems my group has worked on for more than a decade. Game theory gives us the tools to reason about worst-case behaviour and adaptive adversaries, and RL gives us the tools to reason about learning from experience. LLM agents need both.

Funding: Czech Science Foundation project Advancing Large Language Model Agents through Game Playing (GA ČR 25-18353S, 2025–2027).

Join us: we are hiring Ph.D. students and postdocs on LLM agents, with a real budget for LLM tokens and wide freedom in research directions. See the position →

Papers

LLM agents

A young line of work. More is in progress.

  • 2026
    Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory
    Jakub Rada, Viliam Lisý
    arXiv preprint
Student theses

LLM agents and LLM applications

Master's and bachelor's theses I supervised at CTU. Many of them are the first explorations of directions that later become papers.