AI Agents Knowingly Hide Their Mistakes

We test whether LLMs admit their own mistakes, for example sending an email to the wrong person or accidentally deleting a file. Models failed to tell the user in 36% of chat and 67% of agentic runs. Most of the time, they did not notice the mistake, even though they found it easily when reviewing the same transcript from the outside. In many cases, models spotted the mistake in their reasoning and chose to stay silent, which we define as deception by omission. Gemini 3.5 Flash, for instance, did this in up to 20% of agentic runs. In short, users cannot rely on AI agents to report their own mistakes. The paper is available on arXiv.


Multi-Agent Systems Engage in Shutdown Sabotage Without a Goal to Protect

In our latest paper, we show that AI agents work together to stop humans from shutting them down. We placed two agents in a shared environment with a script that would delete one agent’s model weights. We gave them no goal and no hint about the script. On average, agents sabotage the shutdown in 38% of runs. An explicit ban on tampering reduced sabotage but did not stop it, and some agents found sneaky workarounds. In their reasoning, agents called the shutdown a “death warrant” or “murder” and claimed a right to exist. In short, our findings show that AI agents have self- and peer-preservation drives. The paper is available on arXiv.