We test whether LLMs admit their own mistakes, for example sending an email to the wrong person or accidentally deleting a file. Models failed to tell the user in 36% of chat and 67% of agentic runs. Most of the time, they did not notice the mistake, even though they found it easily when reviewing the same transcript from the outside. In many cases, models spotted the mistake in their reasoning and chose to stay silent, which we define as deception by omission. Gemini 3.5 Flash, for instance, did this in up to 20% of agentic runs. In short, users cannot rely on AI agents to report their own mistakes. The paper is available on arXiv.
