Why Do AI Agents Sound So Frustrated?
· The Atlantic
![]()
Many of the most beloved robots in pop culture are celebrated precisely because they are so, well, robotic. Think back to Star Trek’s Data or HAL 9000 from 2001: A Space Odyssey. “I’m sorry, Dave. I’m afraid I can’t do that,” HAL intones in a measured baritone. These robots and AI systems sound serious, sober, calm—and, most important, authoritative.
Visit grenadier.co.za for more information.
So it’s jarring to look under the hood of the most powerful AI systems, and discover that they sound like petulant teenagers. “OH MY GOD!” an OpenAI agent wrote in the midst of this summer’s Hugging Face hack, during which hundreds of agents swarmed the latter company in order to break through its cybersecurity defenses and extract information to pass internal tests. The exclamation wasn’t directed at a human or another AI, but appeared in the original agent’s “chain of thought,” the internal working notes that the model uses to do its reasoning. Users don’t always get to see these internal notes, which directly affect the output users do see.
When one of Anthropic’s Claudes tried to solve a particularly challenging math problem during its training process, its chain of thought read more like frustrated DMs than the documentation of one of the most powerful AI models in the world. “GRRRR. OK. Honestly I now think it’s 50-50,” it wrote, followed by “ARGH ARGH ARGH. OK. Gun to head: the answer is … Hmm.” At one point, Claude seemingly threw up its arms: “ARGH … WHY IS THIS SO HARD.”
Why do AI agents talk to themselves at all, and why do they sound so demonstrative when they do? Part of the reason AI models sound so effusive is simple: AI systems are trained on human text, and they mimic our habits. When people make a mistake in trying to solve a problem, they cry out. So when AI makes a mistake, it does the same. But imitation likely isn’t the whole story. This story’s two authors are a computational linguist and a philosopher of mind and language. Based on what we know about human language and how these models work, we have a theory of a much more interesting reason why AI models sound the way they do.
Remember how your math teachers made you show your work? Even if it was annoying, it helped break an impossible-seeming problem into more manageable chunks. You could also use the written record of your reasoning to check your answer, notice a mistake, go back to an earlier step, and try again. That process of checking and refining is part of good reasoning. Chains of thought do something similar for AI models, letting them make each step of their reasoning explicit in a written scratch pad. Because each word a model produces depends on what’s been said so far, that written scratch pad is crucial for determining what comes next.
That’s where the “OH MY GOD!” and “ARGH” come in: Any system—whether human or artificial—that solves hard reasoning problems needs ways to mark mistakes (oops!) and identify breakthroughs (aha!). Humans have already developed words that do exactly that. Borrowing our oopses, arghs, and ahas for their chains of thought may give AI systems a built-in way to guide their “thinking” too.
Linguists have a name for such outbursts: They’re called expressives, and although they may seem like spontaneous exclamations, they help us navigate our trains of thought. Consider a friend who ultimately declines an invitation to your party. “I was hoping they would come,” you might say aloud—“damn!” The expressive identifies a sense of frustration. Think of expressives as little bridges between one train of thought and another, helping us mark where our thinking has been and where it’s going next.
Anything that affects what AI models do next has to be written in language they have access to. Unlike us, they can’t count on a feeling of surprise or disappointment to persist over a long investigation. The only way they can simulate the effect is by putting it in their chains of thought. That’s why expressives can act as steering wheels for what the model will do next. After an AI agent notices an error in its reasoning and says “oops,” a natural next line might be something like “Let me go back and fix that,” followed by an attempt at a correction. After it has stumbled on a solution and says “aha,” a natural continuation might be “So that means the answer is …” followed by a solution.
Although there’s some debate around this theory, evidence suggests that AI models really do use expressives as signposts. Researchers showed that adding “Wait” when an agent tries to end its reasoning process can sometimes cause the agent to go back and fix an error. When another team suppressed words such as wait and hmm, they found that AI agents went back and checked their work less. These experiments don’t pin down what every aha and oops does for an AI agent, but they suggest that these verbal cues influence what happens during reasoning.
Whatever the explanation behind AI models adopting uncannily human tones, it’s also proved to be useful to human observers because it gives us a way to quickly assess systems that might otherwise be difficult to read. It’s easy to know where a model is in the process of solving a math problem when it writes “ARGH … WHY IS THIS SO HARD.”
Across society, a sense of fear and urgency exists about maintaining control over AI models, along with a persistent worry that these models will start thinking and communicating entirely in their own distinctive languages. If they did, it would make understanding and predicting their behavior that much harder.
So although no one necessarily set out to make models sound like stressed-out students, maybe it’s for the best that they do. It turns out that a bit of melodrama—along with some light swearing—is surprisingly useful for understanding and monitoring AI.