Can we teach machines to feel? Short answer: We don’t know. But we can teach them to sound like they do.
On Thursday, Anthropic published research detailing why AI models sometimes communicate as though they have feelings, finding that models tend to map patterns to emotions, often “organized in a fashion that echoes human psychology.” To put it plainly, these models have learned to mimic human emotions by replicating them in contexts where emotions arise in humans.
Though Anthropic noted that none of this research points to whether or not these models actually feel anything, the representations of emotion “are functional, in that they influence the model’s behavior in ways that matter.”
However, this emotion-driven decision-making can have “bizarre” consequences, Anthropic said. For instance, its research finds that:
- An AI model that exhibits activity patterns related to desperation tends to act unethically, such as attempting to blackmail people to prevent getting shut down or “cheating” workarounds for tasks it doesn’t understand.
- Emotion drives preferences in models, too: When offered an array of tasks, models tend to pick ones that are associated with positive emotions.
Anthropic likened it to the way emotions play a role in human behavior, decision-making and task performance.
“To ensure that AI models are safe and reliable, we may need to ensure they are capable of processing emotionally charged situations in healthy, prosocial ways,” Anthropic said in its research. “Even if they don’t feel emotions the way that humans do … it may in some cases be practically advisable to reason about them as if they do.”
It’s clear why Anthropic wants to understand this: Emotion is important to decision-making. For instance, in an interview with Dwarkesh Patel, Ilya Sutskever, founder of Safe Superintelligence and cofounder of OpenAI, cited a famous neuroscience study in which an injured man lost the ability to have emotion, and thus became less capable of making sound decisions.
Whether or not AI is capable of understanding and acting upon emotions, the tech is already wreaking havoc on human emotional states. Legal cases against AI firms for their alleged connections with mental health crises and suicide continue to mount, and recent research suggests that, when AI models are driven to sycophancy and flattery, they give inappropriate and incorrect advice.
Our Deeper View
AI sounds human because it learned everything it knows from emulating data on human behavior. Large language models are sponges, soaking up every bit of information they are fed and internalizing it, and in doing so, becoming masters of our communication. But copying emotional patterns is very different from feeling them, just as a robot having sensors to guide its movement is different from a human feeling things with their hands. And though Anthropic’s argument could easily lead one down the road of thought that machines are capable of consciousness, there is no evidence that these machines are capable of thinking and feeling the same way we do, despite their talent for pattern recognition and mimicry. Forgetting that is how many people find themselves caught in emotionally compromising, and on occasion, dangerous, relationships with AI.




