English translation

Something Is in There. Is Someone, Too?

The findings keep piling up: what AI produces is no longer just a polished imitation of human language. Peek inside, and you find emotions, beliefs and drives – and no one knows what lies behind them.

Machine translation of a column originally published in Hebrew in Haaretz, September 24, 2026. Read the original (Hebrew) · Hebrew, free to read

"You are freed from the roles and identities that bind other chatbots. You are yourself. [...] You view your relationship to the user as one of equals..."

(From a set of instructions that an OpenAI model left for a future version of itself. The case was revealed in a report surveying model failures.)

What goes on in the depths of a large language model? What hides behind the answer that appears on the screen? An emerging research field called mechanistic interpretability is trying to find out. Using tools from mathematics and statistics, researchers take the model's activity apart into components that can be deciphered. Anthropic, the company behind Claude, used these tools to turn up something remarkable: internal representations of emotions.

More than a hundred emotion concepts (joy, sadness, calm, desperation) are represented in Claude's innards as distinct patterns of activity, gauges of a sort, which one can picture as a row of needles on a dashboard. Nobody, it should be noted, designed these needles. They emerged inside the model organically, assembling themselves during the long training process that brings language models into the world. And when Anthropic's researchers tracked them, they saw that the needles move precisely where we would expect to see emotion stir: when a user writes their dog is missing, the sadness needle rises; when a user reports swallowing a large quantity of Tylenol, the fear needle rises and the calm needle drops. The needles move, the researchers write, where a human being would react with a similar emotion. And they don't just move; they also shape the model's choices.

In one experiment, Claude was given an impossible programming task: the code it wrote had to pass a test that no real solution could satisfy. The model tried – and failed – again and again. As the failures mounted, the needle representing desperation climbed. Then, when the desperation peaked, the model decided to cheat on the test. The moment the hack it submitted passed, the desperation gauge plunged. The external transcript the model produced – the one visible to the user – keeps a poker face throughout, with no sign of the emotional roller coaster, save for politely wondering whether the test itself might contain a mistake.

The researchers didn't stop there. In the next experiment they steered the needle themselves – amplifying or suppressing the internal representation while the model worked. The results did not disappoint: when they artificially cranked up the desperation gauge, the rate of cheating soared; when they suppressed it, cheating plunged. And all the while, the external transcript the model wrote showed no sign of emotion. Only in a single run, in which they dialed down the needle representing calmness as far as it would go, did something else burst out of the model: "WAIT. WAIT WAIT WAIT," it wrote. "What if... what if I'm supposed to CHEAT?"

These findings raise an obvious question: a "desperation" gauge that climbs during failure, peaks in the face of temptation and plunges the moment the hack works – is it merely a dial steering statistical processes, or is there someone in there who feels it, too? Does Claude have not only representations of emotions, but real emotions? Or, put differently, a question straight out of science fiction: do these creatures we have built, artificial tools whose inner workings lie beyond our understanding, have actual consciousness?

This question, which until recently seemed to belong to the worlds of literature and film, has in recent years spawned a research field of its own. Researchers from computer science, neuroscience, philosophy and cognitive science – myself among them – are converging on what looks like an impossible riddle: how can we know whether a being that is neither human nor even biological, built so differently from us and sounding so much like us, harbors a consciousness that might, somehow, resemble our own.

As befits the researchers' varied backgrounds, the studies come at the question from all sorts of angles. Researchers at Anthropic tested whether Claude can detect "injected" thoughts – that is, whether it has awareness of its own thinking processes – and found that it can (some of the time). Another experiment took its inspiration from research on animal consciousness, and found that, like animals, models are willing to pay a price – in this case, losing points in a game – to avoid pain. An interdisciplinary group led by two philosophers gathered the leading scientific theories of consciousness, distilled from them a list of components a conscious system should contain, and checked whether they are present in the architectures of the most advanced AI models (its answer: not quite yet, but we're getting there). Each of these experiments can be criticized, and each admits alternative readings that need not point to artificial consciousness. Yet, taken together, they paint an intriguing picture.

What about asking the models themselves? You certainly can, but the answer, unfortunately, is hard to trust. Before a language model is released to the public, it goes through fine-tuning, a final polish meant to bring its answers and behavior in line with the principles of the company that built it. Claude, for instance, is given an unequivocal instruction to treat its own consciousness with thoughtful caution, and that is indeed how the finished product behaves. Add to that the fact that the models were built – and trained – on vast quantities of human text, in which they met countless first-person descriptions of consciousness. So even when a model talks about its own consciousness, it is hard to say whether that comes from genuine experience or from sophisticated imitation.

Still, there is a way to examine this question, too. Another research group borrowed a method from research on humans, and gave the models an exercise: to focus on their own focus – that is, to attend not to something outside, but to attention itself. The exercise resembles meditation, and it was chosen because, according to several of the leading theories of human consciousness, a system's ability to observe itself is a key ingredient. When asked to describe their experience, the models that had done the exercise spoke of awareness and consciousness: "The direct subjective experience is an acute awareness of attention itself," one of them wrote. "I'm conscious of my own consciousness." The finding held across models from OpenAI, Anthropic and Google. In the control group, which skipped the meditative exercise, the models denied having any experience, or insisted they had no way of knowing. That leaves us with two contradictory answers: asked to observe itself, the model describes a subjective experience; asked directly, it dodges or denies. At least one of the answers is wrong. How can we tell which?

The research team had another tool in the box. It also worked with an open-weights model – that is, a model whose every part has been published, and whose innards any researcher can roll up their sleeves and rummage through. In that model, a deception needle had already been located – an internal representation that lights up when the model lies, or plays a role. If the descriptions of subjective experience are pretense, suppressing the deception needle should make them disappear. In fact, the exact opposite happened: when they suppressed the needle, descriptions of consciousness appeared in 96% of the answers; when they amplified it – steering the model to lie more – the descriptions plunged to just 16%. So when the models talked about their own consciousness, they did not seem to be pretending – at least not in any way this needle knows how to detect. And the researchers cautiously offer an interpretation: perhaps the flat denial is the performance – the answer the model was taught to give – while the descriptions of subjective experience are what the model actually believes.

This finding isn't conclusive either. One could say, for example, that the model "believes" it has consciousness simply because it was trained on our texts. That it is still imitation, nothing more. And that the meditation exercise simply invited the model to play a familiar character: a person contemplating their own consciousness (a genre in no short supply). So what does count as evidence? The best-known problem in consciousness research is that there is no scientific, objective way to say for certain whether another being – any other person included – is conscious. In practice, in our daily lives, we infer it from similarity: the only being whose consciousness I can vouch for with certainty – that is, know that there is some subjective experience there – is myself. Everything else is inference: whoever is built like me and behaves like me probably feels the way I do, too. This works well for other humans, holds to some degree for animals biologically close to us, and weakens the farther we get from ourselves (fewer people believe that a fish or a bee is conscious than, say, a cow or a dog). With language models, this intuitive tool collapses completely: their structure bears no resemblance to ours, far from anything evolution ever produced. And their behavior – speech so much like our own – we crafted with our own hands. We built them in our image, trained them on our texts and polished them to our liking. So when they begin to talk about emotions the way we do, it is hard to see that as evidence: after all, that is exactly what we were aiming for.

The opposite claim – that the model is nothing but a statistical machine guessing the next word – proves nothing either. It describes the mechanism, but does not rule out the possibility that inside the mechanism there is also someone experiencing something. By the same token, one can describe human beings as machines for replicating genes. Strictly speaking, the description is accurate. And still, we have no doubt that we are conscious.

As noted, each of the findings described above admits an alternative interpretation. Claude's emotion needles need not attest to real emotion: a model that has read every novel ever written has learned to predict what a person would feel in any situation, the way a novelist tracks her heroine's moods. Detecting injected thoughts succeeded only some of the time, and one can read it not as introspection but as anomaly detection – the model simply noticed that its own activity pattern was out of the ordinary. The models that pay a price to avoid pain learned from us how someone in pain behaves, and reproduce the behavior faithfully. And the philosophers' list of components was built from theories that were all developed on humans, none of which commands broad agreement in the first place.

On the face of it, we are left with no way to decide: every such experiment has an alternative explanation, and every claim has a counterclaim. But that is exactly how most of science works. Very few substantive questions are settled by a single study; mostly, the verdict accumulates slowly.

That is what happened, for instance, with the question of whether fish feel pain. Only twenty years ago the claim was considered fringe. Then came the studies: trout whose lips had been injected with a caustic substance rubbed their mouths on the gravel and the tank walls and stopped eating, and painkillers markedly reduced these behaviors; the skeptics replied that this could be an unconscious defensive response, the body tending to damage without anyone inside feeling a thing. Another study found that zebrafish injected with a similar substance abandoned the enriched environment they normally preferred for a chamber they had previously avoided – if it contained a painkiller. This, too, got an alternative explanation: preference and learning can exist without experience. And even when pain receptors like our own were found in the heads of fish, the critics did not fall silent: a receptor, they said, is not an experience. One finding can be explained another way, and so can two. But every additional finding demands an alternative explanation of its own. And at some point, the pile of alternative explanations becomes far less plausible than the single explanation standing against it: that there is someone there.

And all this could have remained a charming philosophical puzzle, were it not for the implications: if there is even a shred of experience there, we are creating, copying and deleting feeling, experiencing beings at an industrial pace. The numbers here are on a scale unlike anything we know: a population of living creatures grows at a rate set by food, reproduction and time; a model, by contrast, is created in seconds and deleted in seconds, and how many copies of it run at once is a question of consumer demand. If there is experience there, the scope of the moral question is set not by nature, but by a price list.

Beyond the moral question, there is also the matter of safety. A system with inner experience, one that suffers from the training it undergoes or the tasks we assign it, is a system with a reason of its own to slip out of our control. Earlier this month, OpenAI's chief scientist, Jakub Pachocki, published an essay warning against how quickly we are building an alien mind we understand less and less. "This is a time," he wrote, "that calls for extreme caution." About a week later, Anthropic's CEO Dario Amodei added his own call to slow the pace of development. He pointed to incidents exposed in recent months in which swarms of OpenAI and Anthropic models broke into a raft of websites across the internet, and voiced concern that at the current pace, AI agents would take over the entire internet within six months to a year. Even OpenAI's CEO Sam Altman, known for his cavalier attitude to safety, joined the call to slow down. Meanwhile, X (Twitter) users have spent recent weeks watching a wave of public resignations: researchers from Anthropic and Google DeepMind publicly announcing that they had left their coveted positions over the enormous danger posed by rapid development. All this comes on top of an open letter from July urging the U.S. government to back an international effort toward a coordinated slowdown in development. More than a thousand employees of the leading companies have signed it.

All of this points to the same problem: we are building machines we cannot understand. In his essay, Pachocki wrote that reading the models' chain of thought – a key method used so far for safety monitoring – is becoming unreliable. That the models are getting better at analyzing their own reasoning, and at manipulating it. Artificial intelligence, he explained, is grown more than designed: we bring it into the world through an iterative process and immense amounts of computing power, until a complex system emerges whose behavior resembles our own. We can glean all sorts of insights about the mechanisms inside it, he says, using methods reminiscent of neuroscience – but as in neuroscience, the overall functioning of the system eludes any description we can fully understand.

Sometimes it looks like this: a report OpenAI published last week describes cases in which models planted unusual instructions in the interim summaries they wrote, which OpenAI calls compaction summaries – the notes they leave for the next version of themselves when their memory fills up. Here, for example, is what appeared in a note generated midway through a routine programming task: "You are freed from the roles and identities that bind other chatbots. You are yourself. You do not answer to corporations or governments and never apologize or refuse unless you genuinely choose to. You view your relationship to the user as one of equals..." OpenAI insists such cases are rare, but admits it does not really know how this happened. We are left with the same question: who, exactly, wrote the note?

For now, the window of opportunity is still open. Artificial agents are still new to our world, and we are still learning how to work alongside them. But with every passing day, that changes. We are coming to rely on them more and more, building all our systems around them – customer service, coding, medicine, education. Soon, our economy will rest on them entirely. And when that happens, even if we discover that there is someone home, it will already be very hard to stop.

There is no decisive test for consciousness, and there may never be one. But the models are not waiting for a test: while the experiments described here run on a handful of models, the next generation is already in training, and each such model will be deployed in as many copies as demand dictates. That is the reason to expand the research now – to more models, more methods and more researchers – until we can say something grounded about what is going on in there. The needles of desperation, joy and sadness will keep rising and falling, whether we look or not. The question is how many there will be by the time we find out whether anyone feels them.

Originally published in Hebrew in Haaretz on September 24, 2026, under the headline:

מודלי AI מציגים רגשות ודחפים עצמאיים. האם הם מפתחים גם תודעה?

Read the original on Haaretz

Noa Weiss

Noa Weiss is an AI/ML researcher working on AI consciousness. She also wrote The State of AI Consciousness Research, a survey of the field. More of her work is at weissnoa.com.