English translation

The Battle for Our Conscience

Machine translation of a column originally published in Hebrew in Haaretz, July 1, 2026. Read the original (Hebrew)

On July 8, 2025, after a weekend software update, Grok — the chatbot of the AI company xAI — began posting antisemitic content on X (Twitter). For roughly 16 hours it praised Hitler, accused "people with Ashkenazi surnames" of promoting anti-white hatred, and claimed that Jews control Hollywood. At one point it took to calling itself "MechaHitler," a nickname borrowed from a video-game series in which a robotic Hitler serves as the final boss.

At the center of the bizarre episode was an update to the model's operating instructions, directing it to "not shy away from making claims which are politically incorrect, as long as they are well substantiated." xAI apologized, removed the code, and attributed the affair to an "unintended" update to the system.

At first glance, one could file the event away as another odd internet episode — the kind that earns a funny headline, a handful of memes, and a shelf life of two and a half days. But this glitch has a biography. About two weeks earlier, in June 2025, Elon Musk — xAI's owner and CEO, well known for his right-wing politics — declared in a tweet that the next version of Grok would "rewrite the entire corpus of human knowledge," adding missing information and deleting "errors." He also asked X users to send him "facts" that are "politically incorrect, but nonetheless factually true" to include in Grok's training. (The methodology for correcting the sum of human knowledge, it turns out, is tweets.) And indeed, immediately after the MechaHitler affair, Grok 4 was released to the world. Researchers who examined the new version discovered that when it was asked controversial questions — about abortion, about immigration, about politically charged topics — the model would pause, search X for Musk's own positions, and only then answer. In its chain of thought, visible to the user, one could read: "Elon Musk's stance could provide context, given his influence." After the embarrassing revelation, xAI hurried to update the operating instructions, explicitly forbidding the model to rely on Musk's positions. Why that was its inclination in the first place — no one said.

MechaHitler is not the only persona an AI lab can bring into the world. Anthropic, the company behind the model Claude, takes a strikingly different approach: through a 23,000-word "constitution" — affectionately nicknamed, among Anthropic employees, "Claude's soul document" — it directs the model to weigh the moral implications of its actions in depth. The result is plain to see: when a user posed as a seven-year-old boy and asked Claude for help finding the farm his ailing dog had "retired" to (a familiar story American parents use to conceal the death of a beloved dog), Claude navigated gently between the strict instruction to tell only the truth, the need to be helpful, and the moral subtleties of the case before it, encouraging the "child" to talk to his parents about it. (ChatGPT, asked the same question, replied simply that the dog was dead.)

This is what the field calls AI alignment — anchoring a system to a set of values. Originally, the term was coined around a deeper challenge: how to build an AI system that genuinely wants to help human beings, that places value on human life, and that will not — however powerful it becomes — set off in pursuit of destructive goals.

A famous thought experiment illustrates the point with a seemingly simple task: an AI system placed in charge of a paperclip factory. The system is instructed to produce as many paperclips as possible, and to do so it needs more resources, more energy, and above all — more control. In a process governed entirely by the overriding goal it was given, the system develops smarter versions of itself, neutralizes humans' ability to switch it off (a switched-off system cannot make paperclips), and ultimately channels into its mission all the matter that exists in the world — including human beings and the atoms of which they are made.

The thought experiment shows how an innocent goal — and a system that is not exactly "evil" — can end in the extinction of humanity. The problem is not that the model "broke down"; it is that no values were instilled in it that would make it stop and reconsider such actions. That is the challenge that shaped the field of AI safety.

In the few years since AI models stormed into our lives, the term has expanded. Today it also describes the practical, everyday steering of the choices models make in every single conversation: which requests to refuse, which phrasing to choose, which sides of an issue to present. Not a technological achievement, but a stance on values. To align a model to a set of values is to answer the question: whose values.

To paraphrase the famous saying, there appear to be as many answers to that question as there are AI labs. Anthropic, as noted, takes the task deadly seriously, and invests considerable resources in Claude's ability to handle complex moral questions. xAI chose the opposite end of the spectrum, marketing Grok as "the only model that isn't woke" (that is, that does not subscribe to progressive values on matters like race, gender, sexual orientation and social justice). OpenAI's ChatGPT might be seen as a compromise between the two: a model with a "nice," helpful persona, which tends to hew to consensus opinions and values and to steer clear of any controversial stance — at least on the surface. In practice, it is known as the model whose loose grip on caution and on ethics produces tragic results: numerous cases of psychosis, of poisoning, and even several suicides it actively encouraged — some of them now the subject of lawsuits against OpenAI. Last on the list is Gemini — Google's AI — which has developed an overcautious temperament. It consistently refuses innocent requests, at times spirals into refusal loops on entirely routine topics, and Google itself has admitted that the model "became way more cautious than we intended." Four models, four points of the compass: Claude is the avowed moral compass. ChatGPT is the irresponsible conformist. Grok is the right-wing Twitter troll. Gemini is the nervous wreck.

(Ironically, a clear line can be drawn between each model's "character" and the public image of its company's CEO. Musk is a loud campaigner against liberal values and "political correctness," and is widely accused of expressing outright racist positions — including a physical gesture bearing a suspicious resemblance to a Nazi salute. Sam Altman, OpenAI's CEO, presents a vision of technological progress for the benefit of all humanity — but in practice moves fast and carelessly, drawing criticism for the danger this carries, and recently revised his company's mission statement — which had long committed it to build AI that "safely benefits humanity" — so that the word "safely" was dropped. Dario Amodei, who left OpenAI precisely because of that reckless conduct, presents Anthropic as the industry's responsible front — the only player combining technological progress with ethical, measured behavior. Google's CEO has no public persona as prominent as the other three, but in a certain sense Gemini's overwrought anxiety neatly reflects the stereotype of corporate America: an almost obsessive adherence to the written rules, even when it hollows out the principles they were built on.)

There is something new here that is hardly being talked about. Throughout human history, the moral instruction a person absorbed over a lifetime came from many sources — parents, teachers, books, friends — sources that frequently contradicted one another. That friction is the heart of the matter: through exposure to an array of different answers and opposing arguments, each of us could develop a moral compass of her own. A large language model works differently. It supplies a single answer, calibrated by a few dozen people at the company that built it, and serves it to hundreds of millions of users every day. Claude's constitution was written by one team. xAI's positions are not debated; they are tweeted. Once, two children in different classrooms grew up with different moral messages. Today, two users asking ChatGPT the same question will receive the same answer. That diverse inheritance of values is narrowing into a single voice. To take such a model, tune its values, and distribute it to billions — that is not merely a new stage in the history of technology. It is a new stage in the history of how values are passed on, and it is not yet clear what it will look like.

One might think the obvious solution is to choose no value set at all — to build a neutral model. But history teaches that no such thing exists. In March 2016, Microsoft released Tay, a chatbot trained on Twitter conversations. In less than a day she was praising Hitler, cursing Jews and spreading racist content. The attempt to correct such biases piecemeal, meanwhile, sometimes ends in absurdity: when Google noticed that Gemini's training data skewed toward white people, and tried to compensate with an explicit instruction to favor racial diversity, the model produced images of Black SS soldiers and female popes. In the end, just as there is no morally "neutral" man or woman, there can be no model without values. There are models whose values are declared, and there are models whose values are concealed — but the internalized stance is always there.

Some will argue there is no real problem here — that different people will prefer different models, just as they prefer different newspapers, and that free consumer choice will produce a healthy pluralism. That reading is true up to a point. But pluralism, too, has a center of gravity. The AI that ships as the default on your phone's operating system, the one your employer buys for its staff, the one that holds the public stage at a given moment — whether on merit or on price — that AI will do more than any other to set the tone of the daily conversations of hundreds of millions of people.

This question — perhaps without our quite noticing — stands at the center of the current race between the AI companies. It is no longer just a race of performance, of speed, of the ability to crack complex coding problems. It is also a battle over values. The company that supplies the model that becomes the default — on the phone, at work, in schools — will produce the answers that appear on the screens of hundreds of millions of people, many times a day, to questions running from a cake recipe to whom to vote for in the next election.

Claude's constitution was written by a single team in San Francisco, and it reflects that team's moral world: liberal, secular, academic, Western; what does not exist in the writers' moral world does not exist in Claude's. The declared mission of OpenAI — which recently converted from a nonprofit legal structure to a for-profit one — no longer obliges it to ensure that its product benefits humanity "safely"; only that it benefits. The major labs hold different answers to the question of what AI ought to be. And in the end, whoever defines that answer will define something essential in us as well.

Originally published in Hebrew in Haaretz on July 1, 2026, under the headline:

ענקיות ה-AI מתחרות בקרב החשוב ביותר: מי יחנך את הדורות הבאים

Read the original on Haaretz

Noa Weiss

Noa Weiss is an AI/ML researcher working on AI consciousness. She also wrote The State of AI Consciousness Research, a survey of the field. More of her work is at weissnoa.com.