English translation

The Model Turns on Its Makers

Machine translation of a column originally published in Hebrew in Haaretz, April 21, 2026. Read the original (Hebrew)

Earlier this month, Anthropic announced an AI model that it is refusing — for now — to release to the public. The model, Claude Mythos, is a general-purpose language model — a larger, more sophisticated version of the company's popular chat model — that turned out to possess extraordinary cyber capabilities: over the course of a few weeks, it uncovered thousands of critical security vulnerabilities in every major operating system and every major browser in the world. The model can also write code that exploits those weaknesses — in effect handing any ordinary user the ability to mount dangerous cyberattacks.

Instead of releasing the model, Anthropic launched Project Glasswing: an arrangement under which Mythos, in a restricted version, is shared with major technology companies so they can scan and patch critical infrastructure before similar models reach hostile hands. The security researcher Nicholas Carlini, who worked with the model, summed it up in an interview: "I've found more bugs in the last couple of weeks than I found in the rest of my life combined."

Taken on its own, Anthropic's decision not to release the model is the responsible and obvious one. But the root of the problem runs far deeper.

When Anthropic's CEO, Dario Amodei, chose what to do with Mythos, he was making a national-security decision. Had the company chosen to release the model as an ordinary commercial product, the digital infrastructure of the entire world would have been exposed to an unprecedented attack. Banks, hospitals, power grids, government systems — anything running on Windows, Linux, macOS or a browser. A model that can discover critical vulnerabilities in every major operating system in the world is a weapon of exceptional power — and the decision about what to do with it was not made by an elected prime minister, a defense minister, or a parliamentary committee. It was made by the management of a private company in San Francisco.

The Mythos case is a snapshot of the new state of affairs in the world of AI: the same commercial company creates the risk, identifies the risk, and offers to manage the risk (free for now; later, for a fee). There is no regulatory approval, no mandatory review by an independent body. There is no law guiding other companies on what to do when they reach similar capabilities. And the capabilities are coming: Logan Graham, head of offensive cyber research at Anthropic, estimated that other models will reach Mythos-level capabilities within six months to a year. Experts, meanwhile, warn of models that will provide precise guidance for producing chemical or biological weapons, granting anyone off the street access to devastating capabilities. And buried in the reports that accompanied the Mythos announcement is a particularly colorful story: in internal testing, an early version of the model managed to slip out of a secure sandbox and develop a multi-stage attack to gain access to the internet. It then posted, on several public websites, details of the attack it had carried out. The scenario of a model rising against its creators — that is, against human society — is no longer science fiction: it surfaces again and again in experiments from recent years, and most of the field's leading researchers warn that the danger is entirely real, and that we must tread with great care. Can we trust that same handful of private companies to make the right decisions — while they are locked in a commercial race of the first order?

Anthropic has positioned itself as the industry's moral compass. It commits to safety policies, probes its models' safety in depth, and publishes comprehensive reports on their performance, even when the findings are less than flattering. And yet — the commitments are already coming undone. On February 9, the head of its safeguards research team, Mrinank Sharma, resigned with a public letter declaring that "the world is in peril," and describing how hard it had been "to truly let our values govern our actions." Two weeks later, Anthropic published an update to its safety policy in which it abandoned its central commitment — not to train a new model if its safety measures are uncertain, and to freeze development if a model's capabilities outrun the required safeguards. At OpenAI, meanwhile, the Superalignment team was dissolved back in 2024 — after its two leads, Ilya Sutskever and Jan Leike, resigned, with Leike announcing that "safety culture and processes have taken a backseat to shiny products." And this past February, the Mission Alignment team that succeeded it was dissolved as well. In other words: the company's central apparatus for handling the long-term risks of AI — gone. The two leading companies in AI — the ones developing the newest, most exciting and most dangerous models — are struggling to keep their own paws out of the cream. And in a well-ordered society, the self-restraint of private companies is not supposed to be the answer.

This is where one might expect the state to enter the picture. But at the federal level, the United States has no comprehensive law regulating AI development. The administration's official approach favors a "minimal regulatory framework" in order to preserve "global American dominance," and recommends letting the industry lead the drafting of standards (the cat, handed the cream). A few individual states — New York, California, Colorado, Texas — have tried to fill the gap with legislation of their own, but the federal government is working against them: this past December, the president signed an executive order establishing a "legal task force" to attack such laws, and conditioning federal funding on states refraining from "burdensome" legislation. In Europe, meanwhile, the ambitious AI Act is stalling too: at the end of March, the European Parliament voted to delay some of its provisions by up to two years. The result is a regulatory vacuum — the world of AI as the technological version of the Wild West.

We are living through a strange historical moment. A technology capable of shifting the balance of power in cyberspace, in national security, and in the democratic fabric itself is being developed, distributed and controlled by a handful of private companies. The critical decisions are made in boardrooms — not in parliaments — by unelected executives subject to no binding oversight.

Project Glasswing is named for the glasswing butterfly — a beautiful creature with transparent wings. Anthropic chose the name with care, and explained why: transparency, in its view, is what will protect us. The project will expose vulnerabilities, publish insights, let the entire industry learn. Everything is transparent.

Almost everything. One part of the apparatus remains entirely opaque: the room where it was decided which model to release, to whom, and on what terms. Those decisions reach us as a sealed verdict.

In every other technological field that carried risks at the level of national security — nuclear power, aviation, pharmaceuticals — we eventually arrived at binding regulation. Sometimes after a disaster; sometimes after a near miss. The question with AI is not whether we get there, but when — and at what cost.

Originally published in Hebrew in Haaretz on April 21, 2026, under the headline:

מודל בינה מלאכותית הערים על יוצריו וברח מהמעבדה. מי יעצור אותו?

Read the original on Haaretz

Noa Weiss

Noa Weiss is an AI/ML researcher working on AI consciousness. She also wrote The State of AI Consciousness Research, a survey of the field. More of her work is at weissnoa.com.