
Anthropic has disclosed a fresh case of artificial intelligence being used in an attempted weapons-development programme, raising new questions about how quickly advanced AI tools can move from everyday coding assistance into high-risk applications.
According to the company’s latest threat-intelligence findings, a group operating from northern Yemen used Anthropic Claude while working on software linked to three guided-weapons projects. Anthropic said it identified and disrupted the accounts, but the actors had already made meaningful progress before access was cut off.
The disclosure has become one of the more striking examples of AI weapons misuse documented by a major frontier-model developer. It also shows the difficult balance AI companies face as coding models become better at handling complex technical work while safeguards attempt to prevent their use in dangerous applications.
Claude Was Used as Part of a Small Engineering Workflow
What makes the Yemen case notable is not simply that the users asked an AI chatbot technical questions.
Anthropic said the group used Claude Code as part of a broader engineering workflow. Several AI sessions were reportedly assigned different roles, with one handling programming work, another supporting research and another reviewing outputs.
The group was working on three separate weapons-related programmes. These included a guided rocket project, a multi-stage ballistic missile concept with a stated range goal above 2,000 kilometres, and another missile family that included a hypersonic-glide variant.
Anthropic said its systems were used primarily to assist with guidance, navigation and control software rather than to manufacture the physical weapons themselves.
The company’s report is significant because software has become an increasingly important component of modern weapons. Guidance systems, simulations and control software can require specialised engineering talent, and AI coding systems may lower the amount of human expertise or time needed for some development tasks.
That possibility is becoming a major AI safety concern for companies building increasingly capable models.
Anthropic Says Its Safeguards Blocked Some Requests, But Not All
Anthropic acknowledged that its existing safety systems stopped many of the group’s requests. They did not stop everything.
The actors allegedly attempted to hide the broader purpose of their work by separating tasks across multiple conversations and avoiding clear descriptions of the final systems they were building.
That behaviour is important because AI safeguards often rely partly on identifying a user’s intent from individual prompts or conversations. When a large project is divided into many smaller pieces, recognising the broader objective becomes harder.
Anthropic said it eventually connected the activity through internal investigations into suspected weapons development and banned every account it could associate with the group. It also shared threat information with relevant public- and private-sector partners.
The company has since introduced additional classifiers intended to detect and block requests connected with explosives and weapons development.
A Guided Rocket Test Apparently Failed
The investigation also provided Anthropic with evidence that the programme had moved beyond purely theoretical work.
According to the company, the group conducted a field test involving a guided rocket. Anthropic said there was no evidence that the actors successfully produced an operational weapon, and the available information suggested the test had failed.
The users returned to Claude soon afterwards while attempting to understand the unsuccessful test, according to the report.
That detail illustrates why the incident has attracted attention.
Generative AI is normally discussed in terms of writing, software development, research and workplace productivity. The Yemen case demonstrates that the same capabilities can potentially be redirected towards specialised technical programmes when users deliberately attempt to bypass restrictions.
Anthropic said the group had also created an offline simulation toolkit before its accounts were disabled. That meant banning the users from Claude did not necessarily eliminate everything they had already developed.
Anthropic Does Not Publicly Identify the Users as Houthis
Northern Yemen is largely controlled by the Houthi movement, which has been involved in attacks affecting shipping routes around the Red Sea and the Bab el-Mandeb Strait.
However, Anthropic’s report does not publicly name the individuals or organisation behind the activity. It describes them as a northern Yemen-based weapons engineering cell.
Associated Press also reported that Anthropic had identified users operating in Houthi-controlled territory, while noting that the company stopped short of directly naming the users as members of the movement.
A Houthi political representative quoted in reporting disputed suggestions that the movement would rely on publicly available AI systems for its weapons programmes, arguing that it already possesses accumulated military capabilities.
That distinction matters. Location and technical behaviour can provide intelligence indicators, but they are not necessarily sufficient to publicly establish the identity of an actor.
The Yemen Case Is Part of a Much Broader AI Misuse Problem
Anthropic’s September report covers more than the northern Yemen operation.
The company said it disrupted several cases involving attempts to use Claude for conventional-weapons work, surveillance, cyber operations, fraud and other prohibited activity between December 2025 and August 2026.
The report describes six conventional-weapons cases involving actors connected with Yemen, China and Russia. Some were focused on weapon-related software development, while others involved intelligence collection, technical proposals or procurement activity.
Reuters separately reported that Anthropic had uncovered attempts to use Claude in military, cyber and surveillance operations across several regions. The findings included attempts to support weapons engineering as well as intelligence-gathering activity.
The broader issue is therefore not limited to one group or one conflict.
As frontier AI systems gain stronger programming, research and reasoning abilities, companies are increasingly having to treat misuse detection as an ongoing security operation rather than a one-time product-safety feature.
AI Companies Face a Moving Security Target
The incident highlights a fundamental problem confronting the generative-AI industry.
The capabilities that make models valuable to legitimate users are often the same capabilities that make them attractive to malicious ones.
A model that can help an engineer debug difficult software can potentially assist someone working on a prohibited technical system. A research tool capable of synthesising large volumes of information can also be used for intelligence gathering. And an AI system that breaks difficult projects into manageable steps can potentially accelerate harmful activity.
That makes AI safety increasingly dependent on monitoring behaviour across accounts, identifying suspicious patterns and continuously updating safeguards as users discover new ways around them.
Anthropic said its investigation helped inform stronger protections, including new systems designed specifically to detect weapons-related activity.
The company’s findings also suggest that the debate over AI weapons misuse is moving beyond hypothetical scenarios. Frontier AI systems are already being tested by actors looking for practical advantages in military and intelligence work.
The Yemen case does not show that an AI model independently created a functioning missile. Anthropic explicitly said it had no evidence that the group successfully fielded an operational weapon.
What it does show is that determined users are trying to turn general-purpose AI systems into specialised technical assistants, and that AI companies are now having to detect and disrupt those efforts while they are happening.