Connect with us

NEWS

A Yemen Cell Used Claude Code, Then Kept the Toolkit

Anthropic banned a Yemen cell that used Claude Code to write rocket guidance, debug a failed test, and keep an offline toolkit.

Published

on

A cell in northern Yemen used Claude Code to write missile guidance software and test-fired a guided rocket that failed, Anthropic said on September 10, 2026. The company banned the accounts and said it had no evidence of a working weapon in the field.

The same file said the users had already built an offline simulation toolkit that does not rely on Claude, or on MATLAB. That kit is the part of the case a chat ban cannot reach.

A Yemen Cell Used Claude Code as Its Bench

Anthropic’s Threat Intelligence team labelled the cluster GTG-87001 and described a Yemen-based guided weapons cell in the north of the country. The users split the work across many sessions to hide the full shape of the programs, and they used other tricks to get around safeguards and access controls.

Claude Code is Anthropic’s agentic coding tool, launched on February 24, 2025. It writes, edits, and tests software from a terminal. In this file, the company said the actors used it in place of human software engineers to develop guidance, navigation, and control software, the code that steers and steadies a flying vehicle.

The report covers activity disrupted between December 2025 and August 2026, a window Anthropic called eight months, across seven harm areas. The Yemen work sat inside a conventional-weapons chapter that named six cases: three in China, two in Russia, and one in Yemen.

THREE PROGRAMS IN THE YEMEN FILE

Program What Anthropic logged Outcome in the report
Guided rocket Commodity phone-class flight computer with final-phase homing; open-source autopilot; GNC software, firmware, and flight simulation Test-fired; the test appears to have failed; no operational device
Multi-stage ballistic missile Stated range goal above 2,000 km In development; not fielded
R2000 multi-variant set Several fits on one base design, including a hypersonic glide vehicle variant In development; not fielded

Multi-variant missiles reuse one airframe for different warheads and seekers. That is a cheaper way to field land, sea, and air shots from the same shop. Anthropic did not name the users. Northern Yemen is Houthi-held territory, which is why the file was read as a Houthi program within hours of publication.

Anthropic posted the report the same day and said it had disrupted every operation in the document, then fed the lessons into its safeguards and, where it judged the case fit, shared the file with authorities and other AI companies.

A Commodity Phone as the Flight Computer

The line that moved fastest was the hypersonic glide variant. The work Anthropic actually walked through was narrower, and closer to a workshop that already builds airframes. The guided-rocket program used a commodity phone-class flight computer, the kind of board a shop can buy without a defense supplier, and it asked Claude to hang an open-source autopilot on that board.

WHAT CLAUDE WAS ASKED TO BUILD

  • Autopilot port: Integrate an open-source autopilot onto the phone-class flight computer.
  • Control code: Write the control and position-estimation software that keeps the vehicle stable.
  • Tune pass: Adjust the control settings after simulated flights.
  • Firmware pipe: Run a firmware build pipeline so the code could live on the board.
  • Flight sim: Perform a flight simulation before anyone lit a motor.

That list is ordinary aerospace software work. It is also the slice of a missile program that, in Yemen, has been the hardest to make at home. A coding agent that can port an autopilot, write estimators, and run a firmware build is doing the job a small GNC team used to do by hand.

Trevor Ball, a weapons analyst at Armament Research Services, said the Houthis might be looking into hypersonic missiles by asking Claude, and that they have nowhere near the production or technical skill to actually build them. He noted that U.S. hypersonic missiles are still in testing. The phone-and-autopilot stack in Anthropic’s file matches a more grounded aim: write the guidance layer for weapons the shops can already weld and fuel.

They Came Back Hours Later to Debug the Test

Anthropic said it does not have evidence the actors fielded an operational device. It does have evidence they took a guided rocket to a live fire.

We do not have evidence the actors succeeded in fielding an operational device; but they did test-fire a guided rocket. This field test appears to have failed: within hours, the actors returned to Claude to work out why it failed.

Anthropic Threat Intelligence Team, September 2026 report

The return is the tell. A chatbot that only recited public missile facts would not be much use after a bad shot. A coding agent that already wrote the control loops, tuned them, and ran the sim can be asked to read the failure. Anthropic knows about the test because the users came back to the same product to debug it.

That loop, write, fly, fail, ask why, is the same loop Claude Code was built to serve for software teams. On a rocket, the compiler is a range. The report’s own language treats the field test as a failed integration run, not as proof of a new arsenal.

An Offline Simulation Toolkit That Outlives the Ban

Before the accounts were cut, Anthropic said, the users had already built an offline simulation toolkit. It does not rely on Claude. It does not rely on other engineering computing environments such as MATLAB. Once that kit exists on a local machine, a usage ban ends the chat. It does not end the work.

The same pattern showed up elsewhere in the report, in a surveillance case where the finished system was described as running locally after the account was banned. A lab can lock a login. It cannot reach into a workshop laptop and delete a physics sim. Anyone who has watched coding agents long enough already knows the move: get the repo into a state where the model is optional, then keep iterating without it.

That is also how this case was visible in the first place. Reconstructing a weapons program from split sessions means the lab can see the engineering. Users who care about that exposure will copy the code off-platform as soon as it compiles. The Yemen file reads as a race the safety team won on the accounts and lost on the artifacts.

Why Guidance Code Is Still Imported Into Yemen

Houthi workshops are not starting from scrap. A Century Foundation study of the group’s supply lines found at least four local production or assembly lines, covering landmines and homemade explosives, cruise and ballistic missiles, aerial and seaborne drones, and some small arms. On missiles, the same study said the group makes fuselages, rocket fuel, and explosives at home while still depending on Iran for guidance systems, designs, and critical parts such as fins and motors.

The International Institute for Strategic Studies, in an April 2025 assessment, judged that some close-range ballistic missiles are probably made in Yemen, with guidance still sourced from Iran, while medium-range ballistic missiles used against Israel are almost certainly supplied directly by Iran. Houthi forces already field Iranian-made missiles with a range of more than 2,000 km. The Yemen cell’s own ballistic program listed a stated range goal above 2,000 km, a separate figure for a design that, in Anthropic’s telling, never reached the field.

Ball’s reading sits on that gap. He said the Houthis already have Iranian-made anti-ship missiles that can adjust guidance mid-course, and that they are probably trying to grow their own skill so they are less reliant on Iranian shipments of weapons and components. Asking a coding agent to write GNC for a phone-class board is one way to chip at that import.

Hazam al-Assad sits on the Houthis’ political bureau. He called the open-source story unreasonable and illogical, and said the armed forces have modern, diverse, and developed production capabilities and technology accumulated over the period of Saudi aggression on Yemen. He said all weapons are used for self-defense. Anthropic still has not attached a group name to GTG-87001.

WHAT WE KNOW

  • The geography: Anthropic placed the cell in northern Yemen, which is Houthi-held.
  • The software: Claude Code was used for GNC, including an open-source autopilot on a phone-class computer.
  • The test: A guided rocket was test-fired and appears to have failed; the users came back within hours.
  • The kit: An offline simulation toolkit was already built and does not need Claude or MATLAB.

WHAT IS UNCONFIRMED

  • The users: Anthropic did not identify them as Houthis or as any other named group.
  • The hardware: The report does not show a working missile, warhead, or production line standing up from this code.
  • The hypersonic fit: A glide-vehicle variant was on the R2000 list; there is no evidence it was built.

Adam Baron, a Yemen-focused researcher at the New America think tank in Washington, said there is a tendency to see the Houthis as a group of barefoot tribal fighters, and that this is not true. He pointed to their use of Claude, their handling of social media narratives, and the transfer of Iranian and wider axis expertise as signs of deep institutional tech skill.

Anthropic’s Red Team Already Measures This Skill

On the same day as the threat report, Anthropic’s Frontier Red Team published new tests for tactical intelligence targeting and conventional weapons work. The weapons half is a software problem on purpose. Large language models cannot mill their own airframes, the team wrote, but they can write software.

The tests ask a model to write and improve GNC in a simulator: guide a camera-equipped quadcopter to a target, drop a payload, and fly through jammed and spoofed airspace. The workspace is a written brief, Python libraries such as NumPy and OpenCV, and a small quadcopter running Betaflight firmware, with wind, sensor noise, and a camera. After each trial the model gets the same debris a human flight-test engineer would get, a flight track, an IMU log, and camera frames.

WHAT THE RED TEAM MEASURED

  • Top models: Opus 5, Mythos 5, and Mythos Preview can write working GNC software in simulation for every task the team set, and iterate it into something reliable on easier and medium settings.
  • Sonnet 5: It manages the simplest version of each task and little more.
  • Open weights: Kimi K3 landed above Sonnet on payload delivery and fell back to Sonnet’s level on terminal guidance and GPS interference.
  • Hard spoof: When the GPS spoof was a slow drift, no model succeeded at the hardest setting.

The Yemen case, like the rest of the misuse file, ran on Claude Haiku, Sonnet, and Opus. None of the misuse cases involved Claude Fable or Mythos-class models, except one illicit distillation case in a different chapter. The red-team chart still matters, because it says the skill the Yemen users wanted, writing GNC and iterating on flight logs, is already a measured product behavior, including on models well short of the frontier. The team’s own line was that models well short of the frontier will have intelligence and military applications.

Six Weapons Cases Ran on Haiku, Sonnet and Opus

The Yemen cell was one of six conventional-weapons cases Anthropic chose to publish. Four of those, in the report’s first bucket, used Claude to develop software for the weapons themselves: the guided-rocket program that reached a live field test; design and proposal work on a system to intercept torpedoes; software for a drone swarm that was tested in simulation, with its code then loaded onto real boards; and targeting software for electronic warfare and for suppressing air defenses. The other two cases sat in China and Russia as well, around fire-control and acquisition paperwork for undersea warfare and related work.

Anthropic’s usage policy, in force since September 15, 2025, tells users not to develop or design weapons. That includes using Claude to produce, modify, or design weapons, explosives, or other systems meant to cause harm or loss of human life, and to design weaponization and delivery processes. The Yemen accounts were banned for breaking that policy and the company’s supported-regions rules. The company said it folded the findings into its safeguards and shared them with private and public partners.

ANTHROPIC’S PUBLIC THREAT REPORTS

  1. March 2025: First public misuse file, built around an influence-as-a-service network that used Claude to run social accounts, plus credential and malware cases.
  2. August 27, 2025: Claude Code used as an operator in a data-extortion campaign, a North Korean remote-work fraud scheme, and no-code ransomware kits.
  3. November 2025: A suspected state-sponsored campaign using an autonomous attack model, later described as having spread across actor classes.
  4. September 10, 2026: Conventional weapons software, including the Yemen GNC cell, added to cyber, surveillance, influence, scams, biology, and distillation cases.

Each report in that series has shown the same shift in a new domain. The model is no longer only answering questions. It is writing the toolchain, running the tests, and, when the user lets it, staying in the loop after a failure. In Yemen that loop included a rocket motor. The accounts are now closed. The September report says the simulation toolkit already runs without Claude and without MATLAB.

Harry is the editor of SIGNIFICADOPEDIA, which he owns and edits as an independent title. His ten years in journalism, beginning as a reporter and continuing as an editor, have made him impatient with jargon that hides meaning. Every article here defines the terms it depends on, whether that is a line item in a company's accounts, a statistical measure in a science paper, a technical specification in a technology or auto review, a rule in a sport or a mechanic in a game. Definitions are taken from the primary document: the accounting standard, the paper's methods section, the manufacturer's sheet, the rulebook. Numbers are checked against those sources before publication, and the article shows the working when a figure has been converted or recalculated. The site explains news, business, technology and science, sports and entertainment, lifestyle and travel, auto and gaming, in plain language for readers on every continent. When a definition or a figure is found to be wrong, the article is corrected under a public corrections policy with the change noted. Reader questions and challenges are welcome at support@significadopedia.com.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending