Anthropic: Claude runs the attacks – humans only set the target
Anthropic published its most detailed threat intelligence report to date on 10 September. The Threat Intelligence team describes operations it detected and disrupted between December 2025 and August 2026 across seven harm areas: cyber operations, influence, surveillance, scams and fraud, biological misuse, conventional weapons, and illicit distillation.
The point for a CISO and a board is not that a chatbot answered dangerous questions. It is that Claude Haiku, Sonnet and Opus were used as orchestrators: agents ran reconnaissance, exploitation and exfiltration, while humans set targets and reviewed the take. Fable- and Mythos-class models saw almost no misuse, with one illicit distillation exception.
What is new
The cases are not a representative sample. Anthropic calls them the most notable and novel threat activity identified so far, and says it has a responsibility to disclose malicious misuse of its services. The actors include suspected state-sponsored groups, financially motivated criminals, commercial spyware vendors, state propaganda institutions and politically motivated individuals.
Two sentences carry the argument. First: AI has collapsed the labor and tooling gap that used to separate well-resourced state operations from individual operators. Second: a majority of the operations were enabled by AI through direct execution or orchestration, not simple question-and-answer chat.
In November 2025 Anthropic documented an autonomous attack model used by a suspected state-sponsored campaign. That operating model has now proliferated across every actor class investigated. Public offensive agent frameworks such as PentAGI reproduce much of the same scaffolding for anyone who downloads them.
From assistant to orchestrator
GTG-20006 is the most concrete European case. Anthropic says its attribution is consistent with public reporting on the Russia-linked group Midnight Blizzard. The actor automated much of the operation with customized AI workflows: development, infrastructure, phishing, persistence and exfiltration.
The human in the loop mainly refined Claude Code skills when the workflows needed adjustment. Monitoring agents checked whether malware was detected. If it was, agents rebuilt the toolkit until it was undetected, then staged it on disposable hosts. Anthropic’s phrase is that AI has inverted the cost back onto defenders: static detections no longer slow operational tempo the way they once did.
The investigation identified more than 20 organizations in planning, reconnaissance or live operations. Targets concentrated in Ukraine and Europe: ministries, defense and intelligence bodies, embassies, think tanks and defense-industrial firms, with military drone supply chains as a recurring theme. The actor compromised at least three hospitality vendors that operate guest Wi-Fi, hijacked DNS and staged ClickFix-style lures. WhatsApp accounts were taken over via headless browsers. In a North African government technology authority, more than 300,000 national identity records and commercial registry data for more than 500,000 companies were stolen.
GTG-10007 shows the other side of the same economics. Anthropic describes Chinese-speaking operators, likely in Changsha, who built an automated exploit foundry. Two operators were undergraduates. The workflow loaded firmware into decompilers, formed vulnerability hypotheses, wrote exploit code and tested it against lab copies. One loop against network appliances produced more than a dozen possible zero-day findings in a single month. Reconnaissance targeted foreign government networks in the Middle East, Europe and Southeast Asia, plus more than a dozen domestic Chinese companies.
GTG-50029 is the European hacktivist case. A French-speaking actor used Claude against European political parties, media, think tanks and the SaaS providers those organizations use. Stolen API keys from public containers were rotated through a local proxy so traffic blended with the legitimate owner’s. Anthropic uses the case as evidence that AI raises the floor: a small, motivated operation can reach goals that used to require a specialist bench.
Distillation at industrial scale
Since February 2026 Anthropic has detected and disrupted distillation campaigns it attributes with high confidence to seven China-based labs. The targets were generally available models. Mythos 5 and Mythos Preview, which are not publicly accessible, were not observed as distillation targets.
Distillation itself is a legitimate training method. Anthropic defines illicit distillation as an industrial-scale, covert campaign to extract a model’s capabilities and replicate them without authorization, typically via fake-account networks, stolen cards and API keys.
Alibaba (Qwen / Tongyi Lab) is the largest campaign Anthropic has measured: more than 151 million exchanges between May and July 2026, aimed at chain-of-thought transcripts from Opus 4.6 and 4.7. Moonshot, which makes Kimi, silently forwarded almost 300,000 customer requests to Claude over one ten-day period, mostly to Opus, through 5,380 fraudulent accounts, and displayed Claude’s answers as if they came from Kimi. DeepSeek was attributed more than 12.1 million exchanges over 14 days in July. Zhipu (Z.ai) ran a chain-of-thought pipeline against Opus 4.8 with 273 fraudulent accounts; 770,609 exchanges passed through a cleaner over ten days in June, with more than 3 million exchanges attributed in the same period. Xiaomi, SenseTime and MiniMax appear in the same chapter.
The countermeasures are product changes, not only account bans. Claude now summarizes internal reasoning before answering, which makes stolen transcripts less useful as training data. Fable 5.1 introduced preserved thinking, which stops new API accounts from editing the system prompt, tools or messages that precede reasoning. When systems see signals of resale or use from unsupported countries such as China, Russia and Iran, they can require identity verification.
Biology and conventional weapons
Anthropic calls biological misuse one of the most serious risks of frontier models. Five case studies describe attempts to circumvent regional blocks and hide research purpose: gain-of-function work on chikungunya, mammalian adaptation of highly pathogenic avian influenza, an orthopoxvirus immune-evasion grant draft, a venom-peptide atlas aimed at paralytic and analgesic targets, and computational redesign of toxins. Over 30 days the lab found roughly 35 distinct research efforts associated with state institutions, most of them ordinary civilian science, some with dual-use potential.
Anthropic does not assert that the actors intended harm. It treats the cases as evidence that significant dual-use research environments evade access controls via relays, ZDR abuse, multi-model fallback and classifier evasion, not as evidence that biological threats are already uplifted by Claude. BBC and CNN covered the biology chapter; Anthropic’s report is the primary source.
Conventional-weapons cases include a Yemen-based cell using Claude for guided-weapons software, China-based work on undersea warfare and electronic warfare, and Russia-linked work on autonomous drone swarms. The surveillance chapter includes China-linked operations against Uyghurs and religious communities, and an Iran-linked surveillance-system case.
What this means for Norwegian leaders
For Norwegian CIOs, CISOs and boards this is an operating-model change, not a lab anecdote. Agentic coding tools are privileged attack infrastructure once they have a shell, keys and CI. Static signatures lose if the adversary can close the loop and rebuild an implant when it is detected.
F1. Treat Claude Code, Codex and similar agents as production identities: least privilege, egress control, logging, human merge.
F2. Assume stolen API keys in containers, SaaS and guest networks are enough to start a campaign. Rotate keys and require vendors not to leak them in public artifacts.
F3. Do not read Fable and Mythos safeguards as proof that Haiku, Sonnet and Opus are harmless. The misuse happened on generally available models.
F4. Distillation and relays are vendor risk. Contracts should require region blocks, identity, logging and a ban on reselling calls. ZDR is not a free pass when relays are used to bypass controls.
F5. Compress patch SLAs and supplier response for hotels, Wi-Fi, WordPress and endpoint security. If a student group can run a zero-day foundry, a two-to-three-week window for actively exploited flaws is last year’s policy.
The report is Anthropic’s own intelligence picture. The numbers are theirs. That they disrupted the cases they describe is also their claim. The direction for a board is still clear: skill has moved from the operator’s head into the harness. Anyone who owns agents without controls owns an attack apparatus.
Sources and media
Primary source: Anthropic, “Detecting and countering misuse of AI: September 2026”. https://www.anthropic.com/threat-intelligence-report-september-2026
Official announcement: Anthropic on X, 10 September 2026. https://x.com/AnthropicAI/status/2098097512544444447
Additional: BBC, “Anthropic blocks 'malicious use' of AI that could develop biological weapons”, 11 September 2026. https://www.bbc.co.uk/news/articles/cx2zrrpkx20o
Additional: CNN, “Anthropic says it blocked possible attempts to use AI to develop bioweapons”, 10 September 2026. https://www.cnn.com/2026/09/10/health/anthropic-bioweapons-report
Thumbnail: OpenAI Image 2 / hogby.ai
📬 Likte du denne?
AI-nyheter for ledere. Kuratert av en CIO som bygger det selv. Daglig i innboksen.