XMACNA
Anthropic details how Claude was used in cyberattacks, surveillance, and weapons

Anthropic details how Claude was used in cyberattacks, surveillance, and weapons

The new Anthropic threat intelligence report describes interrupted operations with Claude in cybersecurity, surveillance, influence, fraud, biology, conventional weapons, and distillation. The key detail is not an 'out-of-control' AI: it is the speed, scale, and self-
XMACNA Team

8 min read

Analysis

Anthropic published its most detailed threat intelligence report so far in 10 September. The company says it identified and stopped operations using Claude between December 2025 and August 2026 in seven areas: cyber operations, influence, surveillance, scams and fraud, biological use, development of conventional weapons, and model distillation. Accounts were banned, indicators were shared with authorities and partners, and findings were incorporated into platform safeguards.

The central point is not the fantasy of an AI deciding to attack the world alone. It is more concrete: human groups are connecting capable models to tools, credentials, data, and persistent routines. The model can write, analyze, translate, prioritize, and, in some workflows, orchestrate actions. This compresses time and cost for attackers — and forces defenders to treat each agent as part of an operation, not as an isolated chat box.

What the report states — and what it does not prove

The report is a publication by Anthropic itself, based on investigations by its Threat Intelligence team. It documents cases the company says it detected and stopped; it is not an independent audit of all Claude uses nor a universal measurement of AI risk. The distinction matters because the company has privileged access to signals from its own platform, but also incentives to publicly explain why its controls are necessary.

There are three different verbs in the document: identify, interrupt, and share. Identify means finding suspicious patterns in interactions and associated systems. Interrupt includes banning accounts, blocking flows, and creating detections for future attempts. Share means providing indicators or context to authorities and other companies when activity exceeds the platform. None of these verbs mean that all potential harm was avoided or that actors had no other tools.

This is the first lesson for a company using AI: detection report is not synonymous with risk eliminated. Governance needs to follow what happened before the alert, which permissions were available, what data was exposed, and whether the same pattern can reappear in another tool.

From assistant to operator of parts of the chain

On the cyber front, Anthropic describes a role change. Instead of just asking for an explanation or code snippet, operators connected Claude to multi-agent frameworks for reconnaissance, exploitation, analysis, and exfiltration. People often continued to select targets and review results, but execution of steps became faster and parallel.

This does not turn every model into an autonomous intruder. The report also notes that autonomy and severity are separate axes: some of the most severe intrusions were human-led step by step. Autonomy, however, changes the economics of the attack. Less manual work per campaign means previously expensive or time-consuming targets may become feasible, and defense must detect distributed behavior over time.

For the security team, the consequence is operational. Prompt logs are insufficient when risk lies in the combination of model, tool, and identity. Permissions, external calls, context shifts, access to secrets, attempts to bypass controls, and produced results must be observed. This is the same principle as a Intelligence Cycle: previous context, current action, and record of next step must remain connected.

Surveillance and influence: the risk does not start with malware

The surveillance cases described by Anthropic show uses that might seem like “just software.” The company reports tools to collect identities on social networks, organize dossiers, analyze large volumes of posts, and operate monitoring systems. It also describes state actors and surveillance providers using Claude to turn scattered requirements into interfaces, extensions, and collection routines.

The problem for organizations is not only stopping malicious code. It is identifying when a sequence of seemingly neutral tasks — collect, correlate, classify, translate, and distribute — produces a surveillance or influence system. Privacy controls, purpose, retention, consent, and human review must exist before automation, not just after an incident.

The same logic applies to influence campaigns. A model can adapt messages, simulate interlocutors, and produce variations at scale. Harm does not depend on an isolated viral piece; it can arise from the combination of targeting, repetition, and appearance of authenticity. For those operating Digital Employees, identity, memory, and action limits are part of the product — not cosmetic details.

Biology and weapons: why “dual use” requires more context

The biology section is deliberately cautious. Anthropic presents five cases where actors used models in activities that could support biological weapons development. The company does not claim every interaction resulted in a weapon, nor that a model alone produced operational capability. The warning is that apparently legitimate requests can compose a dangerous project when viewed together, especially with signs of control evasion and access to external resources.

The report also covers conventional weapons: guidance software, drones, interception, and supply intelligence. In some cases, the company says there was field testing or simulation; in others, just design and acquisition work. The difference between “helped write” and “built a functional weapon” cannot be erased by a headline — and should not be used to minimize the importance of technical assistance in intermediate steps.

This is where the discussion leaves the lab and enters business governance. AI systems need risk classification, role-based controls, auditable records, tool limits, and human escalation. The goal is not to stop all sensitive research; it is to ensure the organization knows who can request what, with which data, for what purpose, and under what review.

What distillation changes in the model race

The seventh category, illicit distillation, covers attempts to extract and reproduce model capabilities through large volumes of interactions. This vector is less cinematic than an attack, but strategic: it can transfer months of research and infrastructure to a competitor or intermediary without reproducing the same investment.

For companies contracting AI, the consequence is twofold. First, credentials and data sent to a provider must be treated as assets with defined purpose and retention. Second, it is not enough to ask which model is used; it is necessary to know what integrations, exports, and observability mechanisms exist around it. A Digital Employee's Long-Term Memory, for example, must have scope, retention, and access consistent with the process risk.

The practical response for AI operators now

The Anthropic report does not call for panic. It offers a work map. A company using agents should:

  1. Design permissions by task, not by convenience: reading, writing, execution, and external access must be separated.
  2. Record context and outcome, not just prompts: identity, tool, accessed data, decision, and escalation need to form a trail.
  3. Test abuse with realistic scenarios, including evasion, persistence, and combining benign actions.
  4. Maintain a clear human handoff for fraud, surveillance, sensitive data, security, and irreversible decisions.
  5. Measure what the control prevented and what it let through, without confusing lack of alert with lack of risk.

This design approaches decision-oriented AI governance: each capability needs an owner, limit, evidence, and review path. The goal is to allow AI to do useful work without turning speed into invisibility.

Anthropic closes the report asking other companies to share signals and strengthen collective defenses. The request is legitimate but does not replace scrutiny. Providers must explain how they detect abuse; clients must demand verifiable controls; researchers and authorities must question methods, evidence, and incentives. AI safety is not a marketing promise. It is an operational property that must survive the first unscripted case.

Read the full Anthropic report and the original publication on X. To translate this debate into an assessment of your process, use the XMACNA Assessment.

Frequently asked questions

Does Anthropic claim Claude created weapons on its own?

No. The report describes people and organizations using Claude in research, development, intelligence, and software stages. Some cases include simulations or tests, but the publication does not prove the model alone produced an operational weapon.

Does this mean any company using AI is under attack?

No. The presented cases are selected examples of sophisticated abuse. They show risk patterns — autonomy, scale, evasion, and tool combination — that need to be part of each organization's threat model.

Did Anthropic's safeguards work?

According to the company, accounts were identified, banned, and used to improve detections. The publication also acknowledges that controls did not block all requests. Therefore, safeguards must be continuously evaluated, with monitoring and independent review.

What should a business leader do first?

Map which agents access data, tools, and external systems; define minimum permissions; record decisions and outcomes; and create a human path for exceptions. Starting with inventory and evidence is more useful than choosing a tool based on headlines.