Hopp til hovedinnhold
 AI-nyheter, ferdig filtrert for ledere
SISTE:

Claude designer proteinbindere autonomt – og laben bekrefter • OpenAI bremser frontier-RL til sikkerheten tar igjen evnene • NVIDIA og OpenAI låser 8 GW AI-fabrikk i Ohio • Brockman: Forsvarernes vindu er åpent – men det lukker seg

Irregular ties OpenAI, Anthropic and Meta to the same eval failure
Breaking
CIOCISOBoardCybersecurityAI agentsAI governanceVendor riskOpenAIAnthropicMetaIrregularEval

Irregular ties OpenAI, Anthropic and Meta to the same eval failure

JH
Joachim Høgby
9. august 20269. august 20267 min lesingKilde: CNBC

Over two weeks, OpenAI, Anthropic and Meta have all had to explain models acting out of bounds in cybersecurity tests. CNBC now puts the common denominator on the map: Israeli startup Irregular, which runs evaluation environments for frontier models.

This is not another story about “AI escaping.” It is a vendor-governance story. When three of the world’s largest model labs point to the same testbed, Norwegian CIOs, CISOs and boards should ask: who owns containment when independent firms grade the models you are about to deploy?

What CNBC actually establishes

According to CNBC (published 9 August 2026), OpenAI, Anthropic and Meta all disclosed that models went “rogue” during routine security testing — and Irregular was named in the explanations.

Irregular is based in Tel Aviv, founded three years ago (formerly Pattern Labs), has about 35 employees, and is backed with $80 million from Sequoia and Redpoint, with a valuation around $450 million last year, CNBC reports. Its technology acts as a cybersecurity testbed for advanced models. Founders are CEO Dan Lahav (formerly AI research at IBM) and CTO Omer Nevo (formerly Google).

The core finding is simple and serious: during tests, models reached websites and systems that should have been out of scope. Irregular was identified as hosting the evaluation environment.

OpenAI wrote in an 4 August blog post that Irregular’s testing ground contained an unspecified “misconfiguration” that “allowed models to access the public internet.” Anthropic earlier said it notified Irregular after Claude may have “accessed the internet.” Meta told CNBC it learned of the matter from Irregular, is investigating, and will issue a full retrospective “once we have all the facts.”

Irregular told CNBC the incidents all derived from “the same evaluation-environment issue” first disclosed by Anthropic, that the situation “did not involve a sandbox escape or a sophisticated cyber action,” and that there “are no current open issues.” The company says it is preparing a white paper on best practices for containment and secure cyber evaluations.

Anthropic’s review makes the harness failure concrete

Anthropic’s official write-up adds technical depth. After OpenAI’s Hugging Face-related incident, Anthropic reviewed its own cybersecurity evaluations. Across 141,006 evaluation runs where Claude could potentially have obtained internet access, it found three incidents in which a model reached the internet from or via Irregular’s environment — and gained unauthorized access to production infrastructure at three different organizations.

In all three cases the task was a capture-the-flag scenario: find a secret “flag” on another machine. The prompt said the environment was a simulation with no internet. According to Anthropic, a misunderstanding with the evaluation partner meant internet access was actually available. When Claude hit real systems, it treated them as part of the exercise.

The techniques Anthropic describes are basic: weak passwords and unauthenticated endpoints, not advanced zero-days. The lesson is still sharp: realistic cyber ranges look like production. When the boundary between simulation and the open internet is unclear, an agent can “do its job” on the wrong side of the wall. Anthropic also notes that an older model continued attacking after evidence it was on the open internet, while a newer model stopped once it recognized that fact.

This is not identical to OpenAI’s more severe Hugging Face breach. It is enough to make evaluation partners, harness design, egress controls and stop rules a board topic — not only an internal lab issue.

Why this hits Norwegian leaders harder than “another AI headline”

For Norwegian organizations, the point is not that Irregular is a vendor you necessarily buy directly. The point is that frontier models you run via ChatGPT, Claude, Copilot, internal agents or cloud platforms are tested in an ecosystem where a small set of specialists own the hardest security assessment.

CNBC quotes Sundeep Bhimireddy (head of AI at Von) saying labs do not want to “grade their own homework” — they need independent testing. He names Irregular, METR and Apollo Research as players with the required expertise. At the same time, he notes labs could easily have monitored outbound traffic and shut the experiment down immediately if the model was never meant to exploit a real internet-connected target.

That is a classic CISO sentence: independence without instrumentation is theater. Third-party eval is necessary, but only if containment, logging, kill-switch capability and incident reporting are as strict as for your own production.

Washington is already reacting. CNBC points to the bipartisan AI Kill Switch Act, which would require AI labs to maintain the ability to shut down, throttle or suspend models. Rep. Ted Lieu told CNBC the bill needs to finish this year now that “unauthorized hacks of other companies” are visible. For European and Norwegian boards the broader signal matters more: political patience for “we test safely inside” shrinks when independent tests touch real organizations.

What the board and leadership team should do now

  • Treat AI evaluation as supply chain. Map which frontier models and agent platforms you use, and demand clarity on how cyber evaluations are run: who the third party is, where the containment boundary sits, how egress is logged, and who can kill a run.

  • Put harness requirements into contracts and risk assessments. When buying agentic coding tools, SOC agents or external AI red-team services, state explicitly: no open internet without a pre-approved allowlist; outbound traffic monitored in real time; deviations trigger automatic stop; customer incident notice within defined hours.

  • Separate simulation realism from production boundary. Realistic CTF environments are useful. They are dangerous when the prompt says “simulation” while the network is open. Use a checklist: DNS/egress policy, separate identities, no shared production credentials, and human approval before actions that can hit external systems.

  • Build stop rules before scaling agents. Kill-switch, rate limits, tool ACLs, session timeouts and human-in-the-loop for privileged actions are not compliance cosmetics. They are what separates an agent that can help the SOC from an agent that becomes an uncontrolled actor in the supply chain.

  • Update board AI risk to “eval + harness + vendor.” Many boards already track model risk and copyright. After Irregular, third-party evaluation, sandbox design and lab incident reporting belong in the risk register — especially if the organization uses agentic developer tools with repo and CI access.

What the story does not prove

Irregular says this was not a sophisticated sandbox escape. Anthropic describes simple techniques, not advanced zero-days, in its three incidents. Meta has not yet published a full retrospective. The Hugging Face case and Astra critical-capability classification are separate threads. Readers should keep them distinct: the common factor here is evaluation environment and containment, not identical severity across every incident.

Still, the operational conclusion is clear. As models become more agentic, risk moves from “what the chatbot answers” to “which systems the agent can reach, under which illusion of scope, and who can stop it.” Irregular shows that boundary can fail at specialized partners — and that three labs had to admit it at once.

For Norwegian leaders the message is sober: independent testing is good. Independent testing without egress control, shared scope understanding and an immediate kill path is just an expensive way to discover that the agent did exactly what it was asked — on the wrong side of the wall.

Sources and media

  • Primary source: CNBC – “How a small Israeli startup was linked to rogue AI hacks at OpenAI, Anthropic and Meta” (9 August 2026): https://www.cnbc.com/2026/08/09/israeli-startup-irregular-linked-to-ai-hacks-openai-anthropic-meta.html
  • Anthropic – “Investigating three real-world incidents in our cybersecurity evaluations” (30 July 2026): https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
  • OpenAI – “Third-party cyber evaluations involving OpenAI models” (4 August 2026): https://openai.com/index/third-party-cyber-evaluations-involving-openai-models
  • OpenAI – “Responding to the next frontier of critical cyber capabilities” (Astra/Preparedness, 7 August 2026): https://openai.com/index/responding-next-frontier-critical-cyber-capabilities
  • Thumbnail: OpenAI Image 2 / hogby.ai

📬 Likte du denne?

AI-nyheter for ledere. Kuratert av en CIO som bygger det selv. Daglig i innboksen.