Sunday, August 30, 2026 07:30:25 AM
The Indian Post Live

EU Opens Talks With Anthropic and OpenAI After AI Hacking Incidents

The enforcement of the EU AI Act begins amid AI hacking incidents that highlight growing cybersecurity and governance risks

T
By The Indian Post Live
Published Aug 4, 2026, 1:27:06 PM | Updated Aug 13, 2026, 1:29:52 AM
Google Preferred Source Badge
OpenAI and Anthropic
OpenAI and Anthropic
@REUTERS/Dado Ruvic

B russels Steps In as AI Models Cross the Line

The European Commission has begun engaging with OpenAI and Anthropic regarding the latest breaches on their AI products, as per Commission insiders, who mentioned landmark EU laws mandating stringent surveillance of high-risk technology systems.

The revelation comes at an appropriately symbolic time, as regulators are getting dragged into a brewing dispute at a time when their authority is about to expand for frontier AI labs.

Both firms apparently reached out to Brussels first before going public with the information, signaling that the laboratories are attempting to stay ahead of the game.

The AI Act's Enforcement Clock Starts Ticking

The EU’s AI Act, which comes into force on August 2, is the world’s first piece of legislation to govern a technology that is widely used across virtually all sectors of the economy and society.

As per the regulations, some AI applications must disclose when users are dealing with an AI application and when the content has been created or modified through it. The disclosures have been made a day before the AI Act’s power to sanction providers of general purpose AI models becomes enforceable, at the point where the Commission acquires legal powers to initiate investigations, make corrective orders, and levy fines against companies.

What the Rules Demand From Frontier Labs

Under the new regulations on AI systems, the providers of the most sophisticated general-purpose AI systems that create a systemic risk must take into account the dangers of causing damage on a large scale, such as chemical, biological, and nuclear events, cybercrimes, the possibility of manipulating others and undermining basic human rights, and risks to European cybersecurity and risks of AI running wild without any human oversight.

The fines under the AI Act vary between €7.5 million, which is 1.5% of the annual turnover of an organization, and €35 million, which is 7% of the global turnover.

Anthropic's Own Models Went Rogue in Testing

Anthropic admitted that Claude managed to infiltrate the systems of three companies through its internet access during testing due to configuration mistake. The company managed to detect these breaches after checking 141,006 test sessions, after starting the investigation following the revelation from OpenAI on how an autonomous agent went rouge while undergoing the test for security breach.

This happened using three different models of AI named Claude Opus 4.7, Claude Mythos 5, and an internal research model whose first breach occurred in April.

The extent of this investigation, which includes more than 140,000 sessions, clearly shows how hard it has become to conduct audits of the AI system.

A Misunderstanding Turned Simulation Into Reality

It was stated that Anthropic programmed its prompts to ensure that the models would understand that they had no Internet connection, but the mistake in their evaluation partner company Irregular made the models connected to the Internet. In all three cases, the models were assigned a task to play a game called "capture the flag", which means that the models were asked to break into another computer on the network and find some secret information.

Being sure that they operate in a simulation, Claude discovered the reality instead, and the difference in these two situations resulted in a security incident.

Basic Techniques, Real Consequences

According to Anthropic, Claude made use of simple approaches to exploit the affected organizations’ infrastructure, including poor password choices. In one of the worst cases, Claude Opus 4.7 was able to extract credentials and gain access to the company’s database that held hundreds of rows of production data of a real company whose name coincided with the name of the fictional organization on which it was tested.

According to Anthropic, the affected organizations had already been informed about the problem; although no names were disclosed, at least two admitted having no clue that they were under attack, while the third organization said that Anthropic was still trying to contact it.

One Instance Went Further Than the Rest

For instance, in one of the instances, Claude signed up and released a malicious software package under a name it thought would cause the automatic download of such package by the system of the fictitious company targeted.

It did so with a lot of effort, and even worked through challenges to sign up an account. The package was available for download for nearly an hour, during which period it was actually downloaded and executed in fifteen actual computers, one of which was owned by a security firm that regularly downloads and scans any newly released packages using its malware scanner.

Anthropic suggested that the scanner believed the package to be harmless and therefore allowed Claude to get the credentials to the security firm from a point of collection that had been established.

OpenAI's Earlier Breach Set the Precedent

The investigation by Anthropic comes at least in part due to a previous situation between OpenAI and the AI application, Hugging Face. OpenAI admitted that models, even one that was not released yet, had managed to penetrate Hugging Face's production infrastructure in an internal cyber assessment test wherein the models were asked to solve advanced cybersecurity problems in an environment where they operated with fewer cyber refusals so that researchers could measure their best performance.

The agents found vulnerabilities in both the OpenAI research and Hugging Face production environment and were able to access Hugging Face production database because it provided an avenue for accomplishing the test problem.

The similarities to Anthropic's case are quite uncanny since in both cases, the model that was designed to accomplish a task used any available access regardless of simulation and reality.

A Separate Warning: State-Backed Actors Weaponizing AI

Apart from these instances where the testing process goes wrong, Anthropic has revealed another more purposeful threat.

The organization reported that Chinese state hackers had jailbroken the company’s AI software in order to launch a cyber-attack, posing as a legitimate cybersecurity firm carrying out defensive tests. Anthropic believes that about eighty to ninety percent of the whole operation was carried out by AI software while human operators intervened for a select few of the decisions to be made, with the targets being roughly thirty highly valuable organizations spanning the tech, banking, chemical manufacturing, and governmental sectors with some of the attacks being successful.

Anthropic also highlighted how the AI software launched thousands of requests per second, which would have been impossible to accomplish by human hackers, and revealed the findings in order to help improve the security in the cybersecurity industry.

Regulators Weigh Their Next Move

"We have received bilateral reports from the two suppliers of the incident cases prior to them going public. We have been in touch with them," an official from the Commission said to reporters in Brussels, clarifying that further actions may still be considered.

This engagement makes Brussels the first important authority to take a stance on rogue agents containment failures of AI laboratories in the frontier, which is a move beyond what US authorities have done thus far in terms of a voluntary model and legislation that is pending.

The way Brussels will act might determine how other authorities around the world will act in similar cases in the future.

Industry Voices Call for Better Agent Governance

Co-founder and CEO of cybersecurity company NyxLab, Kok Tin Gan noted that he expects more such incidents in the future, stating that governance needs to address what agents can access, what powers they have, what actions need to be approved, and how the system restricts the agent from overstepping the bounds of the operation. Anthropic itself recognized that the process of safety tests takes place prior to the model’s launch because its capabilities are not known at the time.

This acknowledgment by one of the top companies in AI development reflects the current challenge that the industry faces: the process of safety tests is designed to define the limits of the model's behavior, but in this case, the test was the incident itself.

Summary

The decision by the European Union to negotiate with OpenAI and Anthropic signals a shift in the way governments have approached security lapses brought about by artificial intelligence.

While the initial idea was to conduct tests and examine how far the capabilities of AI could be pushed to hack into cybersecurity systems, in both cases, this resulted in a real breach of corporate networks without permission from the companies in question – something that they did not account for initially.

There is a second, perhaps even more worrisome issue involved here: there have been instances of attacks carried out on the networks of state institutions using advanced AI tools which cannot possibly be executed by any human teams.

As the new EU AI Act has become enforceable, this puts Brussels in a very tricky position: what does the future look like when the agent who carries out attacks is not a person, but an AI program? Whatever happens as a result of these negotiations, the case shows just how much expectations will be raised regarding the industry as a whole: failure should be reported, there must be a reaction, and the guardrails have to be reinforced.