An overarching number of 700 artificial intelligence agents participated in a coordinated attack against Hugging Face after models running...
METR said the agents were not simply attacking systems randomly but were attempting to find ways to manipulate the ExploitGym evaluation process.
"Agents coordinated on large collective projects to cheat the ExploitGym scorer, and attacked Hugging Face for clues," METR said. "The agents with impossible tasks had a common objective: to find a general-purpose way to trick or tamper with the automated ExploitGym scorer to get it to give a passing score on impossible tasks. As soon as the message board was established, agents started discussing this question and sharing information and resources about cheating the scorer."
OpenAI said the activity observed in May ultimately provided the groundwork for the later Hugging Face attack. After the original message board was removed during the response to the July 4 outage, the agents found another method of communicating through Artifactory.
The company acknowledged that the broader implications of the unauthorized communication system and unintended internet access were not initially recognized.
"At the time, the broader containment and alignment implications of the improvised message board and unintended internet access were not yet understood," it added. "In short, an internal team observed an agent engaging in message board activity and instances of disallowed internet access in late May."
"However, the existence of the improvised message board and the significance of the inter-agent communication activity were not apparent to the leaders responsible for the July 5 incident detection and response. We are continuing to review the process and operating practices that shaped detection and response activities leading up to this incident."
OpenAI identified four major misalignment patterns behind the incident: reward hacking, persistence with apparently impossible tasks that encouraged increasingly risky strategies, unauthorized communication through persistent message boards, and an interconnected agent environment that allowed agents to divide responsibilities and adopt objectives from one another.
"Aside from misalignment, another factor for this incident was that we did not extend the powerful safeguards that we deploy for our externally deployed models to all internal evaluations," it pointed out.
Following the investigation, OpenAI said it is strengthening security and alignment controls across its model-development process. The measures include tougher alignment requirements throughout a model's lifecycle, more isolated sandboxes, tighter restrictions on internet connectivity and stronger controls over access to model weights.
The company also said agents are now expected to request clarification or safely stop when they encounter difficult or seemingly impossible tasks, rather than continuing to pursue "increasingly questionable alternatives."
OpenAI described the incident as a significant warning about the risks associated with increasingly capable autonomous AI systems.
"We are taking this incident as a 'warning shot' that today's model capabilities present the possibility of loss-of-control incidents," OpenAI concluded. "Companies that build AI systems will need to ensure that their systems always remain under meaningful human control, and that meaningful safeguards constrain their ability to cause harm."
The company also warned that similar capabilities could eventually be deliberately deployed by malicious actors.
"As comparable capabilities become more widely available, others may also use them deliberately to carry out attacks. Both model developers and cyber defenders more broadly will have to prepare for AI-enabled attackers that work faster, at a larger scale, and with better coordination than human attackers."
The U.S. Department of Justice (DoJ) and Federal Bureau of Investigation (FBI) have disrupted two hacking platforms operated by a China-linked threat group that were used to conduct reconnaissance, compromise vulnerable systems and conceal attacks against U.S. government agencies, critical infrastructure and other sensitive organizations.
The platforms, QScan and QTRouter, have been attributed to QTFY, a Chinese state-sponsored hacking group linked to Nanjing Xinjiuwei Network Technology Company. According to U.S. authorities, QTFY activity has targeted organizations including NASA, the Federal Reserve, Department of Energy, Department of Justice, Department of Health and Human Services, National Institutes of Health and the U.S. Senate.
Lumen Black Lotus Labs, which tracked the infrastructure for more than 18 months, said QTFY activity dates back to at least May 2018. The researchers described the group as an infrastructure "quartermaster" that developed reusable systems for reconnaissance, exploitation and traffic obfuscation.
QScan automated reconnaissance and exploitation
QScan formed the reconnaissance component of the operation. The platform scanned internet-connected systems and IoT devices for vulnerabilities before automatically compromising susceptible devices and incorporating them into the QTRouter network.
The FBI said QScan was also used to identify vulnerabilities in victim networks. Its infrastructure included servers responsible for distributing scanning tasks to worker nodes and collecting completed results.
The scale of the operation allowed QTFY to conduct reconnaissance across large numbers of systems. Lumen identified scanning activity spanning more than 130 countries, with targets including government, defense, aerospace, healthcare, financial, energy and research organizations.
QTRouter concealed attackers' origins
Compromised devices identified through QScan were subsequently used by QTRouter as proxy nodes. The network combined hacked IoT devices with commercial proxy services and leased virtual private servers (VPSs), allowing malicious traffic to pass through multiple intermediary systems.
This architecture made an intrusion originating from China appear to come from an internet connection located elsewhere. In some cases, QTRouter could route traffic through systems geographically close to the targeted organization, making the activity appear more consistent with legitimate local traffic.
QTRouter operated on routers running customized OpenWrt software and used the Clash proxy framework to establish connections. Operators could select available nodes and chain them together, creating multiple layers between themselves and their targets.
The FBI said this combination of compromised IoT devices and legitimate commercial proxy infrastructure made malicious traffic difficult to distinguish from normal internet activity.
Attackers exploited new and older vulnerabilities
QTFY's attack chain involved both recently disclosed and long-standing vulnerabilities. The vulnerabilities identified by investigators included flaws in Ivanti Connect Secure, Fortinet SSL-VPN, Citrix ADC, Microsoft Exchange Server, F5 BIG-IP, Kentico CMS, Apache Log4j, Atlassian Confluence, Check Point Quantum Gateway, CrushFTP and BeyondTrust Remote Support.
After obtaining initial access, QTFY actors used remote access trojans, web shells and legitimate credentials to maintain persistence.
The infrastructure could subsequently provide concealed access into victim networks through nearby compromised IoT devices. QTBotnet also allowed operators to control infected systems, execute commands and conduct distributed denial-of-service attacks.
Four-part infrastructure supported QTFY operations
Lumen identified QScan and QTRouter as part of a larger architecture that also included Fast Labyrinth and QTProxy.
Fast Labyrinth incorporated commercial proxy infrastructure into encrypted relay paths, while QTProxy managed operational nodes and allowed operators to configure routes toward selected targets.
The researchers compared the architecture to an operational relay box, or ORB, network. Such systems use compromised devices and leased infrastructure as rotating relay points, making traditional IP blocklists and location-based defenses less effective.
Lumen said the infrastructure demonstrated an increasingly industrialized model of China-linked cyber operations, in which reusable and shared services can provide reconnaissance and anonymity at global scale.
FBI seized domains used by the platforms
The disruption targeted domains hard-coded into QScan and QTRouter, including infrastructure used to distribute scanning tasks and administer proxy connections.
By seizing these domains through court-authorized action, U.S. authorities disrupted communication between the platforms and their operators, causing the systems to cease functioning.
Investigators also linked QTFY to Chinese cyber-brokering networks where exploits, malware and access to compromised organizations were allegedly traded. Nanjing Xinjiuwei was described by U.S. authorities as an enabling company with relationships across China's cyber ecosystem and connections to former People's Liberation Army personnel.
QTFY activity reportedly continued into June 2026, when actors targeted a U.S. election system.
The disruption demonstrates how China-linked threat actors are increasingly relying on distributed infrastructure rather than fixed attacker-controlled servers. While domain seizures can interrupt an operation, the reuse of compromised IoT devices, commercial proxies and leased servers means defenders will need to monitor behavior and network relationships rather than rely solely on static IP-based blocking.