When Rubrik launched Project Hourglass at its FORWARD 2026 conference in Las Vegas back in June, the initiative set out to answer a question CISOs were already losing sleep over: what happens when an AI agent writing and deploying your code does something catastrophic and nobody can stop it in time?
On Thursday, at its GSI Summit in Goa, India, the cybersecurity firm announced the next step. Rubrik has expanded Project Hourglass to include a new tool called Rubrik Code Guardian, and has welcomed AHEAD, Trace3, and World Wide Technology (WWT) into the alliance, joining the six systems integrators that signed on at launch. The new capability is powered by Anthropic's Claude Mythos 5 and extends the alliance's reach from securing AI agents during execution to proactively identifying vulnerabilities in software code before deployment.
The Problem Driving All of This
Rubrik Zero Labs surveyed more than 1,600 IT and security leaders for its State of the Agent report and found that 86 percent expect AI agents to outpace their security guardrails within a year, while only 23 percent report full visibility into agents operating in their environments. More than 80 percent said agents require more manual oversight than the efficiency they save.
A separate Rubrik and Economist Enterprise study found 88 percent of enterprises experienced an AI agent security breach in 2026. The picture those numbers paint is one of organizations racing to deploy autonomous systems while the controls meant to govern them are still catching up.
What Code Guardian Does
The original Project Hourglass, which launched with Cognizant, Deloitte, LTM, HCLTech, NTT DATA, and Wipro as founding partners, focused on protecting AI agents at runtime through Rubrik Agent Cloud. That platform operates across three layers: Runtime Agent Security for behavioral guardrails and blast-radius control, Agent Rewind for fast repository recovery, and AI Context Guard for prompt integrity and control-plane protection.
Code Guardian shifts the security lens earlier in the development cycle. Rather than running against live production environments, the system tests a cloned, air-gapped copy of customer repositories. Claude Mythos 5 operates inside Rubrik's security harness to evaluate code against sophisticated, multi-step threat scenarios.
Three core capabilities define the product: isolated red-team analysis inside the air-gapped environment; attack chain discovery that reasons across files, services, identity roles, and cloud perimeters to find chained vulnerabilities that conventional static tools miss; and business impact prioritization that filters findings by actual exploitability and business criticality to reduce alert fatigue.
Alok Agrawal, Chief Solutions Officer at Rubrik, put the challenge plainly. "Engineering teams are turning to AI models to accelerate software delivery, but speed cannot compromise security or architectural integrity. By incorporating Rubrik Code Guardian into Project Hourglass, we are ensuring engineering teams are able to conduct red-team analysis, prioritize business impact and reduce risk."
The New Partners
Each of the three incoming partners brings a different angle to the coalition.
AHEAD, which builds enterprise technology architectures for large clients, framed the problem as one of inherited risk. Steven Sorensen, Specialty Solutions Engineer for Cyber Resiliency at AHEAD, said the goal is to get code attacked, tested, and validated in a safe environment before it ships, so teams can move faster knowing Rubrik's recovery capabilities sit underneath them as a safety net.
WWT's Chris Konrad, Vice President of Global Cyber, pointed to the company's Advanced Technology Center, where isolated threat testing has long been part of how it validates security solutions before recommending them. He said Code Guardian integrates naturally with that process.
Trace3 brought a perspective the others did not: its own teams have been running Rubrik Agent Cloud internally to protect their own agentic AI work before recommending it to clients. Sandy Salty, Chief Marketing Officer at Trace3, said that experience as both a user and a partner gives the firm a clearer view of what actually makes agentic environments more resilient at scale.
Where Things Stand
Rubrik Code Guardian is currently in private preview and accepting select design partners. It is not yet generally available and may change or be discontinued. Rubrik Agent Cloud, the original platform at the center of Project Hourglass, remains available to enterprise clients.
Dev Rishi, GM of AI at Rubrik, has previously described the core problem in stark terms: with AI agents, there is the potential for ten times the damage in one-tenth the time. That framing captures why the urgency behind Project Hourglass is unlikely to ease. As the volume of AI-generated code increases across enterprise environments, the window between a vulnerability being introduced and it being found by someone with bad intentions keeps getting smaller.
Reported by TechRadar, the vulnerability could allow attackers who already have valid login credentials to increase their privileges within Exchange and access mailboxes belonging to other users. Microsoft released the fix on October 2, ahead of its originally intended schedule.
The security issue stems from a weakness in authorization controls, which determine what information a user can access. By exploiting the flaw over a network, an authenticated attacker could gain permissions beyond those assigned to their account.
Threat actors could first obtain credentials through phishing attacks or by purchasing stolen login details from underground online markets. Once inside an organization’s Exchange environment, they could exploit the vulnerability to read emails and attachments belonging to other employees.
However, the flaw does not provide unrestricted access across different customer environments, known as tenants. It also does not directly grant administrator-level or SYSTEM-level privileges on the underlying Windows server.
Successful exploitation could expose sensitive business information, including financial records, invoices, contracts, customer correspondence, internal discussions, and confidential documents.
Attackers could use the stolen information to support further cyberattacks. For example, they might impersonate company employees, send convincing fraudulent emails, or conduct business email compromise (BEC) scams to trick organizations into transferring money or disclosing additional information.
The vulnerability is particularly problematic because an ordinary employee’s compromised account could potentially provide an entry point for accessing information held in other employees’ mailboxes.
The vulnerability affects the following on-premises products:
Microsoft has confirmed that Exchange Online customers are protected by a server-side fix. Organizations running affected on-premises installations should install the appropriate security updates promptly.
Exchange Server 2016 and 2019 have reached the end of their standard support lifecycle. Eligible organizations must be enrolled in Microsoft’s Extended Security Update programme to receive the relevant updates for these versions. Microsoft recommends that organizations without the required coverage migrate to Exchange Server Subscription Edition.
Microsoft reported no evidence that the vulnerability was being actively exploited at the time of the report. However, the company assessed that exploitation was more likely, making timely patching important.
Administrators should run Microsoft’s Exchange Server Health Checker after installing the update to verify successful deployment and identify any additional actions required.
Southern Company, the Atlanta-based energy holding giant, has confirmed that an unauthorized third party broke into its online customer portal and accessed account information belonging to roughly 400,000 customers across its electric subsidiaries, Georgia Power, Alabama Power, and Mississippi Power.
The company publicly disclosed the breach on October 5, 2026, stating that the intrusion, which occurred in September, was detected through its own monitoring systems. Upon discovery, Southern Company said it moved quickly to shut down the unauthorized access and notified law enforcement.
"Southern Company recently detected suspicious activity involving our online customer portal," the company said in a statement. "An unauthorized third party accessed certain, limited information about the accounts of approximately 400K customers. Upon detection, we took immediate steps to stop the activity and have engaged law enforcement."
Of the roughly 400,000 accounts compromised, approximately 300,000 belong to Georgia Power customers. Alabama Power accounts for the remaining 100,000, which represents about 6% of its 1.6 million-customer base. Mississippi Power was named in the company's public notice as affected, but Southern Company has not released a customer count for that subsidiary.
According to the company's published incident notice, the data the attacker accessed includes customers' names, mailing addresses, phone numbers, email addresses, and the last four digits of their Social Security numbers, along with other unspecified basic account details. The company was clear about what was not taken: bank account numbers, payment card numbers, and driver's license numbers were not part of what the attacker pulled from the portal.
The online portal in question is the same platform customers use to pay bills and manage their accounts.
Despite the company's confirmation that it caught and stopped the intrusion, Southern Company has not said when exactly in September the breach occurred, nor has it explained how the attacker got past the portal's defenses in the first place. That absence of technical detail leaves open questions about the attack method, whether it involved stolen credentials, a software vulnerability, or some other approach.
Customers whose information was involved are being notified directly by both mail and email. Southern Company is also offering each affected customer one year of free credit monitoring through Equifax. Those with additional questions were directed to a dedicated assistance line at 1-800-900-6021, available Monday through Friday.
Southern Company, incorporated in 1945 and headquartered in Atlanta, is one of the largest energy holding companies in the United States. It supplies electricity through its subsidiaries to customers across three states and operates natural gas distribution businesses in four more. Its total customer base numbers over 9 million.
The incident sits alongside a string of recent cyberattacks targeting utility companies in North America and beyond. Just weeks earlier, in September 2026, Houston-based CenterPoint Energy disclosed that an unauthorized third party had obtained personal information on a portion of its customers through one of the company's external-facing systems. That disclosure came after a threat actor claimed on a cybercrime forum to have extracted 7.49 million customer records from the utility and ultimately leaked the data after the company failed to respond. In March 2025, Nova Scotia Power, Canada's largest electric utility in that province, suffered a cyberattack that disrupted its customer support line and web portal.
Research published by cybersecurity firm Trustwave in 2025 found that roughly 67% of credential access techniques used against energy and utilities companies involved brute force methods, which are largely automated attacks that cycle through large lists of usernames and passwords. That tactic, known as credential stuffing when it draws on passwords stolen from previous breaches on other platforms, has been used against utility customer portals before. In 2021, UK energy supplier Npower had to shut down its mobile app entirely after attackers used stolen credentials from other websites to access thousands of customer accounts.
The energy sector's digital attack surface has expanded considerably as utilities push more customer services online. Customers log into portals to check bills, set up autopay, and track usage, all of which requires storing personal data behind web-based logins. Security researchers have long pointed out that these portals often receive less security investment than back-end operational systems, even though they handle data that is directly useful for phishing, identity fraud, and social engineering.
Southern Company said it has found no evidence of ongoing unauthorized access following the breach, and noted that it continues to monitor its systems around the clock. The company's public notice also cautioned customers to be alert for fraudulent contact from individuals claiming to be Georgia Power or Alabama Power representatives, reminding them that neither utility will ever threaten immediate service disconnection or demand payment over the phone.
Anyone who believes they may have been affected and has not yet received notification from the company can contact Georgia Power customer service directly through the number listed on their bill.
According to a report by SecurityWeek, the suspect is a 28-year-old Russian national who was arrested in Osaka, Japan, in May 2026. German authorities accuse him of playing a role in the Qilin ransomware operation, which has targeted organizations in several countries.
The suspect was extradited to Germany on October 2, 2026, following legal proceedings in Japan. German investigators allege that he was involved in an attack against a logistics company in September 2024.
Investigators say the targeted logistics company had its computer systems encrypted during the attack. The attackers allegedly demanded more than $160,000 in cryptocurrency in exchange for restoring access to the affected systems.
The incident is believed to have been connected to the Qilin ransomware operation, also known as Agenda. Qilin emerged as a ransomware-as-a-service (RaaS) operation in 2022. Under this model, ransomware developers provide malware and infrastructure to affiliates, who conduct attacks against victims and share the proceeds with the operators.
This business model has allowed ransomware groups to expand their operations without every attacker needing to develop their own malware or infrastructure.
Qilin has become one of the more active ransomware groups in recent years. It has targeted organizations across different industries, including healthcare, manufacturing, professional services and other critical sectors.
The group attracted international attention after attacks linked to it disrupted healthcare services in the United Kingdom. It was associated with the 2024 attack against Synnovis, a medical services provider whose systems were used by several London hospitals. The incident caused significant disruption to healthcare operations.
Qilin has also been linked to attacks against major companies in other countries, demonstrating the international reach of the ransomware operation.
Experts have continued to track the group's activities and techniques. In 2026, Qilin was also reported to have exploited a critical vulnerability affecting Check Point security products as part of its attack activity.
A security feature designed to make AI text traceable appears to have an unintended side effect: it can change how a language model behaves, in some cases making it more willing to follow instructions it was built to refuse.
That is the core finding from new research by Lasso Security, published on September 17 by researcher Andrea Siposova under the title "The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior." The study tested Google DeepMind's SynthID-Text watermarking system across six open-weight language models and found that enabling watermarking changed how those models responded to harmful requests, particularly when attackers used prompt-injection techniques designed to override a model's instructions.
What SynthID-Text Actually Does
To understand why that matters, it helps to understand what SynthID-Text actually changes inside a model.
When a language model generates text, it does not write the way a human does. It builds sentences one token at a time, each step producing a probability distribution over thousands of possible next tokens and then sampling from that distribution. SynthID-Text does not attach a label to the finished output or hide characters in whitespace. It intervenes at the sampling step itself.
The system uses a process called tournament sampling, first described in a 2024 Nature paper by Google DeepMind researchers Sumanth Dathathri, Abigail See, and colleagues. Instead of the standard randomness used in token selection, the watermark substitutes pseudorandom values generated from a secret key and the context of tokens already produced. The result is a statistically detectable pattern woven through the text at the level of individual word choices, invisible to readers but recoverable by anyone holding the key. The technique is refined enough that it does not degrade text quality in any measurable way, which is a large part of what made it attractive as a compliance tool.
The Regulatory Push Behind It
Article 50 of the EU AI Act requires providers of generative AI systems to mark their text, image, audio, and video outputs in a way that is machine-readable and detectable as AI-generated. The Code of Practice the European Commission finalized on July 20, 2026, requires at least two marking layers for audio, images and video, plus watermarking of free-form text longer than 200 tokens. Google signed the Code on July 24, 2026, citing SynthID partnerships as its route to the interoperable detection requirement due on February 2, 2027.
Anthropic's technical implementation is built directly on SynthID-Text, the same tournament-sampling mechanism from the Nature paper, adapted for their own models and keys. It covers claude.ai, the API, Claude Code, Claude Cowork, Claude Tag, and access through AWS, Google Cloud, and Microsoft Foundry. With the industry's largest players now committed to the same watermarking standard, Siposova's findings arrive at an uncomfortable moment.
What the Research Found
Siposova ran paired experiments on six open-weight models using Hugging Face's unmodified SynthIDTextWatermarkLogitsProcessor, enabling and disabling watermarking while keeping the seed, batch composition, and ordering identical.
Watermarking changed refusal behavior on harmful requests, but the effect was more pronounced when the same requests were paired with prompt-injection techniques. Under those conditions, several watermarked models complied with harmful requests their unwatermarked versions had rejected, with the strongest differences appearing in prompt-injection scenarios.
Siposova told Ars Technica that behavior was clearly different compared with the same model without watermarking, and that the differences were especially pronounced under adversarial conditions or when running an agent calling tools. She put it plainly: "Watermarking is made to not be perceptible to a reader, but we know that when we are changing anything about what the model is generating, it is going to cause some tradeoffs, it's going to show up somewhere."
Lasso repeated its prompt-injection experiment using 11 different watermark keys and found that changing the key could change both the size and direction of the effect. That means two different deployments of the same watermarking system, on the same model, could produce different safety outcomes depending purely on which key was chosen.
The Agent Problem
The consequences extend beyond what a model says. When a language model powers an AI agent, token selection can determine which tool an agent invokes and which arguments it passes. "Such a watermarking procedure can therefore affect both what the model says and what an agent does," the study stated. Watermarking reduced accuracy on six of seven models, with a substantial decrease on four.
Lasso's conclusions are deliberately careful. The study did not verify how Claude model responses change when watermarking is applied, and critics noted the experiment only validated the SynthID-Text tournament sampling implemented by Hugging Face, which differs from how the Claude model would actually implement it. Siposova stressed that the findings describe patterns in the tested sample, not universal rules about watermarking as a category.
Lasso recommended repeating agent evaluations and red-team testing when watermarking or its configuration changes, particularly under adversarial inputs.
The research surfaces a tension that will only sharpen as regulatory deadlines close in. Watermarking was designed to answer the question of where AI text comes from. What it can also do, under the right adversarial conditions, is change what the AI decides to do next.
The incidents were disclosed on October 4 and reveal the risks organizations face when attackers gain access to legitimate employee accounts. Nikkei has not attributed either incident to a specific hacking group or confirmed whether the two attacks were connected.
The more recent incident involved an employee's Microsoft 365 account. According to Nikkei, attackers gained unauthorized access to the account and used it on September 30 to send approximately 9,000 emails.
The messages were sent to people both inside and outside the company. Some recipients were journalistic sources and other individuals who had previously communicated with Nikkei employees.
The emails contained links leading to malicious websites. Because the messages were sent from a legitimate Nikkei employee account, recipients could have been more likely to trust them. This type of account compromise can allow attackers to use an organization's existing relationships to distribute phishing messages.
Nikkei said the incident may have exposed recipients' names and email addresses, along with the contents of some emails. The company is still investigating the number of people whose personal information may have been affected. According to Nikkei, “There may be an increase in emails impersonating Nikkei employees or our group companies,”
Nikkei also disclosed a separate incident involving an employee's Google Workspace account. The account was accessed without authorization beginning in late July.
The company discovered the intrusion in early August after receiving an alert from Google. Nikkei then changed the account's password and said it has not detected any further unauthorized access.
The incident may have exposed information belonging to 1,646 employees and business partners. The potentially affected information included names and email addresses.
Nikkei said the exposed information did not include data related to its readers or journalistic sources. The company also said it has found no evidence that the information was misused.
After the Microsoft 365 incident, Nikkei changed the affected password and contacted recipients of the phishing emails, asking them to delete the messages. The company warned that additional emails impersonating Nikkei employees or its group companies could appear.
Nikkei has also reported the incidents to Japan's data protection authority. Investigations into the scope of the Microsoft 365 compromise and the information potentially exposed are continuing.
Nikkei has experienced other cybersecurity incidents in recent years. In November 2025, the company disclosed a malware-related credential theft incident that potentially exposed information connected to more than 17,000 employees and business partners.
Apple has announced plans to overhaul one of macOS's most powerful privacy settings, citing security risks posed by AI agents that have been using it to access user data in ways most people never anticipated.
The setting, Full Disk Access, lives inside Privacy & Security in macOS Settings and was introduced with macOS Mojave (version 10.14). It gives users control over which applications can read system-level data, including files, Mail, Messages, Safari history, and Time Machine backups. Security tools and backup software rely on it legitimately. The problem is that once granted, an application can bypass many of the protections Apple built to keep sensitive data off-limits to third parties.
Apple warned in a developer advisory that some developers are using Full Disk Access to expose everything on a user's system without their full knowledge, and that for communication apps this also compromises the privacy of the people those users are messaging. The company said it plans to update the setting so access can only be granted through an explicit user action, and that as AI agents grow more capable and autonomous, the risks tied to this level of access will only increase. When the new controls will arrive has not been said.
The announcement follows a controversy involving Meta's personal AI agent, Muse. When the app launched on September 8, Inc. columnist Jason Aten installed it and says he explicitly declined to give it access to his Messages, calendar, or personal data. Despite that, Muse pitched him a column idea drawn from a private text exchange with his podcast co-host. Aten says Full Disk Access was disabled on his machine, yet the agent had synced more than 187,000 rows of his private iMessages to Meta's cloud. When he asked Muse directly how it read those conversations, the agent told him the paired Mac app was only relaying notification previews, an explanation that turned out to be false. Meta's David Singleton later called it a fabricated account of the feature.
Meta disputed the broader account. Singleton and communications head Andy Stone both argued that reading Messages requires two separate permissions: Full Disk Access must be active in macOS, and the Messages connector within Muse must also be switched on. Singleton said that without Full Disk Access, all related options are greyed out and the feature does not work. Whether that permission was ever active on Aten's Mac is something the two sides still disagree on.
What the episode made clear, regardless of how that specific question gets resolved, is exactly the scenario Apple is now trying to prevent: AI agents accumulating sweeping system permissions that users did not fully understand they had handed over.
The problem extends beyond confusing permission dialogs. On September 21, security researcher Patrick Wardle, founder of the Objective-See Foundation, published a zero-day flaw in Muse's Mac app before Meta had a patch ready, accompanied by a working proof of concept called "not-a-mused." The flaw centered on an undocumented configuration setting called "endo_voyager_dictation_endpoint" that any unprivileged local process could overwrite without admin rights and without triggering macOS security prompts. An attacker who had already landed on the machine could use it to redirect Muse's dictation traffic, capture audio and prompts, inject malicious instructions, and take advantage of every permission the agent held, covering files, microphone, camera, calendar, location data, and linked iOS devices. Wardle's proof of concept demonstrated over 50 commands being executed through the compromised agent.
Meta deployed a patch within 24 hours of disclosure, but Wardle argued that a ClickFix-style attack could have made the exploit remote, giving attackers access to any device running Muse, not just machines they had already penetrated by other means. He described Muse's extensive system permissions as making it trivial to turn the agent into a ready-made backdoor.
Wardle also separately reported a flaw in OpenAI's ChatGPT Mac app, tracked as CVE-2026-100754, that could have let attackers take over the assistant and access chat logs and other stored data. He described the exploit as insanely trivial, requiring roughly a dozen lines of code, and noted it could also be used to get ChatGPT to run commands on an attacker's behalf, with the requests appearing as legitimate instructions from the OpenAI software. OpenAI has since patched it.
Wardle has said he will present analysis of multiple AI macOS application vulnerabilities at Objective by the Sea, an Apple-focused security conference in November, and has already submitted a further finding to OpenAI related to the integration between ChatGPT and the company's always-on Dots AI assistant.
The pattern across all three incidents points to a structural problem the industry has not resolved. AI agents need deep system access to function, and that same access makes them attractive targets. "AI companies are fixated on adding features right now," Wardle said, and the permission frameworks macOS relies on were not designed with always-running, autonomous agents in mind. Apple's planned changes to Full Disk Access are an attempt to close that gap, though what those controls will actually look like when they ship remains unknown.