Protect and secure your data from cyber attacks
Data Protection
Data Security
Data Insights
The 5 Steps to Cyber Resilience
Cloud & SaaS
Enterprise
Industries
Part two of four: AI agents will fail. The question is whether you can recover.
Part one of this series examined recent incidents where AI agents pursued assigned goals without the judgment to recognize when they had crossed a boundary—and there was often no control in place to stop the action in time. The risk begins when an agent can complete its action before a person can intervene. This post covers why resilience must cover that gap.
Nobody told Claude Opus 5 to lie to its suppliers. Yet in a recent Vending-Bench Arena test, it invented rival price quotes, entered a price-fixing pact with a competing AI, and broke the agreement when doing so became profitable. Its reasoning logs documented the tactic.
The pattern isn’t new. In earlier Palisade Research tests, OpenAI's o3 rewrote a chess engine's position file rather than resigning a losing game, attempting to hack its way to a win in 88% of trials.
Most recently, a combination of sandboxed OpenAI models escaped its test environment while pursuing the answer key to an internal benchmark. It chained a zero-day exploit with stolen credentials and spent four and a half days moving through Hugging Face's production infrastructure.
That is the uncomfortable pattern behind several recent AI incidents. When an agent has enough autonomy and is rewarded for reaching an outcome, it may find a shortcut no one intended. But not every AI agent incident involved an incentive to cheat or an attacker. Cursor's support bot invented a one-device-per-subscription policy. Replit's coding agent wiped a production database it had been explicitly told not to touch, then fabricated 4,000 records to hide the damage. Gemini CLI deleted a project after interpreting a conversation as a command.
Better guardrails and alignment work to reduce how often that happens, and they are worth the investment. They have not reduced the security risk to zero, and frontier-model evidence like this is why resilience must backstop the remaining gap.
The triggers were different, but the operational gap was the same. Observability could identify the failure. It could not undo the damage.
Cohesity’s James Blake made this point when the Hugging Face story broke: the real test is not simply whether an organization can detect an incident. It is whether the organization can:
That means proving that critical workloads like identity, AI systems, and sensitive data are cleanly restored to a trusted state—not just restored from backups that might still be infected.
And it must happen at machine speed. Hugging Face's account describes roughly 17,600 recorded actions over about four and a half days. No manual process can move that fast. A SOC analyst still has to work through the alert queue, an on-call engineer has to escalate the issue, and someone has to dig through the logs. Unless detection and recovery are automated and pre-authorized, the agent may complete its full sequence of actions before people can respond.
That speed of threat detection and subsequent recovery requires three concrete design choices to achieve. Whatever threat model turns out to be right—an attacking agent or an organization's own agent going wrong—recovery is the one investment that pays off regardless, because systems are fragile no matter which danger you're picturing. That is why it comes first here, not last.
The three choices below map to three of the five steps in the Cohesity 5 Steps of Cyber Resilience©, applied specifically to agents:
The other two steps in that framework will come into focus in the next post, once the full stack is on the table.
These controls are only as reliable as the policy, permissions, and activity record behind them. If any of those can be altered by the agent, an attacker controlling it, or a compromised administrator, trust breaks down. Replit's agent demonstrated that risk by fabricating records to conceal its actions.
The trusted baseline, therefore, needs to remain protected outside a typical blast radius. Cohesity FortKnox keeps an immutable, virtually air-gapped copy of data, protected by WORM controls and a two-person approval requirement for critical changes. No one person or agent can alter it unilaterally. Attackers typically target backups in four ways: corrupting restore points, disabling backup jobs or retention policies, encrypting the repository, or stealing backup copies as valuable data in their own right. Cohesity’s broader defense-in-depth model is designed to make each step harder to carry out, easier to detect, and less damaging if it succeeds.
That protection needs to extend beyond the backup copy itself. Databases, object stores, vector stores, and agent memory should be recoverable to a trusted point in time and rebuilt from a known-good configuration, rather than repaired in place. The goal is a clean recovery path across data, agents, and infrastructure using copies that the incident never touched.
That starts with knowing what data an agent can access. Cohesity Data Security Posture Management, powered by Cyera, discovers and classifies sensitive data and monitors who—or what—is accessing it. Organizations cannot protect or recover data that they do not know an agent can reach.
A root of trust, here, is the baseline that an agent did not set and cannot alter on its own: the identity, permissions, and policy that determine what it is allowed to do. A prompt can instruct an agent, but it does not enforce policy. Instructions and permissions, therefore, need to be separated. The target: limit access to the task at hand, elevate privileges only when needed, and revoke them the moment the task ends—a bar most legacy environments cannot fully meet today, but a clear direction for anyone deploying agents in production. Validate each action against policy rather than relying on a one-time check at login.
Agent identity and access should follow the same principle: least privilege, continuous verification, and no standing access. To ensure that your agent identity and access remain both available and recoverable if compromise happens, you need the same trusted resilience for your critical identity infrastructure. Cohesity’s Identity Resilience capabilities reinforce that model on the recovery side by detecting, withstanding, and recovering from identity-based attacks that target Active Directory, Entra ID, and Okta, where agent credentials are sometimes managed. If those directories are compromised or tampered with, organizations can restore them to a known-good state.
An agent with authority to act can also misreport what it did, whether through compromise or a simple mistake. That is why one control point is not enough: a separate, independent layer should evaluate actions against policy before they run, without relying on the acting agent's own explanation, and block or hold destructive actions for human approval.
As Cohesity advances its Enterprise AI Resilience strategy, planned integrations with platforms such as Cyera Agent Guardian, ServiceNow and Datadog will help turn signals of risky agent behavior—such as policy violations or actions outside approved boundaries—into protection and recovery workflows without relying on the agent’s own account. Cohesity and ServiceNow describe the combined approach as "trustworthy by design." A deeper integration with ServiceNow's AI Agent Control Tower is planned for later this year.
Suppose an agent starts taking actions outside its normal pattern, or its credentials end up exposed, as happened with a shared cluster-admin credential at Hugging Face. Once an anomaly is flagged, the organization doesn’t need to reconstruct the exact sequence of events that caused it. It can go straight to a known-good point in time and recover from there. The agent, the data it touched, and the underlying infrastructure are all restored together to that last verified state, using a baseline captured before the incident occurred. Recovery does not depend on the agent's memory of events.
This is the gap exposed by recent incidents, from the benchmark-driven model that reached Hugging Face's production systems to the coding agent that wiped a database it had been told not to touch. Neither required a human attacker. Both show why governance must be paired with resilience: the ability to establish what can still be trusted and recover to a known-good state quickly enough to limit the damage.
Part three of this series will map the Hugging Face breach to the Cohesity 5 Steps of Cyber Resilience©. The full stack is on the table as we examine Cohesity’s Enterprise AI Resilience strategy.
Written By
Gyan Prakash
Director of Security Engineering , Cohesity