Capsule Security launches AI shield to stop rogue acts
Thu, 3rd Sep 2026 (Today)
Capsule Security has launched a system designed to stop rogue AI agent actions before they are executed. The product is based on NVIDIA Nemotron models.
The Boston-based cybersecurity company says the system acts as an independent control layer, evaluating an AI agent's intended action immediately before execution so organisations can permit, flag or block the step in real time.
Researchers recorded 98% accuracy on StepShield, which Capsule described as an independent academic benchmark for testing whether security systems can identify and stop rogue agent behaviour before damage occurs. The system also identified violations at the exact step where they happened, a feature Capsule said matters when businesses are handling millions of agent actions.
The announcement reflects a broader shift in cybersecurity as companies deploy autonomous software agents that can access sensitive data, write code, manage infrastructure and interact with other systems with limited human involvement. In that environment, traditional permissions and approval processes may define what an agent is allowed to do, but not whether a specific action is appropriate in context.
Naor Paz, Chief Executive Officer and Co-Founder of Capsule Security, said the risk profile of AI systems is changing as agents become more autonomous.
"The defining AI security risk is no longer only what people can do with agents. It is what autonomous agents can decide to do by themselves," said Paz.
"When software can reason, use tools and take action, a wrong decision can become a real-world incident in seconds. Human trust in AI depends on our ability to stop that action before it happens," he added.
Model testing
Capsule fine-tuned two NVIDIA Nemotron models using training based on real agent traces, human review and adversarial examples designed to distinguish authorised behaviour from rogue actions. NVIDIA Nemotron 3 Ultra was used to support the training process.
Because the models are focused on classification rather than generating full responses, they can sit directly in an agent's execution path with limited delay, according to Capsule. In testing, the models made decisions in as little as 71 milliseconds.
Capsule also said its internal benchmark showed its most accurate detector reached 96.9%, compared with 86% for the strongest third-party model it evaluated. The company added that it reduced the larger model's memory requirements by nearly half without affecting performance, enabling it to run on a single NVIDIA L40S GPU.
Enterprise use
Capsule says its platform is already protecting billions of tokens across millions of agent interactions. Its customers include financial institutions and technology companies, though it did not name them beyond H&R Block.
Phillip Miller, Vice President & Global Chief Security Information Officer at H&R Block, described the technology as a way for security teams to supervise agent activity before unauthorised actions take place.
"AI agents represent a fundamentally new security challenge: they can reason, use tools, and take consequential actions at machine speed. Capsule helps organizations monitor agent behavior in real time and stop unauthorized actions before they execute. This gives security teams the confidence to expand their use of agentic AI while maintaining the security, governance, and accountability their clients expect," said Phillip Miller, Vice President & Global Chief Security Information Officer at H&R Block.
The product is available now. Capsule said it works across AI platforms, agent frameworks, software-as-a-service applications and endpoints, with one control layer intended to sit over existing enterprise architecture.
Capsule was founded by Naor Paz and Lidan Hazout. The company is part of Anthropic's Claude Security Program and was selected for Google's Gemini Startup Forum: Cybersecurity.