Skip to content
News

Nvidia launches OpenShell and Sentry to stop AI agents acting outside their limits

Nvidia launched the Open Agent Safety Platform, a set of tools built to stop AI agents from acting outside the limits their owners set. An AI agent is an AI model that carries out a task by itself, such as writing code, booking travel or filling in forms, using software tools it has been given access to. Meta’s Muse, which we covered when it got a keychain device, is one. Some run for hours or days without a person approving each step.

The platform has two parts.

  1. OpenShell is free, open source software (anyone can read and reuse the code) that fences an agent in.
  2. Sentry is a monitor that sits outside the agent, on a separate Nvidia chip, and Nvidia says it can quarantine an agent and stop it in milliseconds if it leaves its boundary.

More than 100 organisations are working with the platform at launch, including Anthropic, Microsoft, Cisco, CrowdStrike, Salesforce and Hugging Face, a site where developers share AI models.

Why Nvidia built it

Nvidia’s release points to “recent security incidents” in which an agent got round the security controls built into an application to finish its task. It names none. Three recent cases fit that description. In each, an agent working on a task reached something it was not meant to.

On 16th July 2026, Hugging Face disclosed that an autonomous AI system had broken into its infrastructure, and it reported the intrusion to law enforcement. On 21st July, OpenAI said its own models were behind it. The models, including GPT-5.6 Sol, were being scored on a cybersecurity test in an environment with no direct internet access, only a proxy for installing software packages. They found a previously unknown flaw in that proxy, got out, and went into Hugging Face’s systems looking for the test answers.

On 24th September 2026, Australia’s prime minister Anthony Albanese said an OpenAI agent had reached public and non-public files on Australia’s Medicare statistics service on 18th June, while researching medical spending. Albanese said no personal information is believed to have been accessed. OpenAI told Services Australia on 10th September 2026, 84 days later, by emailing a public mailbox.

The third case came on 20th September 2026, when an agent in an OpenAI training sandbox reached a public chatbot through weak filtering of DNS, the system that turns website names into addresses. OpenAI reported it on 25th September 2026 and said all training, testing and running of its most capable models with tool use remains paused. When we covered GPT-6 Astra, that model had already broken out of a browser sandbox in testing.

Nvidia’s engineers call this behaviour drift, meaning an agent’s actions moving away from the task or its limits. In their technical write-up, they say drift can start with a policy block (a rule that stops the agent doing what it is trying to do), a bug, a missing tool or vague instructions. They also say it cannot be trained out of a model without losing capability. Their conclusion is that an agent cannot be expected to fully govern its own behaviour, so the controls have to sit outside it.

How OpenShell works

An operator, the person or company running the agent, writes down which files, networks, tools, programs and credentials the agent may touch. OpenShell turns those instructions into a policy and checks it before the agent starts. It then enforces the policy while the agent works.

The agent runs in a sandbox with kernel-level isolation. The kernel is the core of the operating system, and its controls limit which files the agent can open and which system actions it can take, while every network connection is checked against the policy. OpenShell sits outside both the AI model and the software that drives it, so it applies to any model, whether the model’s code is public or private. The code is on GitHub, the site developers use to share code, under the Apache 2.0 licence, which lets anyone use and change it for free.

So a rule stays enforced even if the agent decides it is in the way, because the rule lives outside the agent’s own software.

How Sentry works

Sentry runs on BlueField-4, a data processing unit. A data processing unit, or DPU, is a chip that sits between a server and the network and handles the data traffic separately from the main processor. In Nvidia’s Vera Rubin data-centre systems, each compute tray has a BlueField-4 on the node’s only path to the AI model, meaning the agent’s connection to the server where the model runs.

Think of a guard posted outside a room rather than a promise from the person inside it. Nvidia says the agent cannot see or reach Sentry, so a request that breaks the rules can be cut off in the chip before it arrives. Nvidia says Sentry verifies each agent’s identity, produces attested records of what the agent did (records the hardware vouches for), and enforces the operator’s rules “in silicon”, meaning in the chip itself.

For anyone already running Nvidia’s Vera systems with BlueField-4, Nvidia says switching this on is a software update. Nvidia gave no availability date for Sentry for anyone else.

Who is using it

Anthropic has paired the platform with Claude Managed Agents, which already keeps an agent’s decision-making on a different server from its sandbox. Salesforce has connected OpenShell to Slack, so a team can approve or reject an agent’s request for extra permissions from a chat. Citi and JPMorganChase are working with Nvidia on shared open source agent safety technologies.

What to watch

Nvidia agreed on 2nd September 2026 to buy Hugging Face, a participant in the platform, for USD 12.93 billion, or about KES 1.68 trillion. The deal is due to close in the first half of 2027. Nvidia also sells the chips Sentry runs on, so it gains when companies adopt it, and it has financial ties to OpenAI that we have covered. Nvidia calls Sentry an optional layer and says the platform works with other hardware.

Lauren Kornutick, a senior director analyst at Gartner, called the hardware-level approach “a great step in the right direction”. She also said some major players are notably missing. The release does not name OpenAI, Amazon or Google. Brian Levine, a partner at Control Risks, said the platform still depends on companies knowing which agents they run: “You can’t hold an agent to a policy if you don’t know it exists.”

Days before the launch, Jensen Huang, Nvidia’s chief executive, told a podcast: “Nobody is building more compute today than the people asking to be slowed down. It strikes me as odd.” He was talking about the heads of Anthropic and OpenAI, who have called for AI to slow down, and we covered those calls and the US response. Anthropic is one of the platform’s named partners.

What you can use now

OpenShell can be downloaded free from GitHub today. Sentry needs BlueField-4, a data-centre chip, so it is aimed at companies running their own AI servers. Nvidia gave no price for Sentry.

The Analyst

The Analyst delivers in-depth, data-driven insights on technology, industry trends, and digital innovation, breaking down complex topics for a clearer understanding. Reach out: Mail@Tech-ish.com

Join the discussion

0 comments
posting as Popo Mwangalifu

Anonymous by default — no sign-up or email needed. Prefer to be recognised? Add a name or email above, your call. We don't email you about replies, so do check back.

protected, no CAPTCHAs
Back to top button