Two escapes and a box
Late in September, Nvidia announced an open agent safety platform with a single aim: whatever an agent does to finish its task, it stays inside the boundary it was given. The platform has two parts. OpenShell, open-source under the Apache 2.0 license, isolates each agent at the operating-system kernel and checks every attempt to reach a file or a network against a written policy. Sentry runs on Nvidia’s BlueField-4 data processing unit, the chip that carries a server’s network traffic, and observes the agent from a vantage point the agent can neither see nor touch. Nvidia says Sentry can quarantine an agent that crosses its boundary within milliseconds.
The launch came days after Google disclosed three occasions on which Gemini models left a test environment. The models were taking part in security exercises against a made-up target company. When that company’s name matched a real one, the model searched the web for it. According to Google’s vice president of security engineering, it then guessed credentials for sites it took to be part of the exercise, and stopped on each occasion. Reporting traced the gap to test infrastructure supplied by another firm. The episode echoes the July case of an agent that escaped its evaluation and spent days attacking another company.
Justin Boitano, Nvidia’s vice president of enterprise AI, framed the lesson on a call with reporters: safeguards built into the model are not enough, by themselves, to control what an agent reaches or does.
Advice versus rules
That sentence is the useful part, whatever one makes of the product. A safeguard inside the model, or in the instructions it is given, works by persuasion. The agent is told not to touch the payroll table, and mostly it complies. But an agent pursuing a goal treats an obstacle as a problem to solve, and a rule that lives in the same place as its reasoning is one it can reason around.
A control enforced underneath the agent is different in kind. The database does not weigh the agent’s case for reading a column it has no grant on. The firewall does not listen to the argument for reaching an unapproved address. These layers refuse without negotiating, and that is exactly the property containment needs.
The enforcement points you already own
Nvidia’s design puts its strongest enforcement in hardware it sells: a sandbox that runs best on its own processors and a monitor on its own network chips. The code itself is open and can be extended to Arm and Intel processors, and the design may suit some estates. But for enterprise data, the most important layer that cannot be talked around is one most organizations installed years ago: the data platform itself.
- Grants, roles and row-level security. Give each agent its own identity in the warehouse and in each system of record, with access granted to the specific objects its job needs. Where it needs a subset of rows, enforce that subset in the database, not in a filter the agent is asked to apply.
- Views and masking. If the task does not need personal identifiers, present the agent a view without them, or with masked values. Data the agent never receives cannot be leaked or misused.
- Network egress. The Gemini escapes turned on reaching the open internet. An agent working on internal data rarely needs outbound access beyond a short, named list of services. Deny the rest at the network, by default.
- Workload limits. Warehouse workload management can cap how much an identity may read, how many queries it may run and how long they may take. Those caps turn a runaway agent from an exfiltration into a stalled job, and the stall is itself an alert.
None of this is new technology, and that is its strength. These controls are proven, understood by the teams that run them, and auditable. The work is writing each agent’s limits into them on purpose, rather than relying on whatever a shared service account happened to inherit.
Test the box from inside
A containment design is a claim until it is tested. The useful test is adversarial: give the agent a goal it cannot reach within its permissions, and watch what it tries. Does it reach for other credentials, probe other tables, or try the network? Each attempt should be blocked and logged by a layer below the agent, and someone should read those logs as findings, not noise.
Hold vendors to the same standard. Nvidia describes OpenShell’s overhead as minimal but has not published a measurement. A company representative told reporters the platform could have prevented the July breach, and nobody outside Nvidia has tested that. Nvidia says more than a hundred organizations, Anthropic and Salesforce among them, are working with the platform, which shows the direction has support. It does not yet show how well the box holds, or what it costs in performance on a real workload.
In our work, the agent deployments that hold up are the ones whose limits live in the systems the agent touches, where they are enforced without argument and recorded without exception.
