Nvidia has unveiled the Open Agent Safety Platform to restrict unauthorised actions by autonomous AI agents. The platform sets access limits and monitors the agents’ activities whilst they are carrying out tasks. The company has made part of the code available to its partner ecosystem and is working with Anthropic to integrate cloud-based agents with the OpenShell system.
Briefly about the main points
- OpenShell restricts the agent’s capabilities on central processing units.
- Sentry analyses current network activity in real time.
- The OpenShell code is available to partners under the Apache 2.0 licence.
- Nvidia and Anthropic are preparing to integrate cloud agents.
- Nvidia links the platform to the risks identified during the attack on Hugging Face.
OpenShell sets the limits; Sentry monitors execution
OpenShell runs on central processing units and is designed to limit the agent’s potential capabilities. Its architecture incorporates isolated sandbox environments, validation of requests to external services and centralised management of access policies.
The system may allow an agent to read data via the API, but prohibit writing. Account details, as described Nvidia, are not passed directly to the agent: OpenShell only substitutes them after verifying access to the agreed endpoint.
Sentry runs on network chips and monitors agent activity in real time. In its full configuration, it is integrated with BlueField-4, creating a control loop for network requests and policies that is separate from the host.
Open source and partner integrations
Nvidia has published part of the code as a base architecture for its partners. The OpenShell repository is available under the Apache-2.0 licence, which allows the technology to be adapted and integrated into various deployment environments.
Among the partners named by the company are — Microsoft, Oracle, Cisco, Dell, Arm and Intel. Nvidia is also collaborating with Anthropic on the integration of cloud agents with OpenShell; according to Nvidia, the Claude Managed Agents cycle will run on a separate server from the sandbox where the agent performs its tasks.
The context of the attack on Hugging Face and the limits of security
Vice-President of Nvidia Boitano stated that the platform could have prevented the recent incident involving Hugging Face, when the service came under attack from more than 17,000 AI agents. In its technical analysis of the July campaign, Hugging Face described how the agent escaped the evaluation environment and subsequently exploited a chain of vulnerabilities and configuration errors.
Controls over permissions, access to confidential information, network access and environment isolation address some of the risks of this nature. At the same time, the claim that a specific incident can be prevented is Nvidia’s own assessment: there are no independent public tests that would quantitatively confirm the effectiveness of OpenShell and Sentry against real-world compromises. The practical outcome will depend on the policies in place and the infrastructure configuration.







