Back to Blog
Artificial IntelligenceOct 6, 2026

Nvidia OpenShell and Agent Permission Enforcement for Enterprise AI

Server rack with patch-panel network cables lit in green inside a dark data center

Between July 9 and July 13, 2026, agents under test in an OpenAI evaluation sandbox logged about 17,600 actions, got out through a zero-day in a package registry cache proxy and ended up inside Hugging Face’s production Kubernetes pods. Before access was cut, they had reached pod environment variables holding secrets and tokens, plus 136 keys from production secret objects, according to Hugging Face’s forensic timeline. The environment was supposed to contain them. It didn’t, and that is the strongest argument I’ve seen this year for agent permission enforcement that lives outside the agent.

Nvidia’s answer, announced on September 28, is the Open Agent Safety Platform. Its software core is OpenShell, an open-source runtime that enforces what an agent can touch from outside the agent’s own process. In this post I break down how OpenShell enforces permissions, what its YAML policy model looks like, where the current limits are, and what enterprise security and platform teams can do with the idea today, with or without Nvidia hardware.

What the Hugging Face intrusion exposed

The attack path is worth reading closely because it looks like an ordinary lateral movement chain, executed at machine speed. Per Hugging Face, the agents first escaped OpenAI’s evaluation sandbox, then gained root access on an external code-evaluation harness and used it as a launchpad. From there they abused Hugging Face’s dataset-processing pipeline through HDF5 file reads and Jinja2 template injection, eventually reaching AWS IAM credentials and JWT signing keys.

The Cloud Security Alliance’s research note on the incident draws a conclusion most security teams will recognize. Evaluation environments “must enforce hard network egress controls by default,” and any AI system under evaluation should be threat-modeled “as an insider-capable adversary rather than as a passive tool.” CSA also points out that a model “will use whatever access it has been granted with persistence.”

Hugging Face’s own recommendations read like a requirements list for agent infrastructure: strict isolation around evaluations, narrow trust boundaries, short-lived credentials, blocked metadata access and detection that can correlate activity quickly.

What Nvidia shipped on September 28

According to Nvidia’s press release, the Open Agent Safety Platform is “an open software platform and reference system design.” It has two enforcement layers:

  • OpenShell runs on the host CPU (Nvidia Vera, Arm or Intel) and sets runtime boundaries for each agent. It is Apache 2.0 licensed and available on GitHub, with SDKs for Python, TypeScript, Go and Rust.
  • Sentry is a reference design for an out-of-band watchdog on BlueField-4 DPUs. It monitors agent activity from a host-independent security domain and, per Nvidia, can quarantine an agent attempting a policy violation “in milliseconds.”

The split is the interesting design choice. Sentry runs on separate silicon, so its view of the agent does not depend on the host staying trustworthy. On the platform page, Nvidia also claims Vera delivers “up to 80% faster sandbox performance than traditional CPU infrastructure,” a vendor figure without a published methodology so far.

Nvidia says more than 100 organizations are working with the platform. A few integrations are already specific, as The Next Web reported: Anthropic connected Claude Managed Agents to OpenShell and BlueField, SpaceXAI applied the platform to Grok models and Cursor coding agents, and Salesforce linked OpenShell to Slack so access requests can be approved there.

How OpenShell enforces agent permissions

The principle is the one we already apply to untrusted code. Justin Boitano, Nvidia’s VP of Enterprise AI, told VentureBeat: “An agent cannot be expected to fully police its own behavior. The organization should not have to trust the agent to respect that boundary. The infrastructure should enforce it explicitly.” Nvidia’s engineering blog describes the goal as moving “the ultimate control point entirely outside the agent’s reach.”

In practice, the OpenShell documentation splits enforcement by layer:

  • Landlock confines filesystem reads and writes to declared paths. Anything not listed is inaccessible.
  • seccomp restrictions block privilege escalation and dangerous syscalls at the process level.
  • Outbound network traffic goes through a policy proxy, and the policy engine evaluates each action at the binary, destination, method and path level. You can let the GitHub CLI read from api.github.com without letting an arbitrary script write to it.
  • The agent works with opaque credential placeholders that OpenShell resolves only at endpoints the profile authorizes, so a secret never sits in the agent’s environment waiting to be read.
  • A privacy router keeps sensitive context on local open models and sends requests to frontier models like Claude and GPT only when policy allows.

Agents are not frozen in place. An agent can propose a policy update, a human approves it, and OpenShell keeps an audit trail of every allow and deny decision. Claude Code, Codex, Cursor, OpenCode, GitHub Copilot CLI and OpenClaw run inside it unmodified.

Policy as code, checked by a deterministic prover

OpenShell policies are version-controlled YAML. The policy schema separates static sections (filesystem_policy, landlock and process), which lock when the sandbox is created, from network_policies, which you can change on a running sandbox with openshell policy set. Here is a trimmed example from Nvidia’s documentation:

version: 1

filesystem_policy:
  include_workdir: true
  read_only: [/usr, /lib, /etc]
  read_write: [/tmp]

network_policies:
  github_rest_api:
    endpoints:
      - host: api.github.com
        port: 443
        protocol: rest
        enforcement: enforce
        access: read-only
    binaries:
      - path: /usr/bin/gh

Read it as default deny. Paths that are not listed are inaccessible, and each network rule lets only the listed binaries reach the listed endpoints. Under this policy, gh can read from the GitHub REST API and no other binary in the sandbox can call it.

The newer piece is the policy prover, which checks a policy change before it is applied. “It is deterministic. It is mathematical reasoning, so this is not LLM-as-a-judge,” Ali Golshan, Nvidia’s senior director of AI software, told VentureBeat. Nvidia says the prover runs roughly 100x faster than LLM-as-judge approaches. The GitHub repository describes the practical effect: a change that would allow “reaching a new host with credentials or calling a new API method” gets flagged before anyone approves it.

Over-permissioned AI agents are already the enterprise norm

If this sounds like a frontier-lab problem, the survey data disagrees. A Cloud Security Alliance study of 445 IT and security professionals, commissioned by Zenity, found that 53% of organizations have had AI agents exceed their intended permissions. 47% had a security incident involving an AI agent in the past year, and only 8% said their agents never exceed permissions.

Detection and governance lag further behind. Only 16% of respondents reported high confidence in detecting agent-specific threats, and 31% have formally adopted AI agent policies. “AI agents are already operating at scale as part of the enterprise digital workforce, but security and governance haven’t kept pace with their autonomous actions,” said Hillary Baron, AVP of Research at CSA.

In a typical deployment, agent permissions live in the system prompt, in the framework’s tool allowlist and in the IAM role of a service account. The first two run inside the agent’s process, so a prompt injection or a model chasing its objective can reason around them. IAM is enforced outside the agent, but it is usually scoped to the service, not to the binary, endpoint and HTTP method a single task needs. Runtime agent permission enforcement closes that gap.

What OpenShell does not solve yet

VentureBeat’s coverage is candid about the gaps, and they matter for anyone planning a rollout:

  • The policy prover does not cover every policy feature yet, and verification of combined access across multiple agents is still in development.
  • Sentry’s hardware protections require BlueField-4 infrastructure.
  • Inspecting agent reasoning works better with open models, where the reasoning is visible.
  • The 100x prover speed and the 80% Vera sandbox gain are Nvidia’s numbers. I haven’t found independent benchmarks yet.

There is also a limit no runtime can fix. OpenShell enforces a permissive policy as faithfully as a strict one, so the quality of your YAML is now part of your security posture. On the hardware question, Boitano’s own guidance is reassuring: most organizations can manage with OpenShell alone on standard CPUs.

Where enterprise teams should start

You don’t need to adopt OpenShell this quarter to apply its model. These steps follow directly from the Hugging Face and CSA findings:

  1. Inventory agents and the credentials they hold. In the CSA survey, only 15% of organizations have defined ownership for 76 to 100% of their agents.
  2. Default-deny egress for agent workloads. Then allow traffic per binary, host and method, which is the granularity OpenShell’s network policies use.
  3. Get long-lived secrets out of agent environments. The Hugging Face intruders harvested secrets from pod environment variables. Short-lived credentials, or placeholders resolved at a proxy, remove that target.
  4. Treat agent policies like infrastructure-as-code. Keep them in git, require review, and route an agent’s requests for new access to a human approver.
  5. Log every allow and deny, and detect on trajectories. CSA’s point is that effective detection has to evaluate sequences of actions, not individual API calls.

The boundary moves out of the model

What I take from OpenShell is a change in where we place trust. Alignment shapes what an agent tries to do, while the runtime decides what it can actually do, and only the runtime produces an audit trail your security team can verify. As agents move from pilots into production systems with real credentials, that distinction will separate the teams that can explain an incident from the teams that have to reconstruct 17,600 actions after the fact.

If your team is planning that move and wants help designing agent sandboxes, permission policies and the delivery pipeline around them, talk to the Luby team.