IA y aprendizaje automático
Securing AI Agents With Runtime Boundaries: What NVIDIA OpenShell Adds
NVIDIA's new agent safety reference design puts policy enforcement outside the agent. Here is how enterprise teams can evaluate the boundary, its limits and the review steps that still matter.
Este artículo está disponible actualmente solo en inglés.
An agent that reads files, calls APIs and runs code needs more than a careful system prompt. If it can reach a production database or send data to an arbitrary host, a mistaken plan can become an actual write or disclosure. The practical security question is which actions the environment will permit when the agent tries them.
On September 28, 2026, NVIDIA described its Open Agent Safety Platform, a reference design combining the open-source OpenShell runtime with an optional Sentry layer on BlueField hardware. Its OpenShell 0.1.0 technical walkthrough details the software controls. The announcement is a useful occasion to examine a wider design decision: put the enforcement boundary outside the agent workload.
What the runtime actually controls
OpenShell runs an agent in a sandbox. NVIDIA says its kernel controls limit file and process access, while a supervisor outside the workload checks outbound requests against policy. A gateway manages sandbox lifecycles and policies. For configured HTTP, GraphQL and Model Context Protocol traffic, the supervisor can allow a read while blocking a write to the same service. NVIDIA also describes keeping real credentials outside the workload and attaching them only to approved requests. See the technical walkthrough for the component boundaries and an example policy.
That separation matters because an instruction such as "only read the repository" is still text for the agent to interpret. A service policy can reject a write request even if the agent generates code, launches a child process or changes its plan. It also gives the security team a place to record an allow or deny decision. This does not prove that an allowed answer is accurate; it constrains the actions available to produce it.
NVIDIA's larger reference design adds Sentry as an independent monitoring and enforcement layer on BlueField-4 systems. The company describes that hardware layer as optional, and says the platform is compatible with other hardware systems. A team evaluating the software runtime should therefore distinguish the controls it can test in its own environment from claims tied to a specific hardware deployment. NVIDIA's architecture description makes that distinction.
A concrete enterprise test
Consider an internal agent that summarizes an incident, reads a ticket and prepares a status update. It needs read access to the ticket and relevant documents. It may need to create a draft. It should not gain permission to close the incident, change customer records or send the update simply because those actions use the same integration.
Start with a task-specific inventory: files, API hosts, methods, paths and credentials. Give the agent the smallest set that completes the work. Keep the send or close step with a human reviewer. Then run representative cases, including an instruction inside a ticket that asks the agent to export the incident log to an external address. The test should show the attempted action, the policy decision and the resulting state. A refusal in the agent's prose is less useful evidence than a blocked request in the runtime log.
This is a general evaluation method, not a claim that OpenShell has passed the scenario in your environment. It also applies to the access and audit questions in our guide to governing Gemini Enterprise agents. The products have different control planes; the shared question is whether the permitted action matches the operator's intent.
Four questions before adopting a runtime
- Where is enforcement complete? Check direct tool calls, shell commands, generated scripts, child processes and remote MCP servers. A policy is only as strong as the paths it actually mediates.
- How are exceptions reviewed? An agent may need a new endpoint midway through a task. Record who approved the change, how long it lasts and whether the new permission expands access for other agents.
- What happens to credentials and logs? Verify that secrets remain outside the agent workload, that service-side permissions still apply, and that audit records contain enough context to reconstruct a decision without exposing sensitive content unnecessarily.
- What does the boundary cost? Measure task completion, blocked legitimate requests, latency and operational burden on your own workloads. A restrictive policy that teams routinely widen is a poor production control.
NVIDIA says OpenShell includes a policy prover that checks modeled permissions against a defined boundary. That is valuable for reviewing policy changes, but its proof concerns what the policy permits. It cannot establish that a summary is correct, that an authorized write is wise, or that a human approval step can be removed. Those still require task evaluation and workflow design.
For a first pilot, choose one agent with a narrow task and a reversible output. Document its access, try to cross that boundary deliberately, and review both the denied requests and the work it completed. Expand authority only when those records support the change.
Artículos relacionados
Google Cloud API Gateway MCP: What to Check Before Exposing REST APIs
Google Cloud can now expose REST operations as MCP tools through API Gateway. Before connecting an agent, review tool discovery, operation scope, authentication and the preview limits.
IA y aprendizaje automáticoEl contenido de IA a escala tiende al promedio. Por qué ocurre y qué hacer al respecto.
Sobre la arquitectura de señales, el diseño de habilidades de ADK y la ingeniería detrás del contenido que resiste la compresión.
¿Trabaja en algo similar?
Sin discursos de venta: solo una conversación práctica con el equipo que construye y opera estos sistemas en producción.
Inicie una conversación