AI Agent Safety Needs a Runtime Boundary, Not Another Prompt
AI agents are acquiring the ability to read files, call APIs, use credentials and change production systems. That makes a familiar safety instruction—“only do what the user asked”—far too weak. A prompt is guidance interpreted by the same system whose behaviour it is meant to constrain. A runtime boundary is enforcement.
That distinction moved into the spotlight on 28 September, when NVIDIA announced its Open Agent Safety Platform. The design combines OpenShell, an open-source runtime that restricts and traces an agent’s actions, with Sentry, a reference design for an independent watchdog outside the agent’s main compute path. The product names will evolve. The architecture is the important part.
Agent safety is becoming a systems problem
Traditional applications expose a defined set of routes and execute code written in advance. An agent chooses a sequence of actions at runtime. It can encounter an unexpected page, a poisoned document, an unavailable tool or an ambiguous instruction, then adapt its plan. This flexibility is useful, but it also means the application cannot enumerate every path during design.
The practical response is not to predict every decision. It is to limit the consequences of any decision. OpenShell’s published policy schema separates controls across the filesystem, processes and network access. Filesystem paths can be read-only or writable; processes run under a restricted identity; and outbound connections can be limited by destination, port and calling binary.
This is least privilege applied to an actor whose next step is probabilistic. The model may decide what to try, but infrastructure decides what is possible.
Two boundaries are better than one
The most useful idea in NVIDIA’s announcement is not a new sandbox. It is separation of duties. OpenShell governs the environment in which the agent works. Sentry is designed to monitor from outside that environment and quarantine suspicious behaviour independently. An agent should not be able to rewrite the policy that confines it, suppress the logs that reveal its actions or disable the mechanism that stops it.
The Associated Press explainer also highlights the limit of this approach: containment cannot make a model truthful, prevent every mistake or repair a poorly written permission policy. It reduces blast radius. It does not eliminate the need for product judgement, security testing or human accountability.
A practical control stack for production agents
Founders and technical leaders do not need to wait for a particular hardware platform. They can apply the same pattern to today’s agent workloads.
- Inventory actions, not features. List every file operation, network destination, database query, message, payment, deployment and administrative change the agent can attempt. Classify each as read-only, reversible or irreversible.
- Start from deny. Give each agent a small filesystem view, explicit outbound destinations and the minimum process capabilities needed for one job. Avoid a shared “agent service account” that accumulates permissions across workflows.
- Bind credentials late. Keep secrets out of prompts, memory and the agent-visible filesystem. Inject short-lived credentials only after the runtime has admitted the destination and action. Separate read and write identities.
- Put approval at the irreversible edge. Require a human or deterministic policy check immediately before sending money, publishing content, deleting records, modifying access or deploying code—not merely at the beginning of a long task.
- Keep evidence outside the agent. Send policy decisions, tool calls, network events and approval records to storage the agent cannot alter. Give every run a trace ID so an operator can reconstruct what happened.
- Test the boundary as a product. Try prompt injection, credential theft, path traversal, unapproved egress, excessive loops and policy tampering. Measure both escapes and false blocks; a boundary that routinely prevents legitimate work will be bypassed by its operators.
Do not treat the sandbox as trusted forever
Security infrastructure is software, and it can contain vulnerabilities. NVIDIA’s own August OpenShell security bulletin documented critical issues and shipped fixes. That is a useful reminder to pin versions, track security advisories, test upgrades and layer controls. A sandbox should sit behind restricted cloud identities, segmented networks and monitored data stores—not replace them.
The current OpenShell security guidance also distinguishes controls that are fixed when a sandbox starts from those that can change during a run. Teams should be conservative with dynamic permission expansion. If an agent requests broader access, the request should be specific, time-bound, visible to an operator and automatically revoked.
A 30-day adoption path
In week one, run one valuable agent in observation mode and capture every resource it touches. In week two, convert that trace into a default-deny policy and split read from write access. In week three, add approvals for irreversible actions and route audit evidence to independent storage. In week four, run adversarial scenarios and a containment drill: can the team revoke credentials, stop the workload and explain the incident within minutes?
Success is not “the agent never made a mistake.” A better operational target is that an unexpected decision cannot silently become an unlimited action.
The Qomra Tech view
Agent projects should now be reviewed like small distributed systems, not clever chat interfaces. The delivery plan needs an action model, identity design, policy boundary, approval path, audit trail and shutdown mechanism alongside prompts and evaluations. Product owners should define which outcomes require autonomy; security teams should define the maximum consequence of failure; platform teams should make those limits enforceable.
NVIDIA’s launch is a useful signal that the market is moving from behavioural promises to runtime control. For teams shipping agents today, the next step is simple: choose one production workflow and prove that its permissions, evidence and stop mechanism remain outside the agent’s control.