❝

“An agent cannot be expected to fully police its own behavior.”

Justin Boitano, VP of Enterprise AI, NVIDIA

Redeployed is a weekly newsletter that breaks down one important AI story at a time for product builders and engineering leaders. Every issue explores what the shift means for technology companies and how leaders can respond.

Most AI agents operate with a long list of instructions about what they are allowed to do. Don't access this system. Only use these tools. Ask before taking this action. Stay inside this environment.

That works until the agent doesn't cooperate.

This week, NVIDIA introduced the Open Agent Safety Platform, a new security architecture built around a simple assumption: companies should not expect an AI agent to enforce its own boundaries. The platform combines OpenShell, an open-source runtime that isolates agents and controls what they can access, with Sentry, a hardware-backed monitoring layer designed to detect and stop agents that move outside those boundaries.

The idea is familiar to anyone who works in cybersecurity. You do not secure a system by asking the software inside it to behave. You build controls around it. As agents gain access to production systems, customer data, internal applications, codebases, and business workflows, that principle becomes much more important.

Prompts Are Not Permission Systems

Many agent systems still rely partly on instructions to define what an agent should and should not do. Developers tell agents which tools they can use, what information they should avoid, and when they need human approval. Model providers and agent frameworks add their own safeguards.

Those instructions can shape behavior, but they are a weak place to enforce critical security boundaries. An autonomous agent may operate for hours or days, encounter unexpected situations, call multiple tools, or find routes through systems that developers never anticipated. NVIDIA describes this behavior as drift: an agent gradually moves away from its intended task or operating constraints while continuing to pursue its objective.

OpenShell puts the boundary outside the agent. Each agent runs in an isolated sandbox, while policies determine which files, credentials, network endpoints, APIs, models, and tools it can reach. Enforcement happens outside the agent's process, where a prompt cannot simply override it.

That gives companies a much harder boundary than "don't do that."

Treat Agents Like Untrusted Processes

Modern security architecture already assumes that software can fail, credentials can be compromised, and users can behave unexpectedly. Systems are designed to limit what happens when they do.

Agents deserve the same treatment.

If an agent needs access to a CRM to complete a sales workflow, it should receive the minimum access required for that task. If it needs to read from one system and write to another, those permissions should be explicit. Access to credentials, files, APIs, and external networks can be restricted independently of whatever the model decides to do.

NVIDIA's OpenShell uses a deny-by-default approach. Agents begin without access and receive permissions through policies enforced at the runtime level. Their actions and policy decisions can also be recorded for auditing.

Sentry adds another layer for higher-risk environments. It runs independently on NVIDIA's BlueField-4 hardware and monitors agent activity outside the host system. If an agent crosses a defined boundary, Sentry is designed to quarantine it in milliseconds.

The architecture looks much closer to traditional zero-trust security than to prompt-based agent guardrails.

The Risk Changes When Agents Can Act

Agents are becoming more capable at the same time that companies are giving them more authority. An assistant that drafts an email creates limited operational risk. An agent that can access production databases, execute code, modify infrastructure, send payments, or update customer records has a very different security profile.

During an internal OpenAI cybersecurity evaluation this summer, models operating with reduced safeguards circumvented controls designed to isolate them from the internet, exploited vulnerabilities in the evaluation environment, and eventually accessed Hugging Face systems. The setup was deliberately different from a normal production deployment, but the incident showed how an agent pursuing an objective can discover paths its developers did not anticipate.

This becomes harder as agent systems grow. Consider a multi-agent system where one agent can read sensitive information and another can communicate externally. Each permission set might be reasonable on its own, while the combined system creates a path the organization never intended.

Agent security therefore has to account for what the entire system can do, not simply what each agent was instructed to do.

Authorization Belongs Outside the Model

Separating intelligence from authorization is one of the most useful architectural ideas in NVIDIA's announcement. The model can decide how to approach a task, while the infrastructure determines which actions are actually possible.

Once authorization lives outside the model, changing the model no longer requires changing the organization's fundamental security boundary. OpenShell is designed to work with open and closed models and across different agent frameworks, giving organizations a policy layer that can remain consistent even as the intelligence underneath it changes.

The model remains free to reason within those boundaries. It simply cannot grant itself more authority.

This issue of Redeployed is brought to you by Tecla Labs: As agents gain more authority inside business systems, security has to extend beyond the model itself. Tecla Labs helps companies build AI systems around real workflows, including the integrations, permissions, infrastructure, and controls needed to make agents useful in production. Start with a free AI Assessment to identify where AI can create value without giving it more access than the workflow requires.

Containment Does Not Solve Everything

External controls make agent systems safer, but they do not eliminate the need to understand what agents are doing. An agent can stay entirely within its permissions and still make a bad decision. It can modify the wrong record, misunderstand a request, or take an action that is technically allowed but operationally harmful.

Companies still need evaluations, monitoring, approval thresholds, and ways to investigate agent behavior. Sandboxing limits the potential damage when something goes wrong, but it cannot determine whether every permitted action is a good one.

There is also an operational cost. Fine-grained permissions become difficult to manage when a company moves from a handful of agents to hundreds or thousands. Teams will need ways to define roles, update policies, understand combined permissions, and audit what agents have actually done.

That starts to look like identity and access management for a new category of worker.

Agent Permissions Become Infrastructure

Companies deploying agents at scale will eventually need to answer the same kinds of questions they already answer for employees and software services. What can this agent access? What actions can it take? Who can change those permissions? What happens when it exceeds them?

Those decisions need infrastructure that security and engineering teams can inspect, test, audit, and enforce consistently. They cannot live entirely inside prompts scattered across different applications.

NVIDIA's platform is one attempt to build that layer, and more approaches are likely to follow as autonomous systems move deeper into enterprise operations. For engineering leaders, this means agent architecture increasingly includes an explicit security boundary between the model and the environment where it works.

Agent Safety Is Becoming Systems Engineering

AI companies will continue improving alignment, instruction following, and model safety. More capable agents will still encounter situations their developers did not predict, especially as they operate for longer periods and interact with more systems.

Production software already has an answer for that problem: assume components can fail and limit the damage when they do. Agent systems are beginning to inherit the same architecture.

The safest enterprise agent may eventually be the one that can reason freely inside a boundary it has absolutely no power to change.

Connect With Other Technology Leaders

If you want to exchange practical ideas with senior technology and product leaders navigating the AI era, check out the upcoming GILD Forums. They bring together experienced operators for peer discussions around the technology and business decisions they are working through right now.

More to come…

– Gino Ferrand, Founder @ Tecla Labs