Every organisation adopting agents eventually writes a document stating what they may and may not do. That document prevents nothing. What prevents things is a policy engine evaluating each tool call before it leaves.
Where the decision point has to sit
There are three possible places, and only one works.
In the system prompt — “you must not delete data”. That is instruction, not control. It competes on equal footing with any text entering the context afterwards.
At the tool server — it sees the call in isolation, unaware of which task is running or how much has been consumed. It can deny by permission, not by context.
At the orchestration layer or gateway — everything lives here: agent identity, delegator, task, execution history, provenance of content read. It is the only point with enough information to decide.
A prompt instruction is a request. A policy evaluated outside the model is a control. The difference shows up exactly when someone tries to subvert the system.
Which rules are worth writing first
- Delegation ceiling: deny if effective scope exceeds what the delegating human could do.
- Context contamination: deny external writes if external content entered this task's context.
- Budget: deny when the call count exceeds the task limit.
- Irreversible verbs: require approval for delete, transfer, publish and grant.
- Destination: deny arguments pointing at private or link-local ranges.
Five rules cover most of the risk described across the last several pieces. And they are rules you write in Rego in an afternoon, not in a quarter.
Why reuse OPA
Because the organisation probably already has it. If you run admission control on Kubernetes or policy in your pipelines, the engine, the review process and the policy deployment path already exist. Extending to the tool-call path adds a rule package, it does not adopt a technology.
And there is a credibility gain: agent policy becomes versioned, reviewed by pull request and testable — three things a PDF will never be.