Enforce narrow permissions outside the model and require approval at consequential action boundaries.
Translate intent into capabilities
An instruction such as analyse treasury balances does not imply permission to transfer assets. Break the workflow into capabilities: read data, prepare a proposal, request a signature and submit an authorised action. OAuth scopes illustrate how an access system can describe limited authority, but their exact meaning depends on the service. A prompt telling an agent to behave carefully is not a substitute for permissions enforced by the tools and accounts it uses.
Design a hypothetical reporting agent
Suppose an agent prepares a weekly treasury report. It needs read-only account information and a place to save a draft, not a private key or trading credential. If the user later requests a payment proposal, the agent can assemble the destination and amount for review without gaining general transfer authority. A separate approval can bind the exact network, asset, destination, amount and expiry before an execution component is allowed to act.
Treat retrieved content as data
An external webpage, document or token description can contain instructions aimed at the agent. Those instructions should not acquire authority merely because the agent encountered them during research. Keep the user's task and the tool's permission checks separate from the material being analysed. For example, an invoice attachment asking an assistant to change the payment destination should trigger verification through an established channel, not automatic acceptance as a new instruction.
Test refusal and recovery paths
Ask what happens when a requested action exceeds a limit, when credentials expire and when a user withdraws permission. Test whether the system actually blocks disallowed destinations instead of merely logging a warning. Keep auditable records without exposing secrets, and provide an accessible stop mechanism. Limits, approvals and simulation reduce particular failure modes but do not guarantee correct behaviour. The strongest design makes the agent's available actions match the task even when its reasoning or input data is wrong.
Sources & context
Sources checked 7 October 2026. This is explanatory coverage, not personalised investment advice. Our corrections policy.