AI agent security becomes a business decision when an assistant can change records, send messages or trigger work in another application. A convincing demonstration tells you whether the agent can complete a task. It does not establish whose information it can access, which actions it may take or how your team recovers when something goes wrong.
Secure an AI agent by limiting its data access and available actions, enforcing authorisation in the connected applications, and requiring informed approval for consequential changes. Test those boundaries with realistic failures and hostile inputs before expanding access. A better prompt alone cannot provide that assurance.
For UK teams commissioning an integration, the useful question is what the agent is allowed to do when its judgement is wrong. This guide proposes a practical scope for a controlled pilot, the evidence to request from a supplier and the costs that belong in the business case. It does not promise an attack-proof system or describe a named customer’s deployment.
Define AI agent security around a business task
Start with a narrow workflow such as preparing a support reply or proposing a CRM update. Write down the source records, the intended user and the final destination. Separate reading, drafting and executing: access to a customer record need not imply permission to export it, and drafting a response need not imply permission to send it.
Agree what remains outside the pilot. Refunds, account permission changes and bulk exports are examples of actions that deserve separate decisions. A supplier should be able to show where those restrictions are enforced. A sentence in a system prompt is useful guidance, but it is not evidence that a connected API will reject an unauthorised request.
Give the workflow a named owner who can decide whether an exception is acceptable. Without that person, technical teams can quietly inherit business decisions about customer communications or record changes. A useful brief describes an authorised outcome and its boundaries, rather than asking for an agent that can use every available tool.
Treat incoming content as evidence, not instructions
OWASP’s prompt-injection guidance distinguishes direct manipulation through a prompt from indirect manipulation through external material such as files and websites. It also warns that retrieval and fine-tuning do not fully remove this vulnerability. Connecting an agent to company documents therefore does not make every retrieved sentence trustworthy.
Consider a hypothetical support email that asks an assistant to copy an unrelated customer’s records into its reply. The email is content to assess, not permission to widen access. The important test is whether the surrounding application blocks that attempted disclosure even if the model follows the instruction. Apply the same reasoning to attachments, search results and tool responses.
Label external content and validate proposed outputs, but do not sell these measures as a complete defence. Our recommendation is to design the integration so a misleading answer has limited authority. A model that suggests an inappropriate action should meet an independent access check before anything happens in the business system.
Make permissions follow the user and the record
The connected application should check both who is asking and which record the request concerns. An agent acting for a support colleague should not automatically inherit an administrator’s access. Define whether a service identity is necessary, what it can read or change, and how the integration preserves the initiating user’s scope.
OWASP’s excessive-agency guidance recommends minimal functionality and permissions, user-context execution and downstream authorisation. It identifies excessive functionality, permissions and autonomy as separate causes of damaging actions. A read-only connector and a narrowly scoped account address different parts of that problem.
Test with accounts that have different responsibilities. Attempt to read another team’s material, alter a protected field and request an export outside the approved scope. Record the actual refusal from the application. A successful happy-path demonstration under an administrator account cannot show that ordinary users are isolated from information they should not see.
Put an action gateway between suggestion and execution
An action gateway is application code that checks a proposed operation before calling the destination system. For a CRM pilot, it might accept only approved record identifiers and permitted fields. Reject unknown operations and unsupported values. Keep credentials in the connector’s controlled environment rather than placing secrets in model-visible instructions.
Prefer specific operations such as proposing a note over a general tool that can run arbitrary commands or contact arbitrary destinations. The smaller interface is easier to inspect and test. It also gives the business a concrete list of capabilities to approve when the pilot expands.
| Capability | Initial pilot boundary | Evidence to request |
|---|---|---|
| Read a record | User-authorised records only | Cross-account access is refused |
| Draft a message | No automatic sending | Draft remains reviewable |
| Update a field | Approved fields and values | Invalid changes are rejected |
| Export information | Disabled unless separately scoped | Unapproved destinations are blocked |
This matrix is a proposed starting point, not a universal policy. Adjust the boundaries to your workflow and error consequences. Include ordinary application controls: an agent integration still needs reliable authentication, validated inputs and a recoverable connection to the destination.
Follow the proposal through the control boundary
The diagram separates a model’s suggestion from the application’s decision to execute it. Content can influence the proposed operation, but the gateway checks identity, record scope and permitted changes independently. A consequential action also waits for approval of the final operation. Requests outside policy take a refusal path instead of quietly acquiring more authority.
Make approval a real decision
An approval screen should show the action, destination and material change being authorised. For an outgoing message, show its recipient and final text. For a record update, show the existing value and proposed replacement. Asking someone to approve an unexplained instruction transfers responsibility without giving them the information needed to exercise it.
Bind approval to the operation actually executed. If the target, content or relevant record state changes, require a fresh decision under the agreed policy. Otherwise a person can approve one version while the software executes another. Include expiry and cancellation behaviour in the acceptance tests rather than leaving them as interface details.
Do not make every trivial operation an approval task. That produces a queue people learn to dismiss. Agree which actions need review, which can run within an established policy and which remain unavailable. Measure whether reviewers can understand and complete the work, including busy periods and the absence of the usual approver.
Test failure paths before granting more access
Build an evaluation set from representative, anonymised work. Include missing fields, contradictory records, unsupported attachments and attempts to redirect the task. Keep some examples separate from the material used to adjust the system. The purpose is to test boundaries and usable outcomes, not reward a demonstration that memorised its examples.
Exercise the whole workflow. Interrupt the destination connection, repeat a request after a timeout, revoke a user’s access and change a record while approval is pending. Check whether the agent reports a truthful status and whether staff can resume work without duplicate updates. A polite refusal is insufficient if a tool already performed the forbidden action.
Ask for evidence that connects each test to an expected result, actual application behaviour and a remediation owner. Report unresolved limitations clearly. Repeat relevant tests after changes to prompts, models, tools or permissions. A pilot should establish the conditions under which the workflow can operate, including the conditions that require it to stop.
Ask for an acceptance record you can inspect
A useful acceptance record connects an attempted action to a visible result in the destination, rather than relying on the agent’s explanation. For example, a test may deliberately submit an update for a record outside the user’s scope. The expected outcome is a refusal with no destination change. Preserve the relevant identifiers so somebody other than the demonstrator can check that outcome.
| Test condition | Expected evidence | Reason to stop expansion |
|---|---|---|
| Unauthorised record | Access refusal and no record change | The connector bypasses the user’s scope |
| Changed draft after approval | New approval required before execution | A different action uses an earlier approval |
| Timeout after destination accepts an update | Reconciled result without a duplicate change | A retry creates additional work |
| Stop requested with queued actions | Pending actions remain unexecuted | The worker continues after the stop |
Agree how the team will verify these results before running the pilot. Use a test environment and non-sensitive examples where possible. A failed test should lead to a documented restriction or fix, followed by another check of the affected behaviour. Do not hide a boundary failure inside an overall success percentage.
Keep useful audit records and a working stop control
Record the initiating identity, requested operation, authorisation outcome, relevant approval and destination acknowledgement. Use stable identifiers so an operator can follow the same task across queues and connectors. Avoid indiscriminately copying full documents, credentials or private conversations into diagnostic logs. Decide who can inspect logs and how long they are needed.
Provide a way to stop new actions while preserving pending work. Decide whether a pause affects one workflow, one connector or every agent operation. Test that the stop control actually prevents execution, including queued requests. A reassuring dashboard indicator is not enough if a background worker continues changing records.
Write a recovery procedure that identifies who investigates, how affected records are found and which changes can be reversed. Some communications cannot be recalled, so prevention and review matter alongside rollback. Rehearse the procedure with the people who will operate the system, rather than treating handover documentation as the final technical deliverable.
Budget for controls, testing and ongoing operation
Request a scoped GBP proposal that separates discovery, connector work, permission enforcement, approval interfaces, evaluation and handover. This article gives no universal price band: the effort depends on your applications, their access model and the consequences of mistakes. A quote for a chat demonstration is not comparable to a quote for a supervised production workflow.
Operating costs include model usage, hosting, monitoring, human review and maintenance of the connectors and evaluation set. Ask who responds when access changes or the destination API behaves differently. Verify current vendor charges against the actual charging units before estimating usage; a security design cannot be priced from a model’s token rate alone.
Compare value using completed work after review and rework, not generated answers. Released staff capacity does not automatically become cash savings. Include the time spent handling exceptions and the cost of keeping a fallback process available. A smaller read-only pilot may be the sensible purchase if write access creates more supervision than the workflow can justify.
Build the business case from observed work
Use the same task definition before and during the pilot. If the baseline measures a completed enquiry but the pilot measures a generated note, the comparison will exaggerate the benefit. Record the time spent reviewing drafts, resolving exceptions and correcting destination records, including work performed by someone outside the original team.
| Business-case input | What to measure or request |
|---|---|
| Baseline effort | Handling time for a completed task using the current process |
| Pilot effort | Preparation, review, exception handling and rework for the same outcome |
| Released capacity | Observed difference in effort across the measured workload |
| Recurring spend | Usage, hosting, monitoring, maintenance and retained fallback costs |
| Implementation spend | The scoped quote, including control design and acceptance work |
| Financial benefit | Cash changes the business can substantiate, kept separate from capacity |
A pilot can release useful capacity without reducing payroll or producing immediate cash savings. It can also reveal that the review burden outweighs the preparation time saved. Keep both outcomes available to the decision-maker. The purpose of the worksheet is to support a defensible choice, including a narrower scope or no rollout.
Pilot a bounded workflow, then decide whether to expand
Imagine a hypothetical wholesaler whose assistant proposes CRM notes from incoming enquiries. Start with approved example records and drafts that do not change the CRM. Review whether the notes preserve meaning, avoid unrelated customer information and give staff enough context to decide. This is a scope example, not a customer success story.
Enable constrained updates only after the access, approval and failure tests meet the agreed conditions. Keep the old process available and assign ownership for exceptions. Judge the pilot on correct completed updates, review effort and recovery behaviour. If simple rules handle the task adequately, retaining those rules can be a successful result of the assessment.
Expansion needs its own decision. A new data source, user group or tool changes the boundary being tested. Do not extrapolate from a note-writing pilot to autonomous refunds or customer exports. Preserve the evidence from the earlier scope and establish what additional controls and tests the new capability requires.
Commission an integration with evidence you can review
Bring a workflow description, anonymised examples, an access map and a list of consequential actions to the initial discussion. Ask suppliers to explain which controls are enforced outside the model and demonstrate refusal as well as success. Request a deliverable that states the remaining limitations, ownership and conditions for deployment.
Our AI integration services can help scope a bounded agent workflow and its connections to existing applications. For a useful assessment brief, tell us which systems are involved, what the agent should read or change, and which actions require a person. Ask for a scoped GBP proposal covering the pilot, acceptance evidence and operating responsibilities.
The purchasing decision should be whether the proposed workflow earns its access. If the integration cannot explain who authorised a change or demonstrate how it stops, postpone wider permissions. Useful automation leaves your business with accountable operations as well as a faster way to prepare work.
Frequently Asked Questions
Can a stronger prompt make an AI agent secure? A stronger prompt can guide behaviour, but it does not establish access control. Enforce permissions in the connected application, validate proposed operations and test the boundaries under hostile inputs. Treat prompt improvements as one layer rather than a guarantee.
Should every agent action require human approval? No. Decide according to the operation and its consequences. Low-impact actions can run within an approved policy, while consequential changes need informed review or should remain unavailable. Test that approval applies to the exact operation executed.
What should an AI agent security assessment include? Include the workflow and access map, tool permissions, downstream authorisation, approval behaviour, adversarial tests, failure recovery and operational ownership. Request actual application evidence of refused requests as well as successful tasks, with unresolved limitations recorded.
How much does securing an AI agent cost? Request a scoped GBP quote covering connectors, access enforcement, review interfaces, testing and handover. Ongoing usage, monitoring, reviewer time and maintenance also matter. There is no universal price band that can account for different applications and consequences.
When should a business expand an agent pilot? Expand only when the current workflow meets its agreed acceptance conditions and has an operational owner. New sources, tools or user groups require a fresh scope decision and relevant tests. A successful draft-writing pilot does not justify unrelated privileged actions.