AI agent security becomes a business decision when an assistant can change records, send messages or trigger work in another application. A convincing demonstration tells you whether the agent can complete a task. It does not establish whose information it can access, which actions it may take or how your team recovers when something goes wrong.

Secure an AI agent by limiting its data access and available actions, enforcing authorisation in the connected applications, and requiring informed approval for consequential changes. Test those boundaries with realistic failures and hostile inputs before expanding access. A better prompt alone cannot provide that assurance.

For UK teams commissioning an integration, the useful question is what the agent is allowed to do when its judgement is wrong. This guide proposes a practical scope for a controlled pilot, the evidence to request from a supplier and the costs that belong in the business case. It does not promise an attack-proof system or describe a named customer’s deployment.

Define AI agent security around a business task

Start with a narrow workflow such as preparing a support reply or proposing a CRM update. Write down the source records, the intended user and the final destination. Separate reading, drafting and executing: access to a customer record need not imply permission to export it, and drafting a response need not imply permission to send it.

Agree what remains outside the pilot. Refunds, account permission changes and bulk exports are examples of actions that deserve separate decisions. A supplier should be able to show where those restrictions are enforced. A sentence in a system prompt is useful guidance, but it is not evidence that a connected API will reject an unauthorised request.

Give the workflow a named owner who can decide whether an exception is acceptable. Without that person, technical teams can quietly inherit business decisions about customer communications or record changes. A useful brief describes an authorised outcome and its boundaries, rather than asking for an agent that can use every available tool.

Treat incoming content as evidence, not instructions

OWASP’s prompt-injection guidance distinguishes direct manipulation through a prompt from indirect manipulation through external material such as files and websites. It also warns that retrieval and fine-tuning do not fully remove this vulnerability. Connecting an agent to company documents therefore does not make every retrieved sentence trustworthy.

Consider a hypothetical support email that asks an assistant to copy an unrelated customer’s records into its reply. The email is content to assess, not permission to widen access. The important test is whether the surrounding application blocks that attempted disclosure even if the model follows the instruction. Apply the same reasoning to attachments, search results and tool responses.

Label external content and validate proposed outputs, but do not sell these measures as a complete defence. Our recommendation is to design the integration so a misleading answer has limited authority. A model that suggests an inappropriate action should meet an independent access check before anything happens in the business system.

Make permissions follow the user and the record

The connected application should check both who is asking and which record the request concerns. An agent acting for a support colleague should not automatically inherit an administrator’s access. Define whether a service identity is necessary, what it can read or change, and how the integration preserves the initiating user’s scope.

OWASP’s excessive-agency guidance recommends minimal functionality and permissions, user-context execution and downstream authorisation. It identifies excessive functionality, permissions and autonomy as separate causes of damaging actions. A read-only connector and a narrowly scoped account address different parts of that problem.

Test with accounts that have different responsibilities. Attempt to read another team’s material, alter a protected field and request an export outside the approved scope. Record the actual refusal from the application. A successful happy-path demonstration under an administrator account cannot show that ordinary users are isolated from information they should not see.

Put an action gateway between suggestion and execution

An action gateway is application code that checks a proposed operation before calling the destination system. For a CRM pilot, it might accept only approved record identifiers and permitted fields. Reject unknown operations and unsupported values. Keep credentials in the connector’s controlled environment rather than placing secrets in model-visible instructions.

Prefer specific operations such as proposing a note over a general tool that can run arbitrary commands or contact arbitrary destinations. The smaller interface is easier to inspect and test. It also gives the business a concrete list of capabilities to approve when the pilot expands.

CapabilityInitial pilot boundaryEvidence to request
Read a recordUser-authorised records onlyCross-account access is refused
Draft a messageNo automatic sendingDraft remains reviewable
Update a fieldApproved fields and valuesInvalid changes are rejected
Export informationDisabled unless separately scopedUnapproved destinations are blocked

This matrix is a proposed starting point, not a universal policy. Adjust the boundaries to your workflow and error consequences. Include ordinary application controls: an agent integration still needs reliable authentication, validated inputs and a recoverable connection to the destination.

Follow the proposal through the control boundary

The diagram separates a model’s suggestion from the application’s decision to execute it. Content can influence the proposed operation, but the gateway checks identity, record scope and permitted changes independently. A consequential action also waits for approval of the final operation. Requests outside policy take a refusal path instead of quietly acquiring more authority.

An AI agent proposes an operation. Application code checks identity, record scope and allowed fields. Actions outside policy are refused; consequential permitted actions require approval before execution and a destination acknowledgement.
The model suggests an action; application checks and any required human review control execution.

Make approval a real decision

An approval screen should show the action, destination and material change being authorised. For an outgoing message, show its recipient and final text. For a record update, show the existing value and proposed replacement. Asking someone to approve an unexplained instruction transfers responsibility without giving them the information needed to exercise it.

Bind approval to the operation actually executed. If the target, content or relevant record state changes, require a fresh decision under the agreed policy. Otherwise a person can approve one version while the software executes another. Include expiry and cancellation behaviour in the acceptance tests rather than leaving them as interface details.

Do not make every trivial operation an approval task. That produces a queue people learn to dismiss. Agree which actions need review, which can run within an established policy and which remain unavailable. Measure whether reviewers can understand and complete the work, including busy periods and the absence of the usual approver.

Illustrative approval screen showing an example CRM record, its current follow-up value and proposed replacement. Approval applies only to the displayed target and final change; application permissions still apply.
Example interface: show the record, existing value and final change before approval.

Test failure paths before granting more access

Build an evaluation set from representative, anonymised work. Include missing fields, contradictory records, unsupported attachments and attempts to redirect the task. Keep some examples separate from the material used to adjust the system. The purpose is to test boundaries and usable outcomes, not reward a demonstration that memorised its examples.

Exercise the whole workflow. Interrupt the destination connection, repeat a request after a timeout, revoke a user’s access and change a record while approval is pending. Check whether the agent reports a truthful status and whether staff can resume work without duplicate updates. A polite refusal is insufficient if a tool already performed the forbidden action.

Ask for evidence that connects each test to an expected result, actual application behaviour and a remediation owner. Report unresolved limitations clearly. Repeat relevant tests after changes to prompts, models, tools or permissions. A pilot should establish the conditions under which the workflow can operate, including the conditions that require it to stop.

Ask for an acceptance record you can inspect

A useful acceptance record connects an attempted action to a visible result in the destination, rather than relying on the agent’s explanation. For example, a test may deliberately submit an update for a record outside the user’s scope. The expected outcome is a refusal with no destination change. Preserve the relevant identifiers so somebody other than the demonstrator can check that outcome.

Test conditionExpected evidenceReason to stop expansion
Unauthorised recordAccess refusal and no record changeThe connector bypasses the user’s scope
Changed draft after approvalNew approval required before executionA different action uses an earlier approval
Timeout after destination accepts an updateReconciled result without a duplicate changeA retry creates additional work
Stop requested with queued actionsPending actions remain unexecutedThe worker continues after the stop

Agree how the team will verify these results before running the pilot. Use a test environment and non-sensitive examples where possible. A failed test should lead to a documented restriction or fix, followed by another check of the affected behaviour. Do not hide a boundary failure inside an overall success percentage.

Proposed recovery flow: a permitted update is accepted but its response times out. Reconcile the operation with the destination; report confirmed completion or pause and investigate an unknown outcome instead of blindly repeating the change.
A timeout can hide a successful update. Reconcile the destination result before deciding what happens next.

Keep useful audit records and a working stop control

Record the initiating identity, requested operation, authorisation outcome, relevant approval and destination acknowledgement. Use stable identifiers so an operator can follow the same task across queues and connectors. Avoid indiscriminately copying full documents, credentials or private conversations into diagnostic logs. Decide who can inspect logs and how long they are needed.

Provide a way to stop new actions while preserving pending work. Decide whether a pause affects one workflow, one connector or every agent operation. Test that the stop control actually prevents execution, including queued requests. A reassuring dashboard indicator is not enough if a background worker continues changing records.

Write a recovery procedure that identifies who investigates, how affected records are found and which changes can be reversed. Some communications cannot be recalled, so prevention and review matter alongside rollback. Rehearse the procedure with the people who will operate the system, rather than treating handover documentation as the final technical deliverable.

Budget for controls, testing and ongoing operation

Request a scoped GBP proposal that separates discovery, connector work, permission enforcement, approval interfaces, evaluation and handover. This article gives no universal price band: the effort depends on your applications, their access model and the consequences of mistakes. A quote for a chat demonstration is not comparable to a quote for a supervised production workflow.

Operating costs include model usage, hosting, monitoring, human review and maintenance of the connectors and evaluation set. Ask who responds when access changes or the destination API behaves differently. Verify current vendor charges against the actual charging units before estimating usage; a security design cannot be priced from a model’s token rate alone.

Compare value using completed work after review and rework, not generated answers. Released staff capacity does not automatically become cash savings. Include the time spent handling exceptions and the cost of keeping a fallback process available. A smaller read-only pilot may be the sensible purchase if write access creates more supervision than the workflow can justify.

Build the business case from observed work

Use the same task definition before and during the pilot. If the baseline measures a completed enquiry but the pilot measures a generated note, the comparison will exaggerate the benefit. Record the time spent reviewing drafts, resolving exceptions and correcting destination records, including work performed by someone outside the original team.

Business-case inputWhat to measure or request
Baseline effortHandling time for a completed task using the current process
Pilot effortPreparation, review, exception handling and rework for the same outcome
Released capacityObserved difference in effort across the measured workload
Recurring spendUsage, hosting, monitoring, maintenance and retained fallback costs
Implementation spendThe scoped quote, including control design and acceptance work
Financial benefitCash changes the business can substantiate, kept separate from capacity

A pilot can release useful capacity without reducing payroll or producing immediate cash savings. It can also reveal that the review burden outweighs the preparation time saved. Keep both outcomes available to the decision-maker. The purpose of the worksheet is to support a defensible choice, including a narrower scope or no rollout.

Pilot a bounded workflow, then decide whether to expand

Imagine a hypothetical wholesaler whose assistant proposes CRM notes from incoming enquiries. Start with approved example records and drafts that do not change the CRM. Review whether the notes preserve meaning, avoid unrelated customer information and give staff enough context to decide. This is a scope example, not a customer success story.

Enable constrained updates only after the access, approval and failure tests meet the agreed conditions. Keep the old process available and assign ownership for exceptions. Judge the pilot on correct completed updates, review effort and recovery behaviour. If simple rules handle the task adequately, retaining those rules can be a successful result of the assessment.

Expansion needs its own decision. A new data source, user group or tool changes the boundary being tested. Do not extrapolate from a note-writing pilot to autonomous refunds or customer exports. Preserve the evidence from the earlier scope and establish what additional controls and tests the new capability requires.

Commission an integration with evidence you can review

Bring a workflow description, anonymised examples, an access map and a list of consequential actions to the initial discussion. Ask suppliers to explain which controls are enforced outside the model and demonstrate refusal as well as success. Request a deliverable that states the remaining limitations, ownership and conditions for deployment.

Our AI integration services can help scope a bounded agent workflow and its connections to existing applications. For a useful assessment brief, tell us which systems are involved, what the agent should read or change, and which actions require a person. Ask for a scoped GBP proposal covering the pilot, acceptance evidence and operating responsibilities.

The purchasing decision should be whether the proposed workflow earns its access. If the integration cannot explain who authorised a change or demonstrate how it stops, postpone wider permissions. Useful automation leaves your business with accountable operations as well as a faster way to prepare work.


Frequently Asked Questions

Can a stronger prompt make an AI agent secure? A stronger prompt can guide behaviour, but it does not establish access control. Enforce permissions in the connected application, validate proposed operations and test the boundaries under hostile inputs. Treat prompt improvements as one layer rather than a guarantee.

Should every agent action require human approval? No. Decide according to the operation and its consequences. Low-impact actions can run within an approved policy, while consequential changes need informed review or should remain unavailable. Test that approval applies to the exact operation executed.

What should an AI agent security assessment include? Include the workflow and access map, tool permissions, downstream authorisation, approval behaviour, adversarial tests, failure recovery and operational ownership. Request actual application evidence of refused requests as well as successful tasks, with unresolved limitations recorded.

How much does securing an AI agent cost? Request a scoped GBP quote covering connectors, access enforcement, review interfaces, testing and handover. Ongoing usage, monitoring, reviewer time and maintenance also matter. There is no universal price band that can account for different applications and consequences.

When should a business expand an agent pilot? Expand only when the current workflow meets its agreed acceptance conditions and has an operational owner. New sources, tools or user groups require a fresh scope decision and relevant tests. A successful draft-writing pilot does not justify unrelated privileged actions.