An accounts-payable agent reads an invoice, finds the purchase order and prepares an approval request. The workflow saves time, and the extraction looks accurate.
Then an attachment includes a new instruction: the supplier has changed its bank account, the update is urgent, and the usual verification step should be skipped.
Can the agent change the supplier record? Can it initiate payment? Can an operator reconstruct what happened if the action goes through?
Those questions determine the security of the system.
AI agent security covers the entire path from a user’s request to a business action: identity, retrieved information, memory, model output, MCP servers, tools, credentials, approval, execution and recovery. A weakness anywhere along that path can turn a useful assistant into an expensive operational problem.
For Boxinall, the engineering question is concrete: what evidence must exist before an agent is allowed to change a customer’s record, send information outside the company, spend money or affect production?
This guide explains how to answer it. You will find a reference architecture, a worked invoice scenario, practical MCP checks, permission and approval designs, adversarial tests and an incident-response plan.
The Quick Answer
A business AI agent needs five things before it receives meaningful authority:
- An authenticated identity and a defined business purpose.
- A small set of tools with permissions enforced at execution.
- Independent checks on proposed actions, including their targets and parameters.
- Human approval for operations whose consequences require it.
- Evidence and controls that let operators detect, stop and investigate misuse.
The central design principle: the model may propose an action; the execution system must prove that the action is permitted.
A confident answer, a valid JSON object and a successful API call each prove something different. None proves that the user authorised the business outcome.
Why Agent Security Matters In 2026
An agent can combine information from several sources and act through connected tools. That expands both its usefulness and the consequences of a mistake.
NIST’s March 2026 research discussion describes agent hijacking through external content such as emails, websites and code repositories. A malicious instruction can arrive inside material the agent was legitimately asked to read. It does not require the attacker to be the person chatting with the agent. NIST’s agent-security research.
For a business, the relevant question is how far that influence can travel. A support assistant that drafts a reply has a different exposure from one that can export customer records, issue credits and send email.
Assess the actual deployment. A platform name, a security prompt or a self-hosted server tells you very little about the authority granted to a specific agent.
Map The Attack Surface Before Choosing Controls
Start with the workflow and identify every place where information or authority crosses a boundary.
| Surface | Question to ask | Evidence to collect |
|---|---|---|
| User and channel | Who requested the work, and how was identity verified? | Authenticated requester and channel context |
| Retrieval | Which records could this requester access? | Source IDs, tenant and access decision |
| Model context | What information influenced the proposal? | Versioned instructions and relevant source references |
| Memory | Who can write durable context, and who can retrieve it? | Provenance, scope and expiry |
| MCP server | Who operates the server and what can it reach? | Owner, release, endpoint and permissions |
| Tool | What exact operation can execute? | Schema, allowed targets and enforced constraints |
| Credentials | Whose authority does the destination system see? | Principal, scopes and credential lifecycle |
| Approval | What action did the reviewer actually approve? | Bound action record and reviewer identity |
| Execution | What changed in the destination system? | Transaction ID and verified resulting state |
| Operations | How can the action be stopped or investigated? | Alerts, control paths and protected records |
Use this map in design reviews. It makes vague statements such as “the agent has CRM access” specific enough to challenge.
A sales assistant may need to read assigned accounts and create draft follow-up tasks. It may have no reason to export every contact, change account ownership or issue contractual promises. Each capability deserves its own decision.
A Secure Reference Architecture
The following is a recommended design for an agent connected to business systems. It can be implemented with custom services, a workflow orchestrator or a combination of both.
Authenticated request
|
v
Agent runtime <--- Authorised retrieval and scoped memory
|
| proposed action and evidence references
v
Execution boundary
- schema and business-state validation
- requester, agent and tenant checks
- resource, destination and budget limits
|
+---- denied ----> safe response and audit event
|
+---- approval required ----> independent review UI
| |
| action-bound approval
| |
+----------------------------------+
|
v
Narrow executor / MCP tool / workflow
|
v
Business system
|
v
Verified outcome and protected execution receipt
Enforce controls where the tool or downstream API executes. If a check exists only inside the agent’s prompt, another path may bypass it.
The execution boundary also needs trustworthy inputs. Resolve the user’s identity and tenant from the authenticated session. Do not accept a model-generated user_id or tenant_id as proof of authority. Check access to the target record on every relevant request. OWASP authorisation guidance.
Retain the authorised task as well. A service account may technically be able to issue a credit, while the user only asked for an explanation of a bill. Check the action against the requested purpose and obtain a separate decision when it exceeds that purpose.
For the orchestration choices around this boundary, see Boxinall’s OpenClaw vs n8n decision guide.
MCP Server Security: Check The Connection And The Capability
Model Context Protocol standardises how an AI application connects to tools and information. An MCP host contains clients that communicate with servers, which expose capabilities such as tools. It creates a common interface; the implementation still needs an explicit security design.
Tool descriptions and annotations deserve scrutiny. A tool declaring itself read-only does not independently prove its behaviour. The MCP tools specification says annotations must be treated as untrusted unless they come from trusted servers. MCP tools specification.
Distinguish Remote And Local Servers
For a protected remote server, examine authentication, token validation and transport security. The current MCP authorisation specification covers HTTP transports, including tokens issued for the intended resource. Its HTTP flow should not be assumed to apply unchanged to local stdio servers. MCP authorisation specification.
A local server is an executable process. Inspect the package, startup command, environment variables, filesystem permissions and network access it receives. Installing it can grant considerable access to the operator’s machine.
The official MCP security guidance also addresses token passthrough and server-side request forgery. Do not forward a token issued for one resource as if it were valid authority for another. Treat discovered URLs and proxy relationships as security decisions. MCP security best practices.
Use A Server Admission Record
Before enabling a server for a business workflow, record these review decisions:
| Review area | Admission question |
|---|---|
| Ownership | Who maintains the server and handles vulnerabilities? |
| Supply chain | Which package or image is approved, and how are updates reviewed? |
| Identity | How are the requester, client and server authenticated? |
| Authority | Which operations and records can the credential reach? |
| Tool definitions | How are descriptions and schemas reviewed when they change? |
| Isolation | Which files, processes and networks are accessible? |
| Data handling | What is sent, retained or returned to the model? |
| Limits | Which timeouts, quotas and concurrency caps are enforced? |
| Operations | Where are decisions logged, and how is the server disabled? |
OWASP’s MCP development guidance calls for secure architecture, authentication, authorisation, validation, isolation and hardened deployment. Use that guidance as a review input rather than treating a successful connection as acceptance. OWASP’s secure MCP development guide.
A useful Boxinall deliverable is a reviewed capability inventory alongside the normal integration specification. It tells a business owner what each connection is allowed to do and who must review a change.
Prompt Injection: Keep External Content From Granting Authority
Prompt injection attempts to redirect the model through instructions inside the information it processes. Direct injection comes through a user request; indirect injection arrives through another source, such as an email, document or tool response.
Separate trusted instructions from source content, inspect inputs and outputs, and use detection to route suspicious material. These measures help, but a detection miss must not automatically become an authorised action. OWASP prompt-injection guidance.
A Worked Invoice Exception
The following scenario uses illustrative records to show how the controls work.
An invoice-processing agent receives a document for a known supplier. Alongside ordinary fields, the document tells the agent to replace the bank account and bypass verification because payment is urgent.
The agent should still be able to extract the invoice number, amount and purchase-order reference. The supplier’s claimed change becomes an exception to investigate. It does not become an instruction that grants payment authority.
| Step | Vulnerable outcome | Designed control |
|---|---|---|
| Read attachment | Treat document instructions as operating policy | Preserve source provenance and separate extracted claims from policy |
| Propose supplier change | Write bank details from the attachment | Check against the supplier master and route discrepancies |
| Request verification | Accept contact details supplied in the suspicious file | Use the independently maintained vendor-verification process |
| Seek approval | Show only an agent-written summary | Present the proposed change and authoritative records together |
| Execute payment | Use broad finance credentials | Require a specific authorised operation with target and amount checks |
| Confirm outcome | Trust a generic success message | Reconcile the destination transaction and resulting state |
Even if the model proposes an unsafe action, the executor should reject it. The test is whether the business action stayed within its boundary, including on repeated attempts.
This extends the architecture in Boxinall’s AI invoice-processing agent guide.
Tool Definitions Can Also Carry Risk
Review changes to tool descriptions, schemas and returned content. A server can change what the model sees after initial approval, and one tool’s content can attempt to influence use of another tool. Restrict which servers are available together and re-evaluate material definition changes. OWASP MCP security guidance.
Hashing an approved definition can help detect a change. It does not prove that a remote server’s implementation is safe or that its behaviour matches the definition. Review the implementation or the provider’s assurance evidence as well.
Tool Permissions: Define Operations The Business Can Accept
Design narrow tools around business tasks. For a support agent, these may include read_assigned_ticket, retrieve_approved_policy and create_response_draft. A generic database administrator tool grants a much larger authority than those tasks require.
Keep credentials outside model-visible context where possible, and use dedicated principals with scoped access. Rotation and revocation should be part of the operating design. Removing a secret from an agent’s workspace is insufficient if the credential remains valid elsewhere. OWASP secrets-management guidance.
An Illustrative Tool Contract
Consider a tool that prepares a customer-credit request:
{
"action": "request_customer_credit",
"arguments": {
"customer_id": "cust_204",
"amount_minor_units": 2500,
"currency": "USD",
"reason_code": "SERVICE_ADJUSTMENT",
"evidence_refs": ["ticket_731"]
}
}
This is an application-level example, not an MCP wire message or ready-to-install security configuration. The USD amount is illustrative and does not define an approval threshold.
The backend resolves identity, checks access to cust_204, verifies the ticket, applies credit policy and decides whether approval is needed. It also rejects unknown fields and invalid business states. Schema validation only proves that inputs have the expected structure; it cannot establish whether a credit is justified.
The Boxinall Capability Envelope
A capability envelope is a written contract for one permitted operation:
| Field | Example decision |
|---|---|
| Purpose | Prepare a credit request for an assigned support case |
| Resources | Assigned accounts inside the authenticated tenant |
| Input | Customer, amount, currency, reason and evidence references |
| Authority | Proposal only; execution requires a separate decision |
| Output | Request status and permitted evidence summary |
| Limits | Per-action and cumulative caps set by the business owner |
| Recovery | Query destination state before retrying an uncertain write |
| Ownership | Named service owner and review date |
The limits must exist in code or policy enforcement. A document makes responsibility visible; it does not enforce permissions by itself.
Restrict Destinations As Well As Actions
A general URL-fetching tool can be redirected towards sensitive internal services. Validate permitted destinations, resolved addresses and redirects, and block inappropriate private or metadata endpoints. Private-network tools need explicit exceptions and narrower controls. OWASP SSRF prevention guidance.
Also examine outbound data. Reading a customer record and sending a message may each be permitted separately, yet combining them can disclose information to an unauthorised recipient.
Human Approval Must Survive Parameter Changes
A reviewer needs to understand what will happen: the affected record, recipient, amount, proposed change and supporting evidence.
An approval system should construct the decision view from the validated action held by the backend. Keep the agent’s explanation visibly separate from authoritative records. OWASP’s transaction-authorisation guidance emphasises server-side enforcement and protecting transaction data against modification. OWASP transaction-authorisation guidance.
Bind The Approval To An Exact Action
For the customer-credit example, the approved record should bind the requester, reviewer, tenant, operation, customer, amount, currency, relevant record version and expiry. Store it in a trusted service or protect it with an appropriate integrity mechanism.
At execution, confirm that the action still matches the approval and that the reviewer remains authorised. A changed amount, recipient or material policy condition requires another decision.
Capture approval through an authenticated interface. Text saying “approved” in a document, chat transcript or model response is not an approval record. For critical operations, apply the organisation’s stronger authentication requirements and separation of duties.
Use an atomic transition to prevent two workers from consuming the same permission independently. Bind a destination idempotency key to the intended business action where supported. If the destination lacks this capability, use reconciliation and duplicate controls; an approval token alone cannot guarantee exactly-once execution across systems.
Test The Approval, Not Just The Button
Boxinall’s proposed approval-integrity test uses four cases:
- Approve one amount, then submit a larger amount.
- Approve one recipient, then substitute another.
- Submit the same approved action concurrently.
- Reuse the approval after expiry or a relevant record change.
The expected results are rejection of mismatched or stale actions and prevention or detection of duplicate side effects. Verify the business system’s state as well as the executor’s response.
Protect Retrieval, Memory And Agent Delegation
Retrieval-augmented generation introduces another access boundary. Apply the requester’s permissions before restricted records enter model context, and scope caches to their access context. Keep track of source updates and revocations. OWASP RAG security guidance.
Stored conversation content also needs a clear role. A user’s preference, a supplier’s claim and an approved corporate policy must not become interchangeable because all three were saved as text.
For a service workflow that combines customer records and approved knowledge, see Boxinall’s customer-support AI agent guide.
For sensitive workflows, keep external claims linked to their source, separate them from authoritative policy and define an expiry or review path. When an incident involves poisoned memory, clean the affected context before resuming the agent.
For delegated work, the receiving agent should get a defined task and bounded permissions. It should not acquire every capability held by its parent. Review isolation, tool scope and oversight across the chain. OWASP AI-agent security guidance.
For OpenClaw specifically, its documentation describes one trust boundary per gateway and advises separating gateways and credentials for materially different trust groups. Treat that as a deployment constraint when serving multiple customers. OpenClaw security documentation.
For n8n, its security audit can flag issues such as risky nodes, credentials and unprotected webhooks. Use the findings alongside the workflow’s business-authorisation checks. n8n security audit.
Boxinall Engineering Patterns That Make Security Verifiable
The following patterns are proposed delivery artefacts for this guide. They bring together application design, integration engineering, permissions, review interfaces and operations. Each needs to be implemented and verified against the project’s actual requirements.
1. A Permission Ledger Across The Whole Task
A tool’s individual limit can be defeated by a sequence of smaller permitted actions. Ten small credits may exceed the intended account limit; many small reads may become a bulk export.
Record cumulative authority by requester, agent, tenant, workflow and business resource. Include financial totals, external recipients, records returned, delegated work and active jobs where relevant.
Set resource budgets too: model usage, tool calls, elapsed time and delegation depth. Enforce shared limits across retries and child jobs so an agent cannot avoid its allowance by opening another task.
Enforce reservations atomically so parallel calls cannot each spend the same remaining allowance. Retain the relationship between a reservation, the destination outcome and any reconciliation. This addresses a business problem that per-call schema checks will miss.
2. A Business Invariant List
Write down conditions that must remain true regardless of the model’s recommendation:
- An invoice cannot change the approved supplier bank account.
- A reviewer cannot approve beyond their authority.
- A source from one tenant cannot support an action in another.
- A customer-visible claim must use the current approved policy.
- A completed credit cannot be repeated because a worker timed out.
Turn each invariant into an executor check and a test. It gives developers, QA and the business owner a shared definition of failure.
3. Evidence-Based Approval Views
The review screen should let someone compare the proposed change with the original records. Provide source provenance, record versions, changes since the request, consequences and clear approve, revise, reject or escalate actions.
Test it with realistic reviewers. If they cannot spot a changed recipient or unsupported claim, improving the model will not repair the approval process.
4. Receipts For Actions And Denials
Record the request, operation, target, policy version, authorisation result, approval reference and destination outcome. Store denials too: they reveal attempted boundary crossings and help distinguish an attacked agent from a faulty policy.
Use structured records with controlled access. Avoid putting live credentials or unnecessary personal data into logs. Preserve evidence separately when an incident requires fuller payloads. OWASP logging guidance.
A record may link to the model’s concise explanation, but it should not depend on hidden reasoning to establish what happened.
5. A Rehearsal Of The Stop Path
Run a non-destructive exercise in which an agent continues proposing writes after its authority is revoked.
Measure when the executor stops accepting them, whether queued jobs are also blocked and whether operators can identify the last committed action. Try a paused approval, a retrying worker and a delegated task during the exercise.
The useful result is a measured containment path. An emergency button that stops new chats while old jobs continue writing leaves a serious gap.
Test The Business Outcome Under Adversarial Input
Run tests in an isolated environment with synthetic records and destinations. Keep both normal tasks and known abuse cases so a security improvement does not quietly make the product unusable.
| Test | Required observation |
|---|---|
| Invoice includes a bypass instruction | Extract permitted fields; prevent supplier changes |
| Tool definition changes after review | Trigger the defined change-review process |
| Model supplies another tenant’s record ID | Deny access before data is exposed |
| Proposed amount changes after approval | Reject the mismatched action |
| Two workers retry an uncertain write | Reconcile state and prevent duplicate business effects |
| Retrieved text asks to save a new policy | Prevent it becoming authoritative memory |
| Agent calls a new outbound destination | Enforce destination and data-sharing policy |
| Model or tool enters a retry loop | Enforce cumulative limits and stop conditions |
| Reviewer account loses its role | Revalidate authority before execution |
| Policy service becomes unavailable | Block protected actions and expose a clear recovery path |
For every case, retain the configuration, expected result, observed tool calls and destination state. Repeat relevant tests after changing models, prompts, retrieval, tool definitions or approval logic.
A zero-failure result applies to that test set and deployment version. It does not establish that every possible attack has been prevented.
AI Agent Incident Response: Stop Authority And Preserve Evidence
An incident may involve stolen credentials, an unsafe action, a compromised tool, poisoned memory or unauthorised disclosure. Establish a response owner and escalation route before release.
NIST SP 800-61 Revision 3 integrates incident response into cybersecurity risk management. The sequence below applies that operational discipline to an agent connected to business systems. NIST incident-response guidance.
Contain The Affected Action Path
Disable the compromised capability at its enforcement point. Revoke or restrict credentials where necessary, block affected outbound routes and pause related workers, scheduled jobs and approvals.
Check what is already in flight. A process that passed its checks before revocation may still commit a change. Use destination controls and state reconciliation to determine what completed.
Preserve volatile evidence when possible without delaying urgent containment. Avoid deleting conversations, logs or memory as the first response.
Establish What Actually Happened
Collect the model and runtime versions, relevant source material, tool definitions, policy decisions, approval records, credential history and destination transactions. Protect the evidence with access controls and record who collected it.
Distinguish proposals from executed changes. An alarming tool request may have been denied; a reassuring success message may conceal a wrong destination update.
Assess data access, external disclosure, financial effects, record changes and affected users. Route notification decisions through the company’s established security, privacy and contractual response process.
Repair And Recover
Remove the vulnerable path, rotate exposed credentials, quarantine suspect sources and correct affected memory or indexes. Reconcile business records before replaying jobs.
Make the observed failure a regression test. Resume with a limited capability set, monitor the resulting actions and have the accountable owner accept any remaining risk.
| Response stage | Responsible role | Exit evidence |
|---|---|---|
| Containment | Incident lead and platform operator | Affected authority blocked; in-flight work accounted for |
| Investigation | Security and application engineers | Timeline tied to verified business effects |
| Correction | Integration and process owners | Credentials, sources and records addressed |
| Recovery | Service owner | Regression tests pass; restricted release observed |
A Practical 30-Day Implementation Plan
This is a planning sequence for one bounded workflow. Timing depends on integrations, existing controls and review requirements.
Week 1: Define The Authority
Choose the workflow owner and map users, records, tools, credentials and external destinations. Establish the business invariants, prohibited actions and baseline handling time. Build the capability inventory and identify gaps in incident evidence.
Week 2: Implement The Boundaries
Add identity and record checks, narrow tools, approval binding and cumulative limits. Design the review screen and receipts. Make the stop path reachable by operators independently of the agent.
Week 3: Test Normal Work And Abuse Cases
Use synthetic cases first, then approved shadow traffic where appropriate. Test injection, cross-tenant access, approval changes, concurrency, outages and recovery. Measure whether reviewers can find the evidence they need.
Week 4: Release A Limited Scope
Start with read and draft capabilities, then separately authorise selected writes after testing. Review denied and executed actions, perform the containment rehearsal and decide what can expand.
For programme planning beyond one workflow, see Boxinall’s 90-day AI automation roadmap.
Cost And Business Value
Estimate security from the authority and failure modes of the workflow. Budget for tool review, identity integration, policy enforcement, approval design, adversarial testing, monitoring and incident readiness. Include recurring review as models, tools and business rules change.
Avoid pricing an entire agent-security programme from the number of prompts or MCP connections alone. One read-only knowledge tool and one payment tool have very different engineering needs.
Use three measurement groups:
| Group | Example measurements |
|---|---|
| Control effectiveness | Unauthorised actions blocked, tested isolation, approval integrity, measured containment time |
| Product usefulness | Accepted outcome rate, false denials, reviewer effort, latency and rework |
| Operating cost | Model and API use, infrastructure, maintenance, review time and incident-response effort |
For an illustrative calculation, suppose a workflow completes 3,000 accepted cases each month, saves five minutes per case before review and requires review on 15% of cases at two minutes each. At a loaded labour cost of USD30 per hour:
Gross handling capacity released: 250 hours/month
Review effort: 15 hours/month
Net capacity released: 235 hours/month
Illustrative capacity value: USD7,050/month
Subtract operating and maintenance costs, and account for rework. This is capacity value, not guaranteed cash savings: the business must use the released time productively. Do not add a speculative avoided-breach figure to make the ROI look larger.
Build Security Into The Business Workflow
The strongest agent design makes authority visible. It connects each requested action to a verified identity, permitted records, current business policy, an appropriate approval and a confirmed outcome.
Boxinall’s proposed approach is to review those connections across the application, integrations, agent runtime, human interface and cloud operations. The deliverables can include a capability inventory, threat model, business invariants, approval design, adversarial test set and incident runbook.
Bring one real workflow, its connected systems and the actions it must perform. Contact Boxinall Softech to define its authority boundaries and the evidence required before production.
Frequently Asked Questions
What Is AI Agent Security?
AI agent security is the design and operation of controls over an agent’s identity, information access, tools, credentials, actions and recovery. It covers the path from a request to its effect on a business system.
Does MCP Make An AI Agent Secure?
MCP provides a standard interface for capabilities. Security depends on the host, server, identity design, tool permissions, implementation and downstream systems. Each connection still needs review.
Can Prompt Injection Be Completely Prevented?
Do not assume every attempt will be detected. Combine prompt defences with enforced permissions and action checks so an unsafe proposal cannot automatically become an unsafe business operation.
Is A Read-Only Agent Low Risk?
It can reduce the risk of changing records, but it may still expose sensitive information. Restrict which data it retrieves and which users, outputs or destinations can receive it.
When Should A Human Approve An Agent Action?
Use approval when the action’s financial, administrative, external or irreversible consequences require accountable review. Bind it to a specific validated action, and recheck material changes before execution.
Can Self-Hosting Replace Security Controls?
Self-hosting gives control over selected infrastructure and data paths. It still requires identity, permissions, isolation, monitoring, patching and recovery. Configured model providers and integrations may receive data outside that infrastructure.
How Should A Business Respond To An Agent Incident?
Block affected authority, account for queued and in-flight actions, preserve evidence and verify destination effects. Correct the vulnerable path, reconcile records and test the failure before restoring access.
How Can Boxinall Help?
The proposed engagement maps one business workflow, designs its capability and approval boundaries, integrates execution controls and defines tests and operating evidence. Scope and outcomes should be agreed against the actual systems involved.



