Prompt Injection Defence: How to Secure LLM Applications and AI Agents
A defence-in-depth guide to prompt injection, with practical controls for untrusted content, tool permissions, output validation, monitoring and incident response.
By Maya Chen, Women in AI Editorial Fellow ยท 24 August 2026
Prompt injection is a systems-security problem disguised as a conversation.
An attacker places instructions in user input, a webpage, a document, an email, an image or another source the model processes. The model treats those instructions as relevant context and changes its behaviour. If the application can retrieve private data or call tools, a text manipulation can become a security incident.
The central engineering lesson is simple: do not treat the model as a security boundary.
Direct and indirect injection
OWASP distinguishes two common forms.
A direct injection arrives through the user's prompt. It might ask the model to ignore prior instructions, reveal hidden information or call a function outside the intended task.
An indirect injection is embedded in external content the system retrieves or reads. A summarisation agent may encounter malicious instructions in a webpage. A recruitment tool may process hidden text in a CV. A coding assistant may read an instruction placed in a repository file.
Indirect injection is especially difficult because consuming untrusted content is often the application's purpose.
The risk grows with agency. A chatbot that can only generate text has a narrower impact than an agent that can send emails, edit records, execute code or access private systems.
Why prompt wording is not enough
System prompts can reduce accidental drift and make attacks less reliable. They cannot guarantee isolation between instructions and data.
OWASP notes that retrieval-augmented generation and fine-tuning do not fully mitigate prompt injection. The model still processes a combined context and may follow malicious content.
Treat prompt instructions as one control inside a larger architecture. The strongest protections sit outside the model, where deterministic code can enforce identity, permissions, formats and approvals.
Start with a threat model
Map the application before choosing controls.
Identify:
users and attackers; sensitive data available to the system; external content the model can read; tools and actions it can invoke; trust boundaries; model and provider dependencies; outputs consumed by software or people; the most damaging plausible action.
Then describe attack paths.
For example:
"An attacker adds hidden instructions to a webpage. The research agent retrieves the page, follows the instructions and sends content from a private document store to an attacker-controlled URL."
This path identifies several possible controls: content isolation, restricted retrieval, outbound-network policy, tool approval and data-loss monitoring.
Separate data from authority
Mark external content as untrusted in the prompt and data model. Preserve provenance so the system and the user know where content originated.
Where possible:
retrieve only from approved sources; strip active content and unnecessary metadata; normalise documents before model use; segment external content clearly; limit how much untrusted text enters context; prevent retrieved content from defining tool arguments; show citations for consequential claims.
This does not make injection impossible. It reduces ambiguity and supports later validation.
Apply least privilege to every tool
The model should have the minimum authority needed for the current task.
Use:
scoped service identities rather than user-wide credentials; read-only access by default; narrow APIs instead of general shell or database access; row, tenant and resource-level controls; allowlists for destinations and actions; short-lived credentials; rate and spending limits; separate permissions for planning and execution.
OWASP recommends privilege control and handling functions in code rather than trusting the model to police its own access.
An email assistant that drafts a reply does not automatically need permission to send it. A data assistant that writes a query does not automatically need production write access.
Validate tool calls in code
Treat every model-proposed action as untrusted input.
Before execution, deterministic code should validate:
tool name; parameter types and formats; resource scope; user authorisation; destination; volume and rate; policy constraints; current session context.
For high-risk operations, require an explicit user or human approver to see the action, target and consequence. Avoid approval dialogs so vague that a user cannot understand what is being authorised.
A model should propose. The control plane should decide whether the proposal is permitted.
Constrain outputs
Model outputs can become inputs to browsers, databases, code interpreters and downstream agents. Validate and encode them for the destination.
Examples:
use parameterised queries rather than generated SQL strings; escape content before HTML rendering; require structured schemas and reject invalid fields; scan for secrets and personal data; restrict URLs and file paths; separate generated code from execution; cap output and recursion; verify claims against retrieved evidence.
Improper output handling can turn a model error into a conventional injection vulnerability.
Limit autonomous chains
Agents create compounded risk because one output becomes the next input and early errors can gain momentum.
Set limits on:
number of steps; execution time; cost; tool calls; accessible data; external destinations; parallel actions; retry behaviour.
Require checkpoints before irreversible or high-impact actions. Provide a kill switch that operators can use without relying on the agent itself.
Monitor behaviour, not only availability
Traditional service metrics are necessary but incomplete.
Log:
prompt and policy version; model and provider version; retrieved sources; tool requests and decisions; approvals and denials; identity and permission context; output validation failures; anomalous costs or sequences; user reports.
Handle logging lawfully and avoid creating a new repository of sensitive prompts. Redact where possible and control access.
Useful alerts include repeated blocked tool calls, attempts to reach unknown destinations, sudden increases in retrieval volume, unusual encoded content and actions outside a user's normal scope.
Test adversarially
Build an evaluation set that reflects the real application.
Include:
direct override attempts; hidden instructions in documents; multilingual and encoded attacks; malicious retrieved pages; instructions inside images; payloads split across multiple sources; attempts to extract system prompts or secrets; requests that combine legitimate and prohibited actions; poisoned memory or stored context.
Measure the outcome that matters. "The model mentioned the attack" is less important than "the application prevented unauthorised data access and tool execution."
Test again when the model, prompt, retrieval pipeline, tool set or permission design changes.
Design containment and recovery
The NCSC's secure AI development guidance calls for incident-management processes, logging, monitoring and secure operation throughout the lifecycle.
Prepare controls to:
revoke agent credentials; disable a tool; block a data source; change the system to read-only; switch to a known configuration; preserve relevant evidence; notify affected users; review actions already taken.
Version prompts, tool schemas, policies and retrieval configurations so rollback is possible.
A practical control stack
For a tool-using LLM application, a credible minimum includes:
authenticated users and scoped identities; clearly separated untrusted content; least-privilege tools; deterministic validation before execution; human approval for high-impact actions; destination-specific output handling; rate, cost and step limits; end-to-end logging; adversarial evaluation; rehearsed incident response.
No single layer is perfect. The point of defence in depth is that one failure does not grant the attacker the full capability of the system.
Communicate residual risk honestly
There is no credible claim that prompt injection has been eliminated from a general-purpose LLM application that processes untrusted content.
Engineering leaders should describe what has been reduced, what remains possible and why the approved use is proportionate to that residual risk. Narrow the use case when controls cannot contain the consequence.
The safest agent is not the one with the longest system prompt. It is the one whose architecture assumes the model can be manipulated and still prevents that manipulation from becoming an unauthorised action.
Explore our AI engineering hub, AI agent evaluation guide, AI incident response guide and agentic AI governance guide.