What prompt injection is
Large language models follow instructions written in plain language, and they struggle to tell the difference between instructions from you and instructions hidden in the content they read. Prompt injection exploits that.
- Direct injection: a user types instructions designed to override the system's rules.
- Indirect injection: the instructions are hidden in a document, web page, email or database record that the AI reads while doing its job. The attacker never talks to your system directly.
For a chatbot, the result might be embarrassing output or a leaked system prompt. For an agent connected to email, files, payments or code, it can mean data sent to an attacker or actions taken in your name.
The wider risk picture
The OWASP Top 10 for Large Language Model Applications is the most widely used reference. Its 2025 edition covers:
| Risk | What it means for leaders |
|---|---|
| Prompt injection | Hidden instructions steer the system. |
| Sensitive information disclosure | Personal data, secrets or confidential content leak in answers. |
| Supply chain | Risky models, plug-ins and third-party components. |
| Data and model poisoning | Manipulated training or reference data skews behaviour. |
| Improper output handling | AI output passed to other systems without checks. |
| Excessive agency | Agents with more permissions or autonomy than they need. |
| System prompt leakage | Hidden instructions and secrets exposed. |
| Vector and embedding weaknesses | Search indexes that bypass permissions or can be poisoned. |
| Misinformation | Confident, wrong answers that people act on. |
| Unbounded consumption | Runaway usage and cost, or denial of service. |
Controls that work
- Least privilege for tools. Give agents only the actions and data they need, with read-only access wherever possible.
- Human approval for consequential actions, such as payments, external emails, deletions and changes to records.
- Treat retrieved content as untrusted. Separate it from instructions, and check outputs before they reach other systems.
- Permission-aware retrieval, so the AI cannot surface content a user is not allowed to see.
- Monitoring and limits on usage, cost and unusual behaviour, with an off switch.
- Adversarial testing before launch and after every significant change to models, prompts, tools or data sources.
Ten questions to ask your team
- Which of our AI systems can take actions, not just answer questions?
- What is the worst thing each one could do if it were tricked?
- Which external content do they read, such as email, web pages or uploaded files?
- Which actions require a person to approve them?
- Can the AI see data the user asking the question cannot?
- Where are our system prompts, keys and secrets, and could they leak?
- How would we know if an agent started behaving unusually?
- Have we tested for indirect prompt injection specifically?
- Which third-party models and plug-ins are we relying on?
- Who is accountable for each AI system's behaviour?
How we help
Our AI Security Assessment threat-models your AI systems, tests them hands-on against these risks, and gives your board a clear summary and your engineers a prioritised fix list, with a retest to confirm.