Insight · AI security

Prompt injection: the risk inside every AI agent.

The moment an AI system can read your data and act on your behalf, a sentence hidden in an email can turn it against you. Here is what leaders need to know, and ask.

Published 5 October 20267 minute read

What prompt injection is

Large language models follow instructions written in plain language, and they struggle to tell the difference between instructions from you and instructions hidden in the content they read. Prompt injection exploits that.

  • Direct injection: a user types instructions designed to override the system's rules.
  • Indirect injection: the instructions are hidden in a document, web page, email or database record that the AI reads while doing its job. The attacker never talks to your system directly.

For a chatbot, the result might be embarrassing output or a leaked system prompt. For an agent connected to email, files, payments or code, it can mean data sent to an attacker or actions taken in your name.

The wider risk picture

The OWASP Top 10 for Large Language Model Applications is the most widely used reference. Its 2025 edition covers:

RiskWhat it means for leaders
Prompt injectionHidden instructions steer the system.
Sensitive information disclosurePersonal data, secrets or confidential content leak in answers.
Supply chainRisky models, plug-ins and third-party components.
Data and model poisoningManipulated training or reference data skews behaviour.
Improper output handlingAI output passed to other systems without checks.
Excessive agencyAgents with more permissions or autonomy than they need.
System prompt leakageHidden instructions and secrets exposed.
Vector and embedding weaknessesSearch indexes that bypass permissions or can be poisoned.
MisinformationConfident, wrong answers that people act on.
Unbounded consumptionRunaway usage and cost, or denial of service.

Controls that work

  • Least privilege for tools. Give agents only the actions and data they need, with read-only access wherever possible.
  • Human approval for consequential actions, such as payments, external emails, deletions and changes to records.
  • Treat retrieved content as untrusted. Separate it from instructions, and check outputs before they reach other systems.
  • Permission-aware retrieval, so the AI cannot surface content a user is not allowed to see.
  • Monitoring and limits on usage, cost and unusual behaviour, with an off switch.
  • Adversarial testing before launch and after every significant change to models, prompts, tools or data sources.

Ten questions to ask your team

  1. Which of our AI systems can take actions, not just answer questions?
  2. What is the worst thing each one could do if it were tricked?
  3. Which external content do they read, such as email, web pages or uploaded files?
  4. Which actions require a person to approve them?
  5. Can the AI see data the user asking the question cannot?
  6. Where are our system prompts, keys and secrets, and could they leak?
  7. How would we know if an agent started behaving unusually?
  8. Have we tested for indirect prompt injection specifically?
  9. Which third-party models and plug-ins are we relying on?
  10. Who is accountable for each AI system's behaviour?

How we help

Our AI Security Assessment threat-models your AI systems, tests them hands-on against these risks, and gives your board a clear summary and your engineers a prioritised fix list, with a retest to confirm.

Sources

FAQ

Quick answers

Anything else? Email contact@sovereignsystemslabs.com.

What is prompt injection?

An attack where instructions hidden in input, documents, web pages or emails steer an AI system into doing something it should not, such as leaking data or misusing a connected tool.

Can prompt injection be fully prevented?

Not with today's models. The practical goal is to limit what a compromised AI system can do, through least privilege, human approval for risky actions, isolation of untrusted content and monitoring.

Is a normal penetration test enough?

Usually not. AI systems need testing designed for their specific risks, such as indirect prompt injection and excessive agency.

Book a free strategy call