Large language model applications increasingly read content they did not write: emails, web pages, uploaded documents and records retrieved from internal knowledge bases. Any of that content can contain text crafted to look like instructions. This is indirect prompt injection, and it is one of the most important risks in the OWASP Top 10 for LLM Applications.
Why it matters more with tools
An assistant that only answers questions can be manipulated into giving wrong or embarrassing answers. An assistant that can send email, query databases or call APIs can be manipulated into taking actions. The impact of prompt injection scales with the permissions of the system.
Defensive principles
- Treat all retrieved and user-supplied content as untrusted data, never as instructions
- Scope tool permissions to the minimum required for the task
- Require human confirmation for consequential actions
- Enforce authorisation outside the model: the model should never decide what a user is allowed to access
- Filter and constrain outputs before they reach other systems
- Log prompts, retrieved sources and tool calls for investigation
Test it like any other attack surface
Include indirect injection cases in your evaluation and red-team suites: poisoned documents, hidden text and instructions embedded in data fields. Re-run them whenever the model, prompt or tool set changes.