What it is
Prompt injection is an attack on language models (LLMs): an instruction hidden in an email, web page or document is read by the assistant or AI agent as if it were a legitimate command. It can push it to reveal data, send messages or perform actions the user did not ask for.
How it works
- 1The attacker places instructions in content the agent will read: hidden text, comments, metadata.
- 2The agent reads the content together with your requests and does not always tell what is data and what is an order.
- 3If it has access to email, files or tools, it can send data out or perform actions.
- 4The more permissions the agent has, the worse the possible damage.
How to spot it
- The assistant does things you did not ask for or cites instructions you did not give
- Links or images with data in the address generated by the agent
- Unrequested actions on email, files or payments
- Different behaviour after reading a certain page or document
How to defend
- Give agents only the strictly necessary permissions (least privilege)
- Always require human confirmation for sensitive actions: sending, payments, deletions
- Treat all external content as untrusted and keep it apart from sensitive data
- Check and limit the addresses and tools the agent may use
- Log and monitor actions, and test the system with attack exercises
If you think you have been hit
- Revoke the agent’s tokens and keys at once and rotate secrets
- Check the logs to see which actions it performed
- Reduce permissions before turning it back on
And there are many, many more
The attacks above are only some of the most common: there are hundreds, and new ones appear every week. If the one that concerns you is not among them, write to me: I will tell you whether it really affects you and how to defend.
Contact meOther attacks
Watch the Shorts on YouTubeMatteo Russo · Updated October 2026