Article overview
The big picture
A language model that drafts text creates a different risk profile from an agent that can search company records, update tickets, run code or communicate with customers. Once AI has access to business tools, a mistaken interpretation can produce a real-world action. Security architecture must therefore address not only what the model says, but also what the surrounding system permits it to do.
The OWASP GenAI Security Project published its Top 10 for Agentic Applications for 2026 and expanded its generative AI security resources in September 2026. These resources illustrate why traditional access controls, secure software engineering and new model-specific testing all belong in an agentic system’s design.
At a glance
Key takeaways
- Treat retrieved pages, documents, messages and tool responses as untrusted data.
- Enforce authorization in code at every tool boundary.
- Require confirmation for consequential actions and maintain meaningful audit trails.
- Test hostile instructions, privilege escalation and unexpected tool behavior.
Threat-model the system, not just the language model
An AI application includes prompts, retrieval systems, tools, identity services, logs and external integrations. Each boundary may expose sensitive data or grant unintended authority. Start by mapping which information the system can read, which operations it can perform and which users or services authorize those operations.
Distinguish helpful conversation context from trusted instructions. A document retrieved for summarization should not be allowed to redefine system rules, grant permissions or instruct the assistant to disclose credentials. The application must enforce these boundaries independently of the model’s natural-language behavior.
Treat prompt injection as an architectural risk
Prompt injection occurs when attacker-controlled or otherwise untrusted material tries to influence a model’s behavior beyond its intended role. An agent might encounter malicious text inside a web page, issue description, email or retrieved document. Because language models interpret instructions and content through similar interfaces, relying on a warning in the prompt alone is insufficient.
Reduce the impact of hostile instructions by constraining tool access, validating outputs and separating data from authority wherever possible. Test adversarial examples during development and after significant changes to prompts, tools or retrieval sources. Avoid treating any one filtering technique as a complete defense.
Put human approval at meaningful risk boundaries
Requiring approval for every low-risk retrieval would make an assistant unusable. Requiring none for high-impact actions would create avoidable exposure. Design approval around consequences: sending external messages, transferring funds, modifying access permissions and deleting records deserve stricter safeguards than drafting an internal summary.
Approvals should show the proposed action, relevant context and expected effects in language a reviewer can understand. The final operation must still be authorized and validated by the application after approval; a human confirmation is not a substitute for software controls.
Monitor behavior and prepare an incident response path
Record important tool invocations, denials, approval decisions and errors while applying appropriate limits to sensitive data retention. Look for unusual sequences, repeated denied actions and unexpected access patterns. Establish a mechanism to suspend tools or revoke credentials when suspicious activity is detected.
Security evaluation should include malformed inputs, untrusted retrieval content, conflicting permissions and integration failures. Keep the testing strategy current as the product gains new capabilities. An agent becomes a different security system whenever it receives an additional powerful tool.
Final thoughts
Conclusion
Secure AI agents depend on ordinary engineering discipline applied to a new kind of interface. Separate information from authority, enforce least privilege, validate every tool action and make consequential behavior reviewable. Autonomy is safest when the surrounding application remains firmly in control.
Common questions
Frequently asked questions
Can a system prompt prevent all prompt injection attacks?
No. Prompts can communicate intended behavior, but secure tool permissions, isolation, validation and adversarial testing are needed to limit the consequences of malicious content.
Does using MCP automatically make AI tools secure?
No. Interoperability does not guarantee appropriate authentication, authorization or safe application behavior. Validate the exact integration and security design.
Explore further