Codelaro
AI

Securing AI agents and connected tools.

The security decisions that matter when AI systems read external information, call tools and take action across business software.

Codelaro Team10 min read (est.)

Article overview

The big picture

A language model that drafts text creates a different risk profile from an agent that can search company records, update tickets, run code or communicate with customers. Once AI has access to business tools, a mistaken interpretation can produce a real-world action. Security architecture must therefore address not only what the model says, but also what the surrounding system permits it to do.

The OWASP GenAI Security Project published its Top 10 for Agentic Applications for 2026 and expanded its generative AI security resources in September 2026. These resources illustrate why traditional access controls, secure software engineering and new model-specific testing all belong in an agentic system’s design.

At a glance

Key takeaways

  • Treat retrieved pages, documents, messages and tool responses as untrusted data.
  • Enforce authorization in code at every tool boundary.
  • Require confirmation for consequential actions and maintain meaningful audit trails.
  • Test hostile instructions, privilege escalation and unexpected tool behavior.

Threat-model the system, not just the language model

An AI application includes prompts, retrieval systems, tools, identity services, logs and external integrations. Each boundary may expose sensitive data or grant unintended authority. Start by mapping which information the system can read, which operations it can perform and which users or services authorize those operations.

Distinguish helpful conversation context from trusted instructions. A document retrieved for summarization should not be allowed to redefine system rules, grant permissions or instruct the assistant to disclose credentials. The application must enforce these boundaries independently of the model’s natural-language behavior.

Treat prompt injection as an architectural risk

Prompt injection occurs when attacker-controlled or otherwise untrusted material tries to influence a model’s behavior beyond its intended role. An agent might encounter malicious text inside a web page, issue description, email or retrieved document. Because language models interpret instructions and content through similar interfaces, relying on a warning in the prompt alone is insufficient.

Reduce the impact of hostile instructions by constraining tool access, validating outputs and separating data from authority wherever possible. Test adversarial examples during development and after significant changes to prompts, tools or retrieval sources. Avoid treating any one filtering technique as a complete defense.

Enforce permissions where tools actually execute

A model should not be able to grant itself permissions by writing persuasive text. Each tool must verify the identity and authority of the requesting user or service. Use narrowly scoped credentials, validate resource ownership and reject actions outside the approved workflow. For example, the fact that an agent can read a customer record should not imply it can issue a refund or modify that customer’s account.

Connected-tool protocols can standardize interactions but do not remove the need for authorization. Review the security requirements of the exact protocol version and transport being deployed, and prevent credential leakage through URLs, logs and inappropriate token forwarding.

  • Avoid broad shared administrator credentials for agent tools.
  • Separate read-only operations from operations that change state.
  • Use short-lived credentials and verify token audiences where applicable.

Put human approval at meaningful risk boundaries

Requiring approval for every low-risk retrieval would make an assistant unusable. Requiring none for high-impact actions would create avoidable exposure. Design approval around consequences: sending external messages, transferring funds, modifying access permissions and deleting records deserve stricter safeguards than drafting an internal summary.

Approvals should show the proposed action, relevant context and expected effects in language a reviewer can understand. The final operation must still be authorized and validated by the application after approval; a human confirmation is not a substitute for software controls.

Monitor behavior and prepare an incident response path

Record important tool invocations, denials, approval decisions and errors while applying appropriate limits to sensitive data retention. Look for unusual sequences, repeated denied actions and unexpected access patterns. Establish a mechanism to suspend tools or revoke credentials when suspicious activity is detected.

Security evaluation should include malformed inputs, untrusted retrieval content, conflicting permissions and integration failures. Keep the testing strategy current as the product gains new capabilities. An agent becomes a different security system whenever it receives an additional powerful tool.

Final thoughts

Conclusion

Secure AI agents depend on ordinary engineering discipline applied to a new kind of interface. Separate information from authority, enforce least privilege, validate every tool action and make consequential behavior reviewable. Autonomy is safest when the surrounding application remains firmly in control.

Common questions

Frequently asked questions

Can a system prompt prevent all prompt injection attacks?

No. Prompts can communicate intended behavior, but secure tool permissions, isolation, validation and adversarial testing are needed to limit the consequences of malicious content.

Does using MCP automatically make AI tools secure?

No. Interoperability does not guarantee appropriate authentication, authorization or safe application behavior. Validate the exact integration and security design.

Explore further

Further reading

All insightsCodelaro insights

Code. Launch. Grow.

Building something? Let’s talk.

Turn what you've learned into a practical digital product. Talk with Codelaro about your goals and next steps.

Code. Launch. Grow.