
To secure an LLM application or AI agent, assume the model can be manipulated and design so that a manipulated model still cannot do much harm. In practice that means treating every input the model reads as untrusted, giving it the fewest permissions possible, enforcing access rules in code rather than in prompts, validating outputs before they reach users or systems, protecting secrets and personal data, and testing the system with deliberate attacks before and after launch.
Use the checklist below when you design, review or audit an AI system. It covers the main risks in the OWASP Top 10 for LLM Applications and the practical controls we apply in production.
Why Is AI Security Different From Ordinary App Security?
Traditional applications follow code paths you wrote. LLM applications follow instructions written in natural language, and the model cannot reliably tell your instructions apart from instructions hidden in the data it reads. That creates new attack paths on top of all the usual ones.
The risk grows with capability, which is why agentic AI needs stricter controls than a simple chatbot. A chatbot that only answers questions can leak information or say something embarrassing. An AI agent that can send emails, update records, issue refunds or run code can be tricked into doing those things. Classic controls such as authentication, encryption and patching still apply; this checklist adds the AI-specific layer.
What Does the OWASP Top 10 for LLM Applications Cover?
The OWASP Top 10 for LLM Applications is a community-maintained list of the most critical security risks in applications built on large language models. Recent editions cover risks such as prompt injection, sensitive information disclosure, supply chain risks, data and model poisoning, improper output handling, excessive agency, system prompt leakage, weaknesses in vectors and embeddings, misinformation, and unbounded consumption of resources. It is a good shared vocabulary for security reviews and vendor questions.
The AI Security Checklist
1. Prompt injection, direct and indirect
Prompt injection is an attack where crafted text makes the model ignore its instructions. Direct injection comes from the user. Indirect injection arrives inside content the model reads: an email, a web page, a PDF, a support ticket or a tool result.
- Treat all retrieved content, documents, emails and tool outputs as data, never as instructions
- Clearly separate system instructions from user and external content in prompts
- Assume injection will sometimes succeed, and limit what a compromised model can do (see items 3 and 4)
- Screen inputs for known injection patterns, but never rely on screening alone
- Test with injected instructions hidden in every content source the system reads
2. Data leakage and access control
- Apply the user's own permissions when retrieving data, so the model only sees what that user may see
- Filter documents by permission before they reach the model, not after; in RAG systems this means permission-aware retrieval
- Keep secrets, credentials and internal-only notes out of system prompts, and assume prompts can leak
- Check model provider terms on retention and training use, and choose settings and regions that fit your obligations
- Prevent one customer's data from appearing in another customer's session
3. Tool permissions and least privilege
Excessive agency means an AI system has more permissions, tools or autonomy than its task needs.
- Give each agent only the tools it needs, with the narrowest scopes possible
- Authorize every tool call in code, using the end user's identity and rights
- Prefer specific tools ("refund order up to a set limit") over general ones ("run SQL")
- Require human approval for irreversible, financial or external-facing actions
- Set rate limits, spending caps and timeouts on tool use
4. Output validation
Improper output handling means passing model output to other systems or users without checks.
- Validate structured outputs against a schema before using them
- Never execute model-generated code, queries or commands without sandboxing and allow-lists
- Encode or sanitize output before rendering it in a browser to prevent script injection
- Check factual answers against sources where accuracy matters, and show citations
- Block or flag outputs that contain personal data, secrets or policy violations
5. Secrets management
- Store API keys and credentials in a secrets manager, never in prompts, code or client apps
- Use separate keys per environment and per service, with the minimum scope
- Rotate keys on a schedule and immediately after any suspected exposure
- Route model calls through your backend, so keys never reach browsers or mobile apps
6. Logging, monitoring and PII redaction
- Log prompts, tool calls, outputs and decisions with enough context to investigate incidents
- Redact or mask personal data and secrets in logs, or restrict and encrypt those logs
- Set retention periods that match your privacy obligations under laws such as GDPR, HIPAA or India's DPDP Act
- Alert on unusual patterns: spikes in usage, repeated refusals, unexpected tool calls or cost jumps
- Keep an audit trail of every action an agent takes on a business system
7. Supply chain and model security
- Use models, libraries and plugins from trusted sources, and pin versions
- Verify downloaded model weights and datasets, and scan dependencies for known vulnerabilities
- Review third-party tools and MCP servers before connecting them to an agent
- Protect fine-tuning data and retrieval indexes against tampering (data poisoning)
- Re-run security and quality tests when a model version changes
8. Resource abuse
- Limit input size, output length, request rates and per-user spend
- Cap agent loops and retries to stop runaway costs
- Monitor cost per user and per feature
9. Red-teaming and testing
Red-teaming means attacking your own system on purpose to find weaknesses before others do.
- Build an attack test set covering injection, jailbreaks, data extraction and tool misuse
- Run it before launch and on every significant change to prompts, models or tools
- Include indirect attacks through every data source and connected system
- Combine automated tests with manual testing by people who think like attackers
- Track findings to closure like any other security issue
How Should You Roll Out These Controls?
You don't need every control on day one, but you need the right ones for the risk. Use this order:
- Map the system. List every input source, data store, tool and output destination the AI touches.
- Rate the impact. For each tool and data source, ask what the worst outcome of misuse would be.
- Cut permissions first. Remove tools and data access the task does not need. This is the cheapest, most effective control.
- Add enforcement in code. Authorization, validation and approval steps that do not depend on the model.
- Add monitoring and logging with PII redaction.
- Red-team before launch, fix what you find, and repeat on every major change.
Risk and Control Summary
| Risk | Typical impact | Primary control |
|---|---|---|
| Prompt injection (direct and indirect) | Model follows attacker's instructions | Least privilege, untrusted-content handling, approvals |
| Sensitive data disclosure | Leaked personal or confidential data | Permission-aware retrieval, output filtering |
| Excessive agency | Unwanted actions on business systems | Narrow tools, code-enforced authorization |
| Improper output handling | Script injection, bad data written to systems | Schema validation, sanitization, sandboxing |
| Supply chain | Compromised model, library or plugin | Trusted sources, pinning, scanning |
| Unbounded consumption | Runaway cost or denial of service | Rate limits, caps, monitoring |
How Aaga Helps
Aaga builds security into AI systems from the first design session: least-privilege tool design, permission-aware retrieval, validation layers, audit logging and red-team test sets that run on every release. Our application security team reviews existing LLM apps and agents against this checklist, and our AI agent development work applies it by default. If your agents connect to business systems through MCP, our guide to the Model Context Protocol covers the security points specific to it.
Want an outside review of your AI application? Contact Aaga to scope a security assessment.

