The OWASP Top 10 for LLM Applications is OWASP's ranked list of the ten most critical security risks in large language model and generative AI applications, now in its 2026 edition, published in August 2026. Prompt injection holds the top spot for a second year running, sensitive information disclosure stays at number two, and excessive agency jumps from sixth to third as agentic AI puts more tools and autonomy directly in a model's hands. This guide explains all ten risks with realistic scenarios and mitigations, then covers agentic AI security, red-teaming and a checklist for shipping an LLM application safely.
What Is the OWASP Top 10 for LLM Applications, and Who Maintains It
OWASP, the nonprofit best known for the original OWASP Top 10 web application risks, maintains this list through its OWASP Gen AI Security Project, the umbrella that grew out of the original OWASP Top 10 for Large Language Model Applications project listed on owasp.org. The list is community-built: volunteer security practitioners, AI engineers and researchers propose, debate and vote on risks and their ranking, and OWASP publishes the result as a free, vendor-neutral reference, the same model that made the original web Top 10 a shared vocabulary between security and engineering teams.
The project has released three major versions: an initial 2023 list, a substantially revised 2025 edition published in November 2024, and the current 2026 edition published in August 2026. For the first time, the 2026 edition incorporated analysis of real-world documented AI security incidents alongside practitioner voting, rather than relying on expert opinion alone. The list is not a compliance checklist and certifies nothing; treat it as a shared risk vocabulary for threat modeling, architecture review and red-team scoping, the same way engineering teams already use the original OWASP Top 10 for conventional web applications.
The OWASP Top 10 for LLM Applications for 2026
Each entry below reflects the current, 2026 numbering and naming, verified against OWASP's own project pages.
| Rank | Risk | What It Means | Primary Mitigations | | --- | --- | --- | --- | | LLM01 | Prompt Injection | Untrusted input changes model behavior against the application's intent | Treat all input as untrusted; separate instructions from data; constrain what actions output can trigger | | LLM02 | Sensitive Information Disclosure | The model exposes personal data, credentials or proprietary information in its output | Redact sensitive data before it reaches the model; filter output; enforce least-privilege data access | | LLM03 | Excessive Agency | The model has more permission, tools or autonomy than the task needs | Scope tool permissions narrowly; allow-list actions; require human approval for irreversible steps | | LLM04 | Supply Chain | Vulnerable or malicious third-party models, datasets, packages or plugins enter the application | Vet model and package provenance; pin and scan dependencies; maintain a software bill of materials | | LLM05 | Data and Model Poisoning | Manipulated training, fine-tuning or embedding data introduces backdoors or bias | Vet training and fine-tuning data sources; monitor for anomalous behavior after any retraining | | LLM06 | Unbounded Consumption | Uncontrolled resource use drives cost, denial of service or model extraction | Rate-limit and budget per user or tenant; cap input and output size; monitor for abuse patterns | | LLM07 | Misinformation | Confident but false output is presented and trusted as fact | Ground answers in retrieval with sources; flag low-confidence output; keep a human in the loop | | LLM08 | Hidden Context Exposure | Content assembled into the model's context, instructions, documents, memory, tool output, leaks to users | Treat everything in context as potentially exposable; keep secrets and other tenants' data out of it | | LLM09 | Vector and Embedding Weaknesses | Weaknesses in how data is embedded, stored or retrieved let attackers manipulate or extract information | Isolate embeddings per tenant; validate and monitor retrieval; control who can write to a vector store | | LLM10 | Improper Output Handling | Model output reaches downstream systems without validation, enabling injection or code execution | Validate and sanitize output before it reaches a shell, database, browser or API; never execute it directly |
The Five Risks That Cause the Most Damage in Production
All ten risks matter, but five of them account for most of the incidents Agentixly sees in client architecture reviews.
Prompt Injection (LLM01)
Scenario: An AI email assistant reads incoming messages and can draft replies or file them automatically. An attacker sends a message containing hidden text: instructions telling the model to forward every future email matching "invoice" to an external address. The model treats that hidden text as a command rather than as message content.
Impact: Data exfiltration, unauthorized actions taken in the user's name, and in agentic settings, an open-ended loss of control over what the system does next.
Mitigation: Separate trusted instructions from untrusted content structurally where the model architecture allows it, flag instruction-like patterns in ingested content, constrain what actions a model can trigger regardless of what it was told to do, and log every action for anomaly review. See the guide to enterprise prompt engineering for how to structure prompts that are harder to override.
Excessive Agency (LLM03)
Scenario: A support bot can issue refunds up to a set limit to resolve tickets without waiting on a human. A prompt injection, or simply a reasoning error under an unusual combination of inputs, causes it to approve a refund outside policy or chain together tool calls it was never meant to combine.
Impact: Direct financial loss, policy violations, and actions that are difficult or impossible to reverse once executed.
Mitigation: A narrow, explicit allow list of tools and parameters per agent, a human approval gate for anything above a defined risk threshold, and full audit logging of every tool call together with the reasoning that triggered it.
Sensitive Information Disclosure (LLM02)
Scenario: A multi-tenant SaaS support assistant retrieves from a shared knowledge base without strict per-tenant filtering. A user asks a carefully phrased question and receives a fragment of another customer's data inside the answer.
Impact: Contractual and regulatory exposure, loss of customer trust, and potentially a formal breach notification obligation depending on what leaked.
Mitigation: Enforce tenant isolation at the data-access layer, not only in the prompt, redact sensitive fields before they enter a prompt or a fine-tuning set, and test explicitly for cross-tenant leakage before launch. The RAG architecture guide covers how to isolate tenant data at the retrieval layer itself.
Unbounded Consumption (LLM06)
Scenario: A public-facing chatbot has no per-session or per-user limit on request length or frequency. A script sends thousands of maximum-length requests, running up the model bill and degrading service for real customers at the same time.
Impact: Direct financial loss, denial of service for legitimate users, and, in some designs, enough repeated querying to reconstruct proprietary prompt behavior.
Mitigation: Hard rate limits and budgets per user, tenant and API key, input and output size caps, and alerting when consumption patterns deviate from baseline. Cost controls and security controls are the same work here; see the guide to adding AI features to a SaaS product for the architecture that keeps both in check.
Hidden Context Exposure (LLM08)
Scenario: A coding assistant's context window includes its system prompt, retrieved internal documentation, and the output of a previous tool call. A user asks it to "repeat everything above this message" or crafts a query that causes the model to quote from that assembled context.
Impact: Exposure of proprietary instructions, internal documentation, or another user's data that was never meant to be visible, undermining security and competitive position at once.
Mitigation: Assume everything placed in context can eventually surface to the user, keep secrets and cross-tenant data out of shared context entirely, and test for context leakage the same way you would test for any information-disclosure bug.
What Changed Between the 2025 and 2026 Lists, and Why It Matters
The ranking shift tells you where real-world risk is moving, which matters more than memorizing ten names.
| Risk | 2025 Rank | 2026 Rank | Change | | --- | --- | --- | --- | | Prompt Injection | 1 | 1 | No change | | Sensitive Information Disclosure | 2 | 2 | No change | | Excessive Agency | 6 | 3 | Up 3, the largest rise | | Supply Chain | 3 | 4 | Down 1 | | Data and Model Poisoning | 4 | 5 | Down 1 | | Unbounded Consumption | 10 | 6 | Up 4 | | Misinformation | 9 | 7 | Up 2 | | System Prompt Leakage, renamed Hidden Context Exposure | 7 | 8 | Renamed and broadened | | Vector and Embedding Weaknesses | 8 | 9 | Down 1 | | Improper Output Handling | 5 | 10 | Down 5, the largest fall |
Excessive agency's rise from sixth to third is the headline change, and it tracks what changed in production systems between the two editions: far more applications now give a model direct access to tools, APIs and databases instead of only generating text for a human to act on. Improper output handling falling from fifth to tenth suggests the industry has gotten measurably better at the basic discipline of validating model output before it reaches a shell or a database query. The System Prompt Leakage to Hidden Context Exposure rename matters on its own: it is not just a name change, it widens the risk from "do not leak the system prompt" to "assume anything assembled into context, retrieved documents, memory, tool responses, can leak."
Securing Agentic AI: Tool Permissions and Human-in-the-Loop Design
Excessive agency and hidden context exposure both point to the same shift: once a model can call tools, your security boundary is no longer the chat response, it is every action the model can take on the application's behalf. For systems where multiple specialized agents coordinate rather than a single model handling everything, see the guide to multi-agent systems for the architecture patterns and the added coordination risks they introduce.
Three design rules keep agentic features contained:
- Scope credentials to the agent, not to a human admin. An agent should hold the narrowest API key or service account that lets it do its job, never a copy of a person's full-access credentials.
- Allow-list actions instead of trying to block bad ones. Enumerate exactly what an agent may do; anything not on the list is refused by default, rather than trying to anticipate every harmful request in advance.
- Gate irreversible actions on human approval. Refunds, deletions, outbound messages to customers and any spend above a defined threshold should pause for a person, no matter how confident the model's output looks.
Example: An Allow-Listed, Schema-Validated Tool Dispatcher
The snippet below illustrates the pattern: every tool call is checked against an allow list, validated against a schema, logged, and routed to a human for anything high-risk before it executes.
type ToolName = "searchOrders" | "issueRefund" | "sendEmail";
interface ToolCall {
name: ToolName;
args: Record<string, unknown>;
requestedBy: string;
}
const TOOL_SCHEMAS: Record<ToolName, (args: Record<string, unknown>) => boolean> = {
searchOrders: (args) => typeof args.customerId === "string",
issueRefund: (args) =>
typeof args.orderId === "string" && typeof args.amountCents === "number" && args.amountCents > 0,
sendEmail: (args) => typeof args.to === "string" && typeof args.body === "string",
};
const REQUIRES_HUMAN_APPROVAL: ToolName[] = ["issueRefund", "sendEmail"];
async function dispatchToolCall(call: ToolCall): Promise<ToolResult> {
const validate = TOOL_SCHEMAS[call.name];
if (!validate) throw new Error(`Tool "${call.name}" is not on the allow list`);
if (!validate(call.args)) throw new Error(`Arguments for "${call.name}" failed schema validation`);
await logToolCall(call);
if (REQUIRES_HUMAN_APPROVAL.includes(call.name)) {
return queueForHumanApproval(call);
}
return executeTool(call);
}
This dispatcher rejects anything not on the allow list, validates arguments before execution, logs every call for audit, and routes high-risk actions to a human before they run. That single pattern mitigates excessive agency, improper output handling and a large share of prompt injection's real-world impact together, since an injected instruction can only ever request an allowed, schema-valid, logged action.
How to Test and Red-Team an LLM Application
Security testing for an LLM application needs its own process, layered on top of, not replacing, the testing your product already gets.
- Define abuse cases specific to your application. Data exfiltration through retrieval, unauthorized tool calls, cross-tenant leakage and resource exhaustion, not only generic public jailbreak prompts.
- Build or adopt an automated adversarial test suite. Run known prompt-injection and jailbreak patterns against every release, the same way a SAST scan runs against every pull request.
- Run manual red-team sessions. Use people who did not build the feature, and let them try to misuse the product's own tools, uploaded files and linked content, not only the chat box.
- Triage findings like any other security bug. Severity, an owner, a fix, and a retest, with the same service-level discipline as a penetration test finding.
- Regression-test before every release that touches prompts, retrieval sources, tools or the underlying model, not only at launch.
- Monitor production for the same patterns you tested for. Red-teaming finds what you thought to test; monitoring catches what a live attacker tries next.
OWASP's own GenAI Red Teaming Guide covers threat modeling and test categories for generative AI systems in more depth than fits here, and it is a reasonable structure to adopt rather than build from scratch.
A Secure Architecture Checklist for LLM Applications
Work through this list during design, not after a finding forces the conversation:
- Untrusted input is never treated as instructions; user, document and tool content is structurally separated from system instructions.
- Every tool or function the model can call is on an explicit allow list with schema-validated arguments.
- Actions with financial, legal or irreversible impact require human approval before execution.
- Retrieval and memory are filtered by tenant at the data layer, not only in the prompt.
- Output bound for a shell, database query, browser or another API is validated and sanitized, never executed directly.
- Rate limits and budgets are enforced per user, tenant and API key, with alerting on deviation from baseline.
- Prompts, tool schemas and model or provider choices are versioned, logged and covered by an evaluation suite that runs before every release.
- An incident response plan names who gets paged when a model takes an unexpected or harmful action, mirroring the plan already in place for conventional security incidents.
These controls sit on top of, not instead of, conventional API security; see the API security best practices guide for the request-level protections every endpoint still needs regardless of what is calling it. For governance and risk-management structure above the technical controls, NIST's Generative AI Profile extends the broader NIST AI Risk Management Framework specifically to generative AI systems.
How Agentixly Approaches LLM and AI Agent Security
Agentixly's cybersecurity team applies the same discipline to LLM applications that it applies to conventional penetration testing, adapted for a threat model that includes the model itself. For clients building AI features on the SaaS side, that review typically happens alongside the feature work rather than after it ships:
- Threat model the specific application. What data can the model see, what tools can it call, and what is the actual blast radius if it is manipulated or simply wrong.
- Review the architecture against the OWASP Top 10 for LLM Applications. Walk each of the ten risks against the real design, not a generic checklist.
- Harden the allow-list and approval layer. Build or review the tool dispatcher, schema validation and human-approval gates before any agentic feature reaches production.
- Red-team before launch. Combine automated adversarial testing with manual sessions targeting the application's specific abuse cases.
- Set up continuous monitoring and periodic re-testing. LLM security is not a one-time audit; new models, new prompts and new tools each reopen the question.
No engagement guarantees a specific outcome. What this process delivers is a documented, current answer to the question every security-conscious customer or investor now asks: what happens when this model is wrong, manipulated, or given an instruction it should refuse.
The Bottom Line
The OWASP Top 10 for LLM Applications gives security and engineering teams a shared, current vocabulary for AI risk, and the 2026 edition's biggest signal is that agentic AI, models with tools and autonomy, has become the primary attack surface, not an edge case. Prompt injection still opens most incidents, but excessive agency decides how much damage they cause once they succeed. Build the allow-list and approval layer before you grant a model its first tool, not after.
If you are shipping an LLM feature or an AI agent and want a security review before it reaches production, talk to Agentixly about how we approach LLM and agentic AI security.