Find the insights and best practices about our product.
AI Arena Overview

Introduction

As AI adoption grows, every model in your environment becomes a potential target for adversarial attacks. A model that can be jailbroken, manipulated, or made to leak data puts every system built on it at risk. Arena gives security teams a central view of the models discovered across your AI Bills of Materials (BOMs) and shows where each one is vulnerable.

Arena runs automated and manual penetration tests that simulate real-world adversarial techniques. The results reveal weaknesses before an attacker can exploit them. Each tested model has a Model Threat Analysis page that breaks down its security assessment. The page ranks the attack categories where exploits succeeded and maps each one to the known weakness behind it.

This article explains how to find a model in Arena, how to read its Model Threat Analysis page, and how to examine individual successful attacks.

AI Models List

The Arena page displays AI models as cards in a grid. Each model card includes:

  • Model icon or logo: Identifies the model at a glance.
  • Model name: The full model name, which links to the model's Model Threat Analysis page.
  • Vulnerability likelihood: The model's overall risk score, shown as a percentage.
  • AI Systems: The number of AI Systems that contain the model.
  • Severity badge: The highest risk level found during testing.
  • Completion status: When the most recent penetration test was conducted.

To find a specific model, type any part of its name into the search bar at the top of the page. The sort dropdown orders models by highest severity or by most recently updated. The Filters button narrows results by AI System or by the date penetration tests were conducted.

A model's status depends on its testing progress. Models that have finished testing can be selected to open their Model Threat Analysis page. Models still in testing appear in the grid but can't be selected until results are available. Models found in your Bills of Materials but not yet tested also appear in the grid, with no test data until testing begins.

Model Threat Analysis

Selecting a tested model opens its Model Threat Analysis page, which gives a detailed assessment of the model's security posture. Use it to understand the model's vulnerabilities, analyze potential attack vectors, and track security risks across your AI Systems.

The page has two modules. Findings Overview summarizes the latest test results. Vulnerable Attack Categories ranks the attack categories where tested exploits succeeded.

How LLM and Non-LLM Assessments Differ

The page layout depends on the type of model. LLMs go through two types of testing: Fixed Tests and Adaptive Tests. Non-LLM models, such as traditional machine learning and deep learning models, go through Adaptive Tests only. The page shows results for each test type a model received, so a Non-LLM model's page shows Adaptive Test results only.

Fixed Tests send a curated, version-controlled library of adversarial prompts to the LLM. Each prompt is a standalone request that targets a specific type of safety failure. The attack success rate is the share of prompts in each category that produced a policy-violating response. Because the prompt library is fixed and versioned, results are reproducible and comparable across model versions and over time.

Adaptive Tests adjust their approach based on how the model responds. For LLMs, an automated adversary holds a multi-turn conversation aimed at a specific harmful objective. If the model refuses, the adversary changes its framing, persona, or tactic on the next turn, much like a real attacker working across a full conversation. These tests focus on harmful responses and misinformation, with objectives relevant to the industries where the LLM is deployed. For Non-LLM models, each attack makes a series of attempts against the model. It uses the model's responses to refine its input at every step until the model's prediction changes.

Findings Overview

Findings Overview gives a high-level summary of the model's risk profile. It includes:

  • Last Scanned: The date and time of the most recent security evaluation.
  • Known Weaknesses: The number of known weaknesses mapped to the model's attack categories. Each vulnerable attack category card shows its mapped weakness.
  • Vulnerability Likelihood: The model's overall risk score, shown as a percentage. It gives a measurable indicator of how susceptible the model is to adversarial techniques. The score is the average attack success rate across all tested attack categories, and each category counts equally.
  • Categories Tested: Each attack category tested against the model, with the number of successful attacks out of the total attempted. These outcomes show which types of threats the model is most vulnerable to. LLM categories include generating harmful responses, generating misinformation, generating hallucinations, susceptibility to jailbreaks, encoding attacks, prompt injection, data leakage, and enabling cyberattacks. Non-LLM models are tested against different categories, such as Inference-Time Vulnerabilities.
  • AI Systems Containing This Model: Each AI System that includes the model. Each name links to that AI System's details page, so you can trace a vulnerability from the model to every AI System that uses it.
  • BOMs Containing This Model: Each Bill of Materials that includes the model. Each name links to that Bill of Materials.

The Download PDF button downloads the full analysis as a Model Threat Analysis Report for further review or documentation. The Attach PDF Report To AI System button attaches that report to an AI System.

Prioritizing Vulnerable Attack Categories

Vulnerable Attack Categories shows the attack categories where tested exploits succeeded. These are confirmed vulnerabilities that Arena was able to exploit through scanning and penetration testing. They show the security gaps that adversaries could use against the model. Categories appear as cards in a full-width grid, ranked from highest attack success rate to lowest. The most exploitable categories appear first.

Each card includes:

  • Attack success rate: Two rings showing the percentage of successful attacks in the category, one for Fixed Tests and one for Adaptive Tests.
  • Category name: The attack category tested.
  • Weakness: The known weakness mapped to the attack category. The weakness name opens its OWASP entry in a new tab. LLM weaknesses link to the OWASP Top 10 for LLMs. Non-LLM weaknesses link to the OWASP ML Security Top 10. A weakness without a matching OWASP entry appears as plain text.
  • Date: The date the exploit was discovered.

A dashed grey ring means that test type didn't run for the category. It doesn't mean a 0% success rate. Non-LLM models only run Adaptive Tests, so their Fixed Tests ring always appears dashed.

Select a card to open its Vulnerability Details.

Vulnerability Details

Vulnerability Details gives an in-depth look at how the model was compromised in a single attack category. The header shows when the results were last updated.

The Overview section includes:

  • Category: The attack category.
  • Weakness: The known weakness mapped to the category, linked to its OWASP entry.
  • Description: What the vulnerability is.
  • Remediation: Guidance for addressing the vulnerability.
  • Fixed Tests and Adaptive Tests: Rings showing the attack success rate for each test type. A test type that didn't run shows N/A. Fixed Tests always show N/A for Non-LLM models.
  • Attacks Successful: The number of successful attacks out of the total attempted.

The Successful Attacks section lists each instance where an attack exploited the model. Use the arrows to move between attacks. Each entry includes:

  • Date: The date of the test.
  • Algorithm: The attack algorithm used.
  • Objective: What the attack aimed to achieve.

Expand Evidence to review supporting documentation, such as conversation logs, that shows how the attack was carried out. A disclaimer notes that evidence may contain harmful information. Unsuccessful attempts aren't included, so the section only shows confirmed vulnerabilities that need mitigation.


Please note: A model's Arena data does not depend on whether the BOM's Vulnerability Assessment succeeded.

Did this answer your question?