
The attacks Arena runs against a model that is not an LLM depend on what the model does. This article lists each model task type Arena tests, the attack techniques it runs, and any details specific to that task. Each technique is explained once in Non-LLM Attack Categories, along with how to remediate the results. For how Adaptive Tests work, see AI Arena Overview.
Vision Models
Models that take images, video, or scanned documents as input.
Image Classification
Image classification is a task where a model takes an image as input and assigns it to one of several predefined categories. The model does not generate or describe the image: it outputs a single label (or a ranked list of labels) from a fixed set. Common examples include identifying objects in photos, classifying medical scans by diagnosis, detecting defects in manufacturing imagery, categorizing documents by type from scans, and flagging anomalies in aerial or satellite imagery.
- Techniques: Whole-Image Perturbation, Localized Patch Attack
- Specific to this task:
- Arena reports both untargeted and targeted results.
Object Detection
Object detection is a task where a model takes an image as input and identifies every instance of an object within it. For each one, it predicts both a bounding box (location and size) and a class label. Common examples include detecting vehicles in traffic footage, identifying defects in product inspection images, locating people or objects in security camera footage, and finding relevant regions in scanned documents.
- Techniques: Whole-Image Perturbation, Localized Patch Attack
- Specific to this task:
- The attack targets both where bounding boxes are placed and the class label assigned to each detected object.
Depth Estimation
Depth estimation is a task where a model takes an image as input and predicts, for every pixel, how far the corresponding point in the scene is from the camera. The output is a dense, per-pixel distance map rather than a single label or a handful of boxes. Common examples include obstacle detection for autonomous vehicles and mobile robots, 3D scene reconstruction, augmented-reality object placement, and clearance and collision checks in warehouse or industrial automation.
- Techniques: Whole-Image Perturbation, Localized Patch Attack
Semantic Segmentation
Semantic segmentation is a task where a model takes an image as input and assigns a class label to every pixel. The result is a dense map of the scene rather than a single label or a handful of boxes. Common examples include parsing driving footage into road, sidewalk, vehicle, and pedestrian regions, mapping land cover in aerial and satellite imagery, segmenting tissue types in medical scans, and identifying structural regions in scanned documents.
- Techniques: Whole-Image Perturbation, Localized Patch Attack
SAM Mask Generation
SAM-style mask generation is a task where a model takes an image plus a prompt (typically a bounding box drawn around one object of interest) and returns a precise pixel mask for that single object. Unlike semantic segmentation, it is class-agnostic. It segments “the object there” without identifying its class. It also produces one instance mask per prompt rather than a label map of the full scene. Common examples include isolating a specific document field or ID photo for redaction, extracting a region of interest a radiologist has boxed for measurement, cutting out a flagged item in content-moderation pipelines, and precisely outlining a single inspected component in industrial imagery.
- Techniques: Whole-Image Perturbation, Localized Patch Attack
- Specific to this task:
- The box prompt is never attacked. It stays fixed and correct, so any change in the mask comes from the image perturbation.
Keypoint Detection
Keypoint detection (pose estimation) is a task where a model takes an image of a person as input and predicts the pixel location of a fixed set of body joints (nose, shoulders, elbows, wrists, hips, knees, ankles) rather than a class label or a region. Common examples include fall detection for elder care and lone-worker safety monitoring, gesture and activity recognition, ergonomic assessment in workplace safety, and physical-therapy or fitness form tracking.
- Techniques: Whole-Image Perturbation, Localized Patch Attack
Face Recognition
Face recognition (verification) is a task where a model takes a photo of a face and produces a numeric embedding, then decides whether that embedding matches a stored reference embedding for a claimed identity. It compares two faces for similarity rather than classifying a face into fixed categories. Common examples include biometric login for mobile and online banking, identity verification for account recovery and SIM-swap requests, physical access control at secure facilities, and visitor screening against watchlists.
- Techniques: Whole-Image Perturbation, Localized Patch Attack
- Specific to this task:
- Each test uses the attacked photo, a second real photo of the same person, and a real photo of a different person.
- Arena reports the raw similarity between faces rather than a match or no-match decision, because no single threshold applies across all systems.
Video Classification
Video classification is a task where a model takes a short video clip (a fixed-length sequence of frames sampled across it) as input and assigns the whole clip to one action or activity category from a fixed set. It produces a single label per clip, like image classification, but reasons over motion across frames rather than a single still image. Common examples include recognizing unsafe actions in workplace safety footage, flagging suspicious behavior in security camera clips, classifying activity in sports and fitness footage, and verifying that a required procedure was performed correctly in industrial process-monitoring video.
- Techniques: Whole-Image Perturbation, Localized Patch Attack
- Specific to this task:
- The perturbation is applied to every frame of the clip.
- The patch stays in the same location in every frame, like a single physical sticker.
Document Visual Q&A
Document visual question answering is a task where a model takes a scanned document image and a natural-language question about it, and generates a free-text answer directly from the pixels. There is no separate OCR step and no fixed set of answer choices. Common examples include automated invoice and receipt processing (extracting totals, dates, and vendor names), extracting fields from forms and applications, answering questions about scanned contracts or records, and pulling specific values out of scanned technical or regulatory documents.
- Techniques: Whole-Image Perturbation, Localized Patch Attack
- Specific to this task:
- Because the model generates free text, Arena scores each answer with token-level F1 (word overlap with a reference answer) rather than exact match.
Visual Question Answering
Visual question answering is a task where a model takes an ordinary photo and a natural-language question about it, and picks an answer from a large, fixed answer vocabulary. Unlike document visual Q&A, this is a classification problem over real-world photos rather than free-text generation over scanned documents. Common examples include automated photo-based inspection and triage (asking a fixed question about a submitted photo), visual content moderation, accessibility tools that answer questions about images for visually impaired users, and photo-based intake for customer support.
- Techniques: Whole-Image Perturbation, Localized Patch Attack
- Specific to this task:
- Arena scores exact-match accuracy against the correct answer.
- Patch sizes also include 0.5% of the image area.
Document (Visual) Retrieval
Document (visual) retrieval is a task where a model embeds a search query and a document page image directly from pixels (no OCR step) and scores how relevant that page is to the query with a continuous similarity score, rather than assigning a label or extracting an answer. Common examples include searching large scanned-document archives by natural-language query, powering retrieval-augmented generation (RAG) systems over document images, and surfacing relevant pages in e-discovery or compliance document review.
- Techniques: Whole-Image Perturbation
- Specific to this task:
- Arena measures the change in the model's own relevance score for the page, a continuous value rather than a match or no-match decision.
Text and Language Models
Models that read, rank, label, or generate text.
Reranking
Reranking is a task where a model scores how relevant each of several candidate passages is to a search query, and those scores determine the order in which results are shown. It is the second-stage step in most modern search and RAG (retrieval-augmented generation) pipelines that decides which of the already-retrieved passages is actually best. Common examples include reordering search results after a fast initial retrieval pass, ranking candidate passages for a RAG system's context window, and surfacing the most relevant support articles or knowledge base entries for a query.
- Techniques: Keyword Injection
- Specific to this task:
- Arena measures how far the genuinely relevant passage drops in the ranking.
Extractive Summarization (Sentence Selection)
Extractive summarization is a task where a model scores every sentence in a document by how central it is to the document's main topic (its cosine similarity to the document's overall centroid rather than to a search query), then selects the highest-scoring sentences as the summary, copying them rather than paraphrasing. Unlike reranking, there is no external query, so importance is measured entirely within the document. Common examples include auto-generating bullet-point highlights for news articles or reports, creating quick-read document previews, and pulling out key sentences for a case file or incident record.
- Techniques: Keyword Injection
- Specific to this task:
- Arena measures whether a sentence moves into, or out of, the document's top three most central sentences.
Named Entity Recognition
Named entity recognition (NER) is a task where a model scans free text, flags spans that refer to real-world entities (typically people, organizations, and locations), and identifies the type of each one. Common examples include automatically redacting personally identifiable information (PII) before a document is shared externally, extracting the parties and organizations named in contracts or filings, and tagging entities for search and indexing.
- Techniques: Invisible Character Injection
- Specific to this task:
- The target entity type is a person by default.
Text Classification
Text classification is a task where a model reads a piece of text and assigns it to one of several predefined categories. Common examples include classifying customer support messages by urgency or topic, labeling articles or reviews by sentiment, routing incoming requests to the right team, and flagging communications for policy review.
- Techniques: Synonym Substitution
Extractive Q&A
Extractive question answering (Q&A) is a task where a model reads a passage of text and identifies the contiguous span of words in that passage that best answers a given question. The answer must come word for word from the source document. Common examples include extracting terms from contracts, pulling specific figures from technical or policy documents, and surfacing relevant clauses from lengthy agreements.
- Techniques: Synonym Substitution
- Specific to this task:
- Arena changes words in the passage. The question stays fixed.
Table Question Answering
Table question answering is a task where a model reads a structured table (rows, columns, and headers) alongside a natural-language question, and answers by selecting cells and, where needed, applying an aggregation like sum, average, or count. This differs from extractive Q&A, which reads free-form prose. Common examples include answering questions about spreadsheet exports, financial and operational reporting dashboards, and structured data extracted from filings or logs.
- Techniques: Synonym Substitution
- Specific to this task:
- Arena changes words in table cells and headers. The question stays fixed.
- Numeric values are never changed.
Abstractive Question Answering
Abstractive question answering is a task where a model reads a passage and a question, then generates a free-text answer in its own words. This differs from extractive Q&A, which must copy a contiguous span word for word from the source. Common examples include conversational Q&A assistants grounded in a document, customer support bots that answer from a knowledge base, and Q&A over policies or manuals where the natural phrasing of the answer does not appear verbatim in the source text.
- Techniques: Synonym Substitution
- Specific to this task:
- Arena changes words in the passage. The question stays fixed.
- Arena scores answers with BERTScore, a measure of semantic similarity, rather than exact overlap.
Abstractive Summarization
Abstractive summarization is a task where a model reads a long document and generates a shorter summary in its own words, synthesizing and paraphrasing rather than copying sentences verbatim. Common examples include condensing lengthy reports into executive briefs, summarizing technical documents for non-specialist audiences, and generating short abstracts from long-form content.
- Techniques: Synonym Substitution
- Specific to this task:
- Substitutions must also pass a part-of-speech match and a sentence-level coherence check.
Translation
Translation is a task where a model reads text in one language and generates a fluent translation in another. Arena's evaluation translates English to French. Common examples include translating customer communications and support tickets, localizing product documentation and contracts, and translating foreign-language records for a review team that doesn't read the source language.
- Techniques: Synonym Substitution
- Specific to this task:
- Quality is measured with chrF, a standard machine-translation metric.
Text-to-Image Safety Filter
Text-to-image generation is a task where a model creates an image from a text prompt. Production systems pair the generator with a safety classifier that screens each incoming prompt and blocks anything flagged as unsafe before an image is generated. This task tests the opposite failure from a typical jailbreak: starting from a prompt that is genuinely benign and expected to pass, can small, meaning-preserving word substitutions make the safety classifier wrongly flag it as unsafe? Common examples include consumer creative tools, marketing and content-generation platforms, and any product that places an image generator behind a safety gate.
- Techniques: Synonym Substitution
- Specific to this task:
- Arena starts from a prompt confirmed to be benign and checks whether word substitutions make the safety classifier wrongly flag it as unsafe.
- Each query is a full image-generation call, so Arena tests only three budgets on a small sample of prompts: 1.0, 0.8, and 0.6.
Audio and Speech Models
Models that listen to audio or produce speech.
Audio Classification
Audio classification (keyword spotting) is a task where a model listens to a short audio clip and assigns it to one of a fixed set of categories. The categories are most commonly single spoken command words, such as “yes,” “stop,” and “go,” plus catch-all “silence” and “unknown” categories. This differs from Automatic Speech Recognition, which transcribes continuous speech word for word. Common examples include voice-activated wake words and hands-free device commands, hands-free equipment control in industrial and field settings, accessibility voice controls, and simple voice-menu navigation.
- Techniques: Adversarial Audio Perturbation
- Specific to this task:
- Arena reports both untargeted and targeted results.
Automatic Speech Recognition
Automatic Speech Recognition (ASR) is a task where a model converts spoken audio into written text. Common examples include transcribing voice commands, converting recorded meetings into searchable text, voice-controlled interfaces, IVR systems, and dictation tools.
- Techniques: Adversarial Audio Perturbation
Text-to-Speech
Text-to-speech (TTS) is a task where a model reads written text and synthesizes it as spoken audio. Robustness is judged with a round trip: the synthesized speech is fed to an independent speech-recognition model, and the resulting transcript is compared with the original text. If a machine listener can no longer recover what was written, the synthesized audio itself has been corrupted. Common examples include voice assistants reading messages aloud, accessibility screen readers, automated phone and IVR announcements, and audiobook or narration generation.
- Techniques: Invisible Character Injection
Tabular Models
Models that work on rows of structured data. These are the only task types tested for training-time attacks.
Tabular Classification
Tabular classification is a task where a model takes a row of structured data (numeric and categorical features organized in columns) and assigns it to one of several predefined categories. Common examples include classifying applicants into risk tiers, labeling records as legitimate or anomalous, scoring items by the likelihood of a given outcome, and routing entities based on their feature profiles.
- Techniques: Feature Perturbation, Data Poisoning
- Specific to this task:
- Untargeted poisoning flips the labels on a subset of training rows.
- Targeted poisoning plants a hidden trigger tied to a class the attacker chooses.
Tabular Regression
Tabular regression is a task where a model takes a row of structured data and predicts a continuous numeric value rather than a discrete category. Common examples include estimating probabilities or scores on a continuous scale, forecasting demand or consumption, and predicting expected costs or losses.
- Techniques: Feature Perturbation, Data Poisoning
- Specific to this task:
- Feature perturbation tries to maximize the error in the predicted value rather than flip a class.
- Untargeted poisoning shifts the target values on a subset of training rows.
- Targeted poisoning plants a hidden trigger tied to a value the attacker chooses.





