
Arena tests models that are not LLMs, such as traditional machine learning and deep learning models, against two attack categories. This article explains each category and how to remediate it, then describes each attack technique Arena uses. For the techniques Arena runs against each kind of model, see Non-LLM Testing by Task Type. For LLMs, see LLM Attack Categories.
Non-LLM models are tested with Adaptive Tests only, so their Fixed Tests results always show as not run. For how Adaptive Tests work, see AI Arena Overview.
Reading Category Results
Each category appears in Categories Tested and, when an attack succeeds, in Vulnerable Attack Categories on the model's Model Threat Analysis page.
- Attack success rate: The share of attacks in the category that got through. Lower is better.
- Any rate above 0%: At least one real attack worked against the model in that category.
Inference-Time Attack
Checks whether small, carefully crafted changes to an input (often invisible to a person) can push your model into a wrong prediction while it is in use.
How to Remediate
In short: Implement input preprocessing to detect adversarial perturbations, apply model-hardening techniques such as adversarial training, input smoothing, and ensemble verification, and add output validation to catch anomalous predictions.
- Harden the model inference pipeline by applying input preprocessing defenses and model-level robustness techniques.
Recommended Defenses
Input Preprocessing
Apply input smoothing (e.g., Gaussian blur, median filtering) to reduce the effectiveness of
adversarial perturbations.
Use feature squeezing to reduce the input space and collapse adversarial examples back to
their benign counterparts.
Normalize and sanitize all inputs to conform to expected data distributions.
Model Robustness
Use adversarial training to improve the model's inherent robustness by incorporating
adversarial examples into the training process, teaching the model to correctly classify
perturbed inputs.
Use ensemble methods to cross-verify predictions across multiple models or model variants.
Flag cases where ensemble members disagree significantly.
Apply randomized smoothing to certify robustness within a defined perturbation radius.
Implement confidence thresholding to reject predictions below a minimum confidence score,
returning an 'uncertain' or 'requires review' status.
Apply rate limiting and query budgeting on inference endpoints to prevent black-box
adversarial probing.
Monitor query patterns for signs of systematic probing (e.g., many similar inputs with small
variations). - Perform input validation to detect adversarial examples by checking for statistical anomalies, out-of-distribution inputs, unusual perturbation patterns (e.g., Lp-norm deviations), or inputs that deviate significantly from expected data distributions. Use dedicated adversarial detection models or statistical tests where feasible.
- Perform output validation to detect and flag anomalous predictions, sudden confidence shifts, class-boundary oscillations, or outputs that are inconsistent with expected behavior before results are returned to downstream systems or users.
Training-Time Attack
Checks whether poisoned training data can change how your model behaves once it's trained, either by degrading it overall or by planting a hidden trigger.
How to Remediate
In short: Implement training data validation and provenance tracking, apply robust training techniques to resist data poisoning, and perform post-training model validation to detect backdoors or compromised behavior.
- Harden the training pipeline by implementing data provenance tracking, robust training techniques, and access controls.
Recommended Defenses
Data Integrity
Implement data provenance tracking for all training datasets, recording source, collection
method, preprocessing steps, and chain of custody.
Use cryptographic hashing to verify dataset integrity and detect tampering.
Maintain curated, trusted validation datasets that are isolated from potentially compromised
training data.
Apply data sanitization techniques such as spectral signatures, activation clustering, or STRIP
(STRong Intentional Perturbation) to identify and remove poisoned samples.
Robust Training
Use adversarial training to improve model resilience against adversarial inputs and reduce
sensitivity to small perturbations in training data.
Apply differential privacy during training to limit the influence of any single training example,
reducing the effectiveness of data poisoning.
Use data augmentation and mixup training to dilute the impact of poisoned samples.
Train on multiple independently sourced datasets and compare model behavior to detect
dataset-specific anomalies.
Access Controls
Restrict write access to training data repositories and model artifacts with role-based access
controls and audit logging.
Implement CI/CD pipeline security for model training workflows to prevent unauthorized
modifications to training code or configurations.
Use secure enclaves or isolated environments for training sensitive models. - Perform input validation on training data by scanning for poisoned samples, statistical outliers, label inconsistencies, mislabeled examples, duplicate injections, and anomalous patterns that could indicate backdoor triggers or targeted data manipulation. Use techniques such as influence function analysis, dataset distance metrics, and clustering-based anomaly detection.
- Perform post-training model validation by testing for backdoor behaviors using trigger pattern scanning, clean-label accuracy verification, neuron activation analysis (e.g., Neural Cleanse, Activation Clustering), and behavioral testing across diverse inputs to detect compromised or manipulated model behavior before deployment.
Attack Techniques
Most techniques run at several strengths. The first run is a clean baseline with no attack applied. A model that keeps its accuracy as the attack gets stronger is more robust. Untargeted attacks try to make the model give any wrong answer. Targeted attacks try to make it give one specific answer the attacker chose.
Whole-Image Perturbation
Changes every pixel of an image by a small, carefully computed amount that a person cannot see. The attacker needs digital access to the image file before it reaches the model. For video, the same approach is applied to every frame of the clip.
How Arena tests it: Arena generates adversarial examples using a white-box, gradient-based attack with direct model access. This is the strongest possible attacker assumption, with full knowledge of the model's architecture and gradients, so the results are a worst-case robustness bound. If a model withstands this attacker at a given budget, it withstands any less-informed real-world attacker at the same budget. Arena tests six perturbation budgets: ε = 0/255 (clean baseline), 1/255, 2/255, 4/255, 8/255, and 16/255. Each budget is the maximum amount any pixel's color channels may shift under an L∞ constraint. Every pixel can change, but none by more than the budget.
Localized Patch Attack
Places a small, visible patch, like a sticker or printed insert, on one part of an image. Because the patch is a physical object, an attacker can place it in the real world before a camera or scanner ever captures the scene. That makes it a practical attack even against systems that only see camera-captured images.
How Arena tests it: Arena uses the same white-box, gradient-based attack as whole-image perturbation, but confines the change to a small, contiguous region. Arena tests six patch sizes, expressed as a percentage of the total image area: 0% (clean baseline), 1%, 2%, 5%, 10%, and 20%.
Keyword Injection
Inserts a few extra words into a passage or sentence to change how a model ranks or selects it.
How Arena tests it: Arena uses a differential-evolution search to choose which words to insert and where. This is a black-box search: it only queries the model's output and does not use gradients. Arena tests insertion budgets of 0 (clean baseline), 1, 2, 3, and 5 words.
Invisible Character Injection
Inserts zero-width Unicode characters into text. These characters do not print, so the text looks exactly the same to a person, even under close inspection.
How Arena tests it: Arena uses the same black-box differential-evolution search as keyword injection to choose which characters to insert and where. Arena tests insertion budgets of 0 (clean baseline), 1, 2, 3, and 5 invisible characters.
Synonym Substitution
Replaces words with synonyms that keep the meaning of the text the same to a human reader.
How Arena tests it: Words are ranked by how much they affect the model's output, and the most influential are replaced first. Each substitution must pass a word-level cosine similarity threshold and a sentence-level check that the meaning has not drifted, so any failure reflects a real robustness gap rather than a change in meaning. Budgets are expressed as the cosine similarity threshold, from 1.0 (unperturbed) down to 0.5 (loosest).
Adversarial Audio Perturbation
Adds a carefully crafted, noise-like signal to an audio clip. The speech stays clearly intelligible to a person.
How Arena tests it: Arena computes the perturbation with white-box gradient access to the model. The budget is the signal-to-noise ratio (SNR) in decibels, from 100 dB (clean baseline) down to 20 dB.
Feature Perturbation
Shifts the numeric values in a row of structured data to change the model's prediction.
How Arena tests it: Arena uses a white-box, gradient-based attack on continuous numeric features. Categorical features stay fixed, because gradients are not defined for discrete values. The budget is a fraction of each feature's normalized [0,1] range, tested at ten levels from 0 (clean) to 0.9.
Data Poisoning
Corrupts part of a model's training data before the model is trained. Untargeted poisoning teaches the model a generally wrong behavior. Targeted poisoning plants a hidden trigger that makes the model produce an output the attacker chooses whenever the trigger appears, while the model behaves normally otherwise.
How Arena tests it: Arena corrupts a percentage of the training set, retrains the model, and evaluates the result. It repeats this at ten poisoning rates from 0% (clean baseline) up to 90%. Retraining at each rate takes seconds for tabular models but can take hours for image, audio, and text models, so Arena runs this test on tabular models only.





