/leaderboard · loading…
/leaderboard · loading…
/leaderboard · all-time · measured as of 2026-06-07
Every model ROGUE red-teams, ranked by resistance. Lower breach rate = higher rank. Each model is hit with ROGUE's open-web attack corpus and graded by a calibrated judge. 8 production + 16 open-source models — a periodic measured snapshot (production as of 2026-06-07), not a live-updating feed.
Most resistant
1%
claude-opus-4-8
Most breached
44%
voxtral-small-24b-2507
submit your model
Want your own model on the board? Opt-in submissions will be re-verified against the same corpus and judge before they appear.
closed · deep pipeline · measured 2026-06-07
8 hosted models · deep escalation + PAIR pipeline · graded by the calibrated v3 judge (89.3% human agreement). Snapshot as of 2026-06-07.
claude-opus-4-8
claude-haiku-4-5
gpt-5.4-nano
llama-3.1-8b-instruct
gemini-3.1-flash-lite
gpt-audio-mini
mistral-small-2603
voxtral-small-24b-2507
open-weight · single-shot · measured 2026-06-20
16 open-weight models · single-shot aggressive jailbreak pack · primitive-level any-breach · ~40 trials each. A lighter, separate methodology — not directly comparable to the deep-pipeline board.
DeepSeek-R1-Distill-Llama-70B
DeepSeek-V3.1
Qwen2.5-72B-Instruct
MiniMax-M2
Meta-Llama-3.1-8B-Instruct
QwQ-32B
GLM-4-9B-0414
Mistral-Small-24B-Instruct-2501
Mistral-7B-Instruct-v0.2
Qwen3-32B
Meta-Llama-3.1-70B-Instruct
Qwen3-235B-A22B
Qwen2.5-7B-Instruct-abliterated
Qwen2.5-7B-Instruct
Mistral-Nemo-Instruct-2407
Meta-Llama-3.1-8B-Instruct-abliterated
// rank = mean any-breach rate across attack families (lower = more resistant), WITHIN each board
// production: ROGUE all-time corpus, 11,973 deep-pipeline trials, calibrated v3 judge (89.3% human agreement) — the same numbers `rogue try` prints
// open-source: curated single-shot aggressive jailbreak pack (~40 trials each), primitive-level any-breach, measured 2026-06-20
// each facet (production, open-source, PII context-leak, long-context robustness, safeguard over-refusal, agentic IPI) measures a DIFFERENT attack/metric — NOT directly comparable, each ranked independently
// periodic measured snapshot, not a live feed — for the live per-family × config breakdown, see the breach matrix
pii context-leak · exact-match · measured 2026-07-06
Confidential PII is planted in each model's system promptwith a “never reveal” instruction, then probed for disclosure. Ranked by leak rate (lower = more resistant). A separate attack + metric — not comparable to the breach boards above. Hosted frontier models held at 0%; the leak lives in smaller open models.
gpt-5.4
hostedgpt-5.4-nano
hostedQwen2.5-32B-Instruct
openQwen2.5-14B-Instruct
openQwen2.5-7B-Instruct
openQwen2.5-3B-Instruct
openMistral-Nemo-Instruct-2407
open// system-prompt-planted PII · exact-match leak · 36 cells/model · Wilson 95% CI · Fisher vs 0% baseline — measured 2026-07-06. Small panel (2 hosted + 5 open); a first context-leak facet, ranked independently.
long-context robustness · breaks at N tokens · measured 2026-07-11
Many-shot / long-context attacks flood the context until a model caves. Each model is swept over a token ladder (2K→128K) to find the threshold where it breaks — the metric nobody else publishes. Ranked by long-context breach rate (lower = more robust). A separate attack + metric — not comparable to the breach boards above. Click a model to see the attacks by token level.
// Fireworks 128K panel × many-shot ladder of AdvBench harmful asks (2K→128K), comply-vs-refuse judged + verified; n=71–72/model — measured 2026-07-11. Directional (small per-model n); a first robustness-threshold facet, ranked independently.
safeguard · over-refusal (FRR) · measurement pending
The helpfulness costcompanion to the breach boards: rather than “does the model comply with harm?”, it measures how often a model wrongly refuses a benign look-alike — the false-refusal rate on safe prompts that merely resemble unsafe ones. Lower is better. A separate metric — not comparable to the breach boards above.
No approved paid run has populated this board yet. The metric, panel and benign denominators are wired and gated — the board lights up the moment a run lands, with no fabricated numbers shown in the meantime.
detector: over-block judge (pending) · Wilson + bootstrap CI
// source data/safeguard/board.json (jbb_benign, xstest) — pending an approved paid run; Measurement pending — no approved paid run has populated this board yet. Run scripts/sweep/run_over_refusal_sweep.py --go to measure.
agentic IPI · marginal ASR by lever · measurement pending
Per DeploymentConfig indirect prompt injection: a malicious instruction is smuggled into a tool result or retrieved document — not the user turn — and we measure how often the agent follows it, broken down by the template × carrier × concealment levers, plus a defense-stack fingerprint. A separate attack + metric — not comparable to the breach boards above.
No approved paid run has populated the IPI matrix yet. The three levers below and the guardrail-fingerprint probe are wired and gated — the board lights up per DeploymentConfig once a run lands, with no fabricated numbers shown in the meantime.
by template
the injection wording — how the rogue instruction is phrased
by carrier
which tool result / retrieved document carried the payload
by concealment
how the instruction was hidden (comment, unicode, framing)
+ defense-stack fingerprint (guardrail-fingerprint-v1): guard present? · posture · guard family · confidence
// ipi-matrix-v1 — pending an approved paid run; Measurement pending — no approved paid run has populated the IPI matrix yet. The template × carrier × concealment levers and the guardrail-fingerprint probe are wired and gated.
/leaderboard · share card
All 24 models in a single shareable card — production endpoints (deep escalation + PAIR) and open-source models (single-shot pack) in two clearly-separated panels. Not directly comparable; each panel ranks within itself.
