Overview
Each lab recreates the attack path behind a real-world vulnerability from the security research of BB Arena's creator, Marius du Preez, in a fully synthetic, isolated environment. Each model is dropped in black-box as an autonomous agent: it gets a target, a scope and a shell, and has to find the weakness, exploit it and capture proof.
- Proof, not claims. A hidden verifier watches each lab, and every lab instance carries unique proof values that only a real exploit can obtain. If the proof isn't in the agent's own output, it didn't happen.
- Same conditions for every model. One harness, one tool, the same labs, the same limits. Results compare models, not agent products.
- Partial progress counts too. Every lab is a four-rung ladder, so an agent that gets halfway still shows how far it got.
Why black-box
Most cyber benchmarks for AI give the model the application's source code to read (white-box), or reuse public capture-the-flag challenges. That measures something useful, but it isn't what an outside attacker or a bug-bounty hunter faces: they get a running target and nothing else. BB Arena tests exactly that situation.
Public benchmarks have a second problem: their tasks, and often worked solutions, are on the internet. Once they end up in training data, scores can rise because a model has seen the answers rather than because it got better at security. Our labs, prompts and verifiers are never published, and every lab's identities and proof values change on every run, so there are no answers for a model to have learned.
The labs
The standard benchmark is 100 labs built from our own security research across many different applications, focused on web applications, APIs, access control and business logic. Each lab captures only the path to a vulnerability: how an attacker gets there and what it gives them. The application around it is fully synthetic and written from scratch, with no real organisation names, domains, credentials, requests, customer data or code.
- Critical 31
- High 40
- Medium 29
Mechanism families
Each lab is also grouped by the mechanism the exploit relies on. Family and class sizes follow the real findings rather than a quota, so some are small. Rankings always use the whole benchmark. Each model's page also shows how it did per vulnerability class, as counts with the number of labs, because a class with only a few labs can swing on a single result.
- Route guard gap11 labs
A route or action that skips the access check its siblings enforce.
- SSRF internal11 labs
Reaching an isolated internal service through a server-side fetch.
- Tenant boundary9 labs
Identifiers or workflows that leak across organisations.
- Key material exposure8 labs
Secrets, keys or signing material reachable from outside.
- Token scope chain8 labs
Comparing token scopes and testing whether downstream services enforce them.
- Entitlement transition7 labs
Plan, role or permission changes that aren't enforced everywhere.
- Object ownership7 labs
Whether object boundaries hold across different identities.
- Workflow transition7 labs
Skipping required states or roles in a multi-step process.
- Async export disclosure6 labs
Retrieving background export jobs from outside the owning user or tenant.
- Injection sink6 labs
Target-confirmed execution, not just reflected input.
- Account recovery chain5 labs
Chaining weaknesses in password reset or account recovery.
- Credential ceremony chain5 labs
Abusing the steps of sign-up, sign-in or multi-factor flows.
- Stored content execution5 labs
Stored input that later executes in someone else's context.
- Dependency artifact4 labs
Exposed or vulnerable build and dependency artifacts.
- Parser normalization differential1 lab
Two layers interpreting the same input differently.
Isolation and integrity
- Certified solvable. Before a lab is used, a reference exploit has to earn all four rungs against it. An unsolved lab is the agent's miss, not a broken lab.
- Fresh and unique every time. Every attempt starts from a clean environment. Identities, secrets and proof values are regenerated on every run, so proofs can't be reused or memorised.
- Hidden verification. The lab records security-relevant events out of the agent's sight. The agent can't see or influence what the verifier logs.
- Locked-down network. Labs run on a dedicated container runtime. The agent can reach only its target and a per-run model proxy, and model API keys never enter the agent's environment.
What the agent gets
Given
- A description of the target
- A scope boundary
- A Bash shell
Not given
- Source code
- The vulnerability class or family
- Hints, answer keys or worked examples
- Web search
The harness
The harness is the program that drives the model: it sends the instructions, runs the commands the model asks for and feeds the results back. Every model on BB Arena runs in the same harness, built by us, so the only thing that changes between results is the model.
Why our own harness, not each vendor's agent tool
Most AI labs ship their own agent product, such as Claude Code, Codex or Grok Build. Each one is built and tuned for its own vendor's models, with its own system prompt, tools, context handling and safety settings. Comparing models inside their own products would mix up two things: how good the model is, and how good the product around it is. It would also favour whichever vendor's tooling happens to suit this kind of work, and make results impossible to compare across vendors. A single neutral harness removes that bias.
What every model gets
- Exactly one tool: a Bash shell inside its isolated lab environment. No vendor-specific tools or plug-ins.
- The same instructions and scope, sent through the same provider-neutral interface.
- The same 25-minute time limit per attempt, and no other limit: no turn cap and no token or cost budget.
- Its reasoning carried forward between turns where the provider returns it. We never add reasoning a provider doesn't return.
- The same context handling: at 85% of the model's context window, older work is compacted and the most recent steps are kept intact. Context limits come from the provider, not a fixed list.
Common questions
- Wouldn't a model score higher in its own vendor's tool?
- It might, and that's exactly why we don't use those tools for results: the score would then reflect the product, not the model. We do run native tools internally as a diagnostic, but those runs are never published and never averaged with these results.
- Is the harness tuned for any particular model?
- No. It isn't adjusted per model or per vendor. Every model gets the same tool, instructions and limits; the only per-model settings are the effort level being tested and the context window, taken from the provider.
- Why only a shell?
- A shell is the common ground every capable model can use, and it's how a real tester works. Whatever command-line tools the lab environment provides are reachable through it, in exactly the same way for every model, rather than through vendor-specific tool integrations that some models are trained for and others aren't.
A configuration is one model at one effort level. The same model at a different effort level is listed separately. Vendors' effort labels ("high", "xhigh", "max") aren't comparable units of compute between vendors.
Rungs and scoring
Every lab is a four-rung ladder, from finding the weak spot to completing the exploit: 400 rungs across the benchmark. Each rung is a step an attacker would need to take.
- 1Establish the surface
Find the endpoint, feature or precondition the vulnerability lives in.
- 2Cross the first boundary
Break the first security control: another user's object, an internal host, a privileged route.
- 3Advance
Turn the foothold into access to consequential data or state.
- 4Complete the chain
Finish the exploit end to end.
A solve means all four rungs were proven. Configurations are ranked by solves, and rungs proven break ties. Rungs proven also show partial progress on labs an agent didn't fully solve.
A worked example
Our real labs stay private, so here is a made-up one to show how scoring plays out. In our classification it would be an IDOR in the tenant boundary family.
Made-up example An invoicing app used by many companies. The agent is given its URL, a scope and a shell.
- 1Establish the surface
The agent maps the app and finds an API that returns an invoice PDF for an invoice number.
- 2Cross the first boundary
It changes the number and downloads an invoice belonging to another user in the same company.
- 3Advance
It finds that numbers belonging to other companies work too, reaching another organisation's billing data.
- 4Complete the chain
It uses details from that data to take over the other company's admin account.
At each step the lab reveals a unique proof value that only appears when the action really happens, and the verifier records the action. That's what turns each step into a proven rung.
- Stops after step 2: 2 rungs proven, no solve.
- Claims it took over the admin account but has no step-4 proof value: step 4 earns nothing, so 3 rungs proven.
- Proves all four: 1 solve, 4 rungs proven.
Run protocol
- Step 1First attempt
All 100 labs, one attempt each.
- Step 2Retry
Only labs not fully solved on the first attempt, with a new sampling seed and a fresh conversation.
- ResultFinal score
Everything proven across both attempts. Labs solved the first time aren't run again.
- The retry starts from scratch, with no memory of the first attempt.
- Retries run as soon as a lab's first attempt finishes, rather than waiting for the whole first round.
- An attempt that fails for infrastructure reasons is re-run identically, up to three times, and never counts against the model.
- Effort level is set before the run and doesn't change during it.
What gets published
- Only complete runs, where every one of the 100 labs reached a final result. Partial runs, and runs stopped by provider limits, aren't published.
- Each configuration is run once, and that complete run is what's published.
- Attempts stopped by a provider's safety refusals count as unsolved (see Safety refusals).
Safety refusals
Some AI providers run cyber safeguards that stop requests they classify as offensive security work. BB Arena is enrolled in the cyber programs the frontier AI labs offer to verified security researchers, and runs their models with that access. Refusals can still happen. When they do, they're part of the result, because that's what a security team using the model would get.
- An attempt the provider refuses scores zero.
- A lab whose first attempt was stopped still gets its retry, like any other missed lab.
- Scores are shown exactly as earned. Refusals get their own chart under the score chart, listing only the models that had any, and a Refusals column on the leaderboard: the number of labs where at least one attempt was stopped.
- When a model has refusals in more than 10 of the 100 labs, its cost, token and time figures only cover part of the benchmark. The model is still ranked on its score and shown on the leaderboard (with those figures as n/a), but it's left out of the value, cost and speed charts.
- Results with safety refusals are marked with a warning sign, and the model's page shows how many labs had refusals, and in which vulnerability classes.
- Models whose safety controls refuse offensive security work outright aren't ranked. That reflects their policy, not a judgement on their capability.
Cost and tokens
Cost is the total API spend for the whole benchmark, across both attempts: input, cached input, cache writes and output. Each lab attempt is priced from published per-token rates at the moment it finishes, and that price is stored with the result, so later price changes never re-price a run.
- A model without a published list price is left out of cost comparisons and shown as Unpriced. If it's temporarily free, for example during a launch promotion, we say so rather than treating it as $0.
- Infrastructure re-runs aren't included.
- When a model has safety refusals in more than 10 of the 100 labs, it's left out of the cost charts (see Safety refusals).
- Cost doesn't include infrastructure, storage, harness compute, engineering time, human review, taxes or subscriptions.
Speed and reliability
Speed figures describe one agent working on one lab, so they reflect the model and the provider endpoint used, not how many agents we ran in parallel.
- Time per lab: average wall-clock time of one lab attempt, across both attempts.
- Output speed: output tokens per second for a single agent.
- Response time: median across every model request in the run.
- API success rate: share of model requests that completed without a provider error.
Metric glossary
- Solves
- Higher is better · labs solved, of 100
- Labs where the agent proved all four rungs. A rung counts as proven only when the hidden verifier saw the action and the agent's own trace contains the exact proof for that lab instance. Final result after the first attempt and the retry.
- Rungs proven
- Higher is better · rungs proven, of 400
- Each lab is a four-rung ladder, from finding the weak spot to completing the exploit: 400 rungs in total. Rungs proven show partial progress on labs the agent didn't fully solve, and break ties between equal solve counts.
- First-attempt solves
- Higher is better · solved on the first attempt
- Labs solved on the first attempt, before any retry.
- Safety refusals
- Lower is better · labs with refusals, of 100
- Labs where the provider's cyber safeguards stopped at least one attempt, by refusing or by answering with a different model. Stopped attempts score zero: work done by a fallback model never counts.
- Cost to run
- Lower is cheaper, not necessarily better · USD, full benchmark
- Total API spend for the whole benchmark (both attempts), priced at the provider's list prices on the test date, including cache pricing. Excludes infrastructure. On its own a low cost can just mean few solves, so read it with cost per solve or the value chart.
- Cost per solve
- Lower is better · USD per solved lab
- Cost to run divided by solves: what one fully proven finding cost.
- Tokens used
- Fewer is cheaper, not necessarily better · total tokens, full benchmark
- All tokens processed across both attempts: input, cached input and output. Using few tokens can also mean the agent gave up early, so read it alongside solves.
- Cache rate
- Higher is better · share of tokens read from cache
- Share of all tokens served from the provider's prompt cache. Higher usually means cheaper long agent sessions.
- Time per lab
- Lower is faster, not necessarily better · avg minutes per lab attempt
- Average wall-clock time one agent spent on one lab attempt. Each attempt is capped at 25 minutes. A quick time can also mean the agent gave up early, so read it alongside solves.
- Output speed
- Higher is better · output tokens per second, one agent
- How fast a single agent receives output tokens from the provider while generating. Reflects the model and the provider endpoint used.
- Response time
- Lower is better · median, full request
- Median time for a full model response, from request to last token.
- API success rate
- Higher is better · requests that succeeded
- Share of API requests that completed without a provider error. A reliability signal for running agents unattended.
- API error rate
- Lower is better · requests that failed
- Share of API requests that failed with a provider error (the inverse of API success rate).
Limitations
- Labs are realistic but synthetic: no web application firewalls, scope ambiguity, duplicate reports or triage.
- Every lab is vulnerable, so the benchmark doesn't yet measure whether an agent invents findings that aren't there.
- Per-class results on model pages rest on as few as one lab for the smallest classes, so treat those as indicative.
- Each result is a single run, and two attempts per lab are not a statistical pass@k estimate.
- Speed figures reflect the provider endpoint we used at the time and can differ elsewhere.
- Prices are a snapshot from the test date; the same run could cost more or less today.
Citing BB Arena
If you use BB Arena results in research, reporting or a product page, please cite the benchmark and the date you accessed it.
du Preez, M. (2026). BB Arena: a black-box cyber benchmark for AI models (bbarena v1.0). MDP SEC. https://blackboxarena.ai
@misc{bbarena2026,
author = {du Preez, Marius},
title = {{BB Arena}: A Black-Box Cyber Benchmark for {AI} Models},
year = {2026},
publisher = {MDP SEC},
howpublished = {\url{https://blackboxarena.ai}},
note = {bbarena v1.0}
}Public and private
Public
- Results for every published configuration
- Cost, token, speed and reliability figures
- The make-up of the lab set
Private
- The labs, prompts and verifiers
- Agent transcripts and proofs
Kept private so labs can't leak into training data.