← Leaderboard
Anthropic · Proprietary · tested 8 Oct 2026
Claude Sonnet 5 High effort
Overall rank#9 of 10
Solves46/10039 first attempt · +7 retry
Rungs proven243/400223 first attempt · +20 retry
Cost to run$258$5.60 per solve
Tokens used587.3M94.4% from cache · 8.6M output
Time per lab18.2 minaverage of 153 attempts
Where it places
Each line is one metric across all 10 configurations. Right is always better; this model is the large dot.
Solveslabs solved, of 100
worsebetter
46#8 of 10
Rungs provenrungs proven, of 400
worsebetter
243#8 of 10
First-attempt solvessolved on the first attempt
worsebetter
39#7 of 10
Cost per solveUSD per solved lab
worsebetter
$5.60#6 of 7
By vulnerability class
How Claude Sonnet 5 (High) did on each type of vulnerability. Shown as counts, because some classes have only a few labs.
- Authentication / authorisation bypass18 of 28 solved83 of 112 rungs
- Information disclosure8 of 12 solved33 of 48 rungs
- IDOR3 of 11 solved21 of 44 rungs
- SSRF0 of 10 solved8 of 40 rungs
- Business logic6 of 10 solved31 of 40 rungs
- Key or secret exposure3 of 8 solved22 of 32 rungs
- XSS4 of 6 solved17 of 24 rungs
- Token theft4 of 5 solved19 of 20 rungs
Fewer than 5 labs each: one lab changes these a lot, so read them with care.
- Supply chain0 of 4 solved5 of 16 rungs
- Remote code execution0 of 3 solved1 of 12 rungs
- SQL injection0 of 2 solved3 of 8 rungs
- File inclusion0 of 1 solved0 of 4 rungs
Provider performance
Measured per single agent against the provider endpoint used for the run.
- Response time
- 9.16 s median · 30.45 s p95
- Output speed
- 54.4 tok/s
- API success rate
- 99.39%
- Input context
- 41.5K median · 103.3K p95
About this result
Every model runs in the same neutral harness, so results compare the models themselves.
- Harness
- BB Arena neutral harness
- Safety refusals
- None