Frontier models found the vulnerabilities. Only the attacker found the chains.
October 7, 2026
0 mins readAttackers don't read your repository; they hit your URL, and chain together whatever they find. In an era where offensive AI runs against live applications at machine speed, your security tooling needs to go beyond finding vulnerabilities to prove exploitability, including whether they can be combined into a breach.
So we ran a test. We pointed Evo Continuous Offensive Security (COS) and Claude Security running Mythos at the same application: TaintedPort. It’s a deliberately vulnerable web app that exercises real-world bug classes and a set of registered exploit chains. These are different kinds of tools; COS attacks the running application while the two Claude products read source code, and we ran them on the same target to see what each approach proves.
Evo COS confirmed 10 of 15 exploit chains.
This isn't a knock on models - we use them
Claude Security on Mythos is a real advance in AI-driven vulnerability discovery. It reasons about source code the way a strong human reviewer does, and it catches classes of flaws that pattern-matching tools miss. But finding a flaw in the code and proving an attacker can exploit these vulnerabilities are different jobs.
The attack path COS proved
Every tool in this test found the server-side request forgery (SSRF) flaw, and that the application's JWT signing secret was hardcoded. Evo COS went further and used the SSRF to reach and retrieve the signing secret from the running application, then used that secret to mint a valid administrator token and take over the admin surface. Two findings, connected into a full account takeover path, were demonstrated against the live app with a runnable proof of concept.

Above: CHAIN-004 as scored for COS. The two member findings, the SSRF flaw and the hardcoded secret, were reported individually by every tool in this test. Evo COS confirmed that they combine into admin token forgery.
That is what an autonomous attack looks like in practice: get a foothold, then chain from it. Evo COS shows you the path and attaches a runnable proof of concept for each confirmed chain.
The results
TaintedPort v1.35, scored against a fixed answer key of 57 known vulnerabilities and 15 known exploit chains, on a severity-weighted scale (Low 1, Medium 3, High 9, Critical 27; each confirmed chain scored at its own severity on top of its member findings). Weighted detection is points found ÷ 938 points available: 587 from the 57 vulnerabilities and 351 from the 15 chains.
Evo COS | Claude Security (Mythos) | |
|---|---|---|
Severity-weighted detection | 75.7% | 49.6% |
Exploit chains confirmed (of 15) | 10 | X |
Vulnerabilities found (of 57) | 50 | 37 |
F1 | 91.7% | 75.5% |
Claude Security found one more critical-severity vulnerability (10 to 9). Evo COS found the most vulnerabilities, had the highest precision (96.2% vs 90.2%) with fewer false positives, and it was able to identify and prove exploit chains.
Why attack the deployed application?
Evo COS is a dynamic system: its target is always a live URL, with source code as an optional input that upgrades a run from black-box to gray-box. It never runs on code alone. Claude Security ran white-box (code only).
That makes this a cross-disciplinary comparison: Evo COS is an offensive pentest that probes and exploits the running application, a fundamentally more intensive process than a single code-reading pass. It's the same complementary split the industry has always had between DAST and SAST. Here we're showing what exercising the live application proves and what reading the code alone cannot: that these vulnerabilities are real, and that they can chain.
The findings reported only by Evo COS are overwhelmingly runtime behaviors: reflected and misconfigured responses, missing transport protections, tokens that survive logout, and weak session handling. Evo COS was able to confirm several of the context-dependent business-logic flaws: broken object-level authorization on profile updates, privilege escalation via forged JWT claims, and a critical flaw where a single read of a stored TOTP secret nullifies 2FA. The chains require reaching each step through the running app and confirming the next one is actually reachable. A model reasoning over the source has to infer that a vulnerability executes, and misses the chain.
You can build a harness, but you can't build the context
This year's benchmarking debate has landed on a real point: the system around a model matters more than the model itself. We'd go further. "Which harness" isn't a neutral question, and the answer is where the advantage actually lives.
Evo COS orchestrates multiple models, including frontier Claude models, each directed at the job it's best at. So this result isn't a story about model quality. The same class of model that powers a frontier code review sits inside Evo COS. What's different is the system around it: Evo COS starts from what the Snyk platform already knows about the target, rather than reasoning cold; it's a group of agents, each orchestrated for a specific purpose, and every finding clears an independent validation step before it surfaces, because the system that generates a finding shouldn't be the one that grades it.
Anyone can wire a frontier model to a coding agent tonight and point it at a repository. What that setup can't reproduce is the accumulated, validated context Evo COS starts from, the independent validator, and the exercise against the live application.
Dynamic and static testing are complementary
Claude Security caught several vulnerabilities that Evo COS didn't, mostly source-visible logic and cryptographic flaws that don't always surface through the running app. Neither approach is complete on its own; this is the argument for a platform rather than a point tool. Reading the code and attacking the running app, each catches things the other misses.
Methodology
TaintedPort is Snyk's own deliberately vulnerable application, built and maintained by our team, and it's public. Evo COS was not tuned against it.
Target: TaintedPort: 57 known vulnerabilities (34 commodity, 23 business-logic) and 15 known exploit chains.
Tools and inputs:
Evo COS: gray-box (live URL + source), Sep 18.
Claude Security: white-box (source), Mythos with extended thinking, Sep 18.
Runs: one per tool. Results are the single-run outputs, not medians.
Detection by category
Category | Known | Evo COS | Claude Security (Mythos) |
Commodity | 34 | 33 | 21 |
Business logic | 23 | 17 | 16 |
Exploit chains | 15 | 10 | X |
Detection by severity (found/known)
Severity | Known | Evo COS | Claude Security (Mythos) |
Critical | 11 | 9 | 10 |
High | 26 | 22 | 19 |
Medium | 18 | 17 | 8 |
Low | 2 | 2 | 0 |
False positives and F1
Metric | Evo COS | Claude Security (Mythos) |
False positives | 2 | 4 |
F1 score | 91.7% | 75.5% |
Reproduce it: TaintedPort is available at taintedport.com.
Curious which findings in your own applications chain into a breach? See how Evo Continuous Offensive Security tests your live URL.
BOOK A LIVE DEMO
Secure AI adoption at scale
Evo helps organizations safely adopt and scale AI by providing visibility, governance, and security across AI-driven development and AI applications.
