A reproducible benchmark of Cybseco's deterministic engine against a versioned, hand-labeled corpus: precision, recall and F1 per CWE, per language and per analysis layer, with every layer we do not measure named and explained.
0.0%
Precision
0.0%
Recall
0
Corpus cases
121 CWE · 22 lang
6
False positives
931 true positives
Every case, every flow
Each point is a corpus case; each line traces untrusted input to the sink, colored by vulnerability class. Filled = detected, dashed = honest frontier.
Every layer Cybseco ships is listed here, measured or not. Each figure covers only the layer on its row. Read the full methodology
Layers this benchmark does not measure
Cybseco ships these. No number above describes them, and here is why.
Dependency vulnerabilities (SCA) — Produced by external scanners that read a live vulnerability database. Their output changes when the database changes, with no change to Cybseco, so a precision figure measured today would describe the database rather than the engine and would not reproduce tomorrow.
Standalone secret scanning — Gitleaks is an external binary and is not installed for this benchmark. Hardcoded credentials found by Cybseco's own analyzers are measured, inside the code and CI/CD layers, under CWE-798 and CWE-532.
Third-party SAST (Semgrep, Bandit) — An external ruleset Cybseco orchestrates but does not author. Measuring it would report Semgrep's accuracy, not Cybseco's.
Licence compliance policy — A policy decision about a project's licences, not a judgement about its code. The corpus labels vulnerabilities, so a licence finding has nothing to be right or wrong against here.
LLM-assisted analysis — Its output depends on a model and a prompt rather than on the engine, and it is not part of the deterministic result this benchmark measures.
Graph, attack paths, decision engine — A different kind of claim: these produce paths and plans, not findings, so precision and recall over labelled lines cannot express them. The methodology page explains what could be measured, and which of these cannot honestly be reduced to a percentage at all. Publishing a number before the method is settled is how a benchmark stops being evidence.
What is measured
Layer
Corpus cases
Safe cases
TP
FP
FN
Findings measured
Precision
Recall
Application code (18 languages)
1306
497
810
0
0
810
100.0%
100.0%
Terraform (HCL semantics)
43
10
33
1
0
34
97.1%
100.0%
Cloud IaC rules (AWS, Azure, GCP)
26
7
19
0
0
19
100.0%
100.0%
Kubernetes manifests
20
6
14
0
0
14
100.0%
100.0%
Dockerfile
16
6
10
0
0
10
100.0%
100.0%
Docker Compose
15
6
9
0
0
9
100.0%
100.0%
CI/CD pipelines (7 platforms)
30
8
22
1
0
23
95.7%
100.0%
LLM and AI agent integration code
24
10
14
4
0
18
77.8%
100.0%
Findings measured is true positives plus false positives: the number of findings each precision on that row was computed from. A layer with few findings carries a correspondingly wide margin of error in either direction. 100.0% over nine findings and 100.0% over eight hundred are the same number and not the same evidence.
The LLM and AI agent integration code layer analyses an application's own integration code for LLM and agent mistakes. It is not a measure of any AI inside Cybseco: this benchmark runs the deterministic engine with no model in the loop.
Known false positives
Findings the engine produces on code that is not vulnerable. They are counted against the precision above, not excluded from it.
CWE-798cybseco-cicd — `DEPLOY_TOKEN = credentials('deploy-token')` is the documented Jenkins idiom for *not* committing a credential: it binds a value stored in the credential store. The rule sees NAME = <call> and reports a hardcoded secret.
CWE-284cybseco-terraform — The rule reads source_address_prefix and destination_address_prefix into one list and fires if any entry is `*`. Here the wildcard is the destination, which on an inbound NSG rule means 'anywhere inside the protected scope'; the source is a bastion subnet, 10.1.0.0/24.
CWE-78cybseco-llm — The completion is used as a key into a table of commands the developer wrote, and a key outside the table is refused before the call. Nothing the model emits becomes a command. This is the recommended mitigation for the vulnerability the rule looks for.
CWE-89cybseco-llm — The SQL text is a constant; the model only selects which named query to run, and the value travels as a bound parameter. The rule fires on a completion and a query execution appearing in the same function.
CWE-918cybseco-llm — The host is compared against a frozen allow-list and the request is refused otherwise, so metadata endpoints and internal services are unreachable. The rule does not see the check between the completion and the request.
CWE-200cybseco-llm — `page_token` is a pagination cursor. The sensitive-name pattern matches the word `token` regardless of what it holds.
Coverage by CWE
Detection rate per vulnerability class, with the false positives counted rather than assumed away.
CWE-8942 TP · 1 FP
SQL injection
CWE-7833 TP · 1 FP
Command injection
CWE-2226 TP · 0 FP
Path traversal
CWE-91826 TP · 1 FP
SSRF
CWE-32721 TP · 0 FP
Weak crypto
CWE-73220 TP · 0 FP
Weak permissions
CWE-31919 TP · 0 FP
CWE-7919 TP · 0 FP
XSS
CWE-79819 TP · 1 FP
CWE-25017 TP · 0 FP
CWE-29517 TP · 0 FP
TLS verification
CWE-32617 TP · 0 FP
CWE-50217 TP · 0 FP
Unsafe deserialization
CWE-60117 TP · 0 FP
Open redirect
CWE-28416 TP · 1 FP
Access misconfiguration
CWE-33816 TP · 0 FP
Weak randomness
CWE-34716 TP · 0 FP
CWE-61116 TP · 0 FP
XXE
CWE-20915 TP · 0 FP
CWE-32115 TP · 0 FP
CWE-61315 TP · 0 FP
CWE-64315 TP · 0 FP
XPath injection
CWE-91615 TP · 0 FP
CWE-9415 TP · 0 FP
Code injection
CWE-11714 TP · 0 FP
CWE-31214 TP · 0 FP
CWE-43414 TP · 0 FP
CWE-53214 TP · 0 FP
CWE-61414 TP · 0 FP
CWE-102113 TP · 0 FP
CWE-11313 TP · 0 FP
CWE-133313 TP · 0 FP
ReDoS
CWE-50113 TP · 0 FP
CWE-38412 TP · 0 FP
CWE-52112 TP · 0 FP
CWE-77612 TP · 0 FP
CWE-8812 TP · 0 FP
CWE-94312 TP · 0 FP
CWE-133611 TP · 0 FP
Template injection
CWE-31111 TP · 0 FP
CWE-28510 TP · 0 FP
CWE-40010 TP · 0 FP
CWE-47010 TP · 0 FP
Unsafe reflection
CWE-9010 TP · 0 FP
CWE-94210 TP · 0 FP
CWE-1169 TP · 0 FP
CWE-6689 TP · 0 FP
CWE-9159 TP · 0 FP
CWE-959 TP · 0 FP
Code injection
CWE-3528 TP · 0 FP
CWE-2697 TP · 0 FP
CWE-3627 TP · 0 FP
CWE-6397 TP · 0 FP
CWE-6937 TP · 0 FP
CWE-7707 TP · 0 FP
CWE-11046 TP · 0 FP
CWE-3066 TP · 0 FP
CWE-5226 TP · 0 FP
CWE-8626 TP · 0 FP
CWE-13575 TP · 0 FP
CWE-3295 TP · 0 FP
CWE-3775 TP · 0 FP
Insecure temp file
CWE-1204 TP · 0 FP
CWE-3594 TP · 0 FP
CWE-4944 TP · 0 FP
CWE-7784 TP · 0 FP
CWE-1903 TP · 0 FP
CWE-2153 TP · 0 FP
CWE-4893 TP · 0 FP
Debug enabled
CWE-733 TP · 0 FP
CWE-8293 TP · 0 FP
CWE-10502 TP · 0 FP
CWE-1192 TP · 0 FP
CWE-12842 TP · 0 FP
CWE-13212 TP · 0 FP
CWE-202 TP · 0 FP
CWE-2482 TP · 0 FP
CWE-3202 TP · 0 FP
CWE-3302 TP · 0 FP
Weak randomness
CWE-5982 TP · 0 FP
CWE-7032 TP · 0 FP
CWE-7042 TP · 0 FP
CWE-7492 TP · 0 FP
CWE-9392 TP · 0 FP
CWE-982 TP · 0 FP
File inclusion
CWE-10791 TP · 0 FP
CWE-11881 TP · 0 FP
CWE-12201 TP · 0 FP
CWE-1251 TP · 0 FP
CWE-13271 TP · 0 FP
CWE-1341 TP · 0 FP
CWE-151 TP · 0 FP
CWE-161 TP · 0 FP
CWE-1701 TP · 0 FP
CWE-2001 TP · 1 FP
CWE-2141 TP · 0 FP
CWE-2421 TP · 0 FP
CWE-2561 TP · 0 FP
CWE-2661 TP · 0 FP
CWE-3081 TP · 0 FP
CWE-3221 TP · 0 FP
CWE-3541 TP · 0 FP
CWE-3671 TP · 0 FP
CWE-4041 TP · 0 FP
CWE-4191 TP · 0 FP
CWE-4571 TP · 0 FP
CWE-4971 TP · 0 FP
CWE-5521 TP · 0 FP
CWE-6211 TP · 0 FP
CWE-6621 TP · 0 FP
CWE-6641 TP · 0 FP
CWE-6651 TP · 0 FP
CWE-6671 TP · 0 FP
CWE-741 TP · 0 FP
CWE-7541 TP · 0 FP
CWE-771 TP · 0 FP
Command injection
CWE-8071 TP · 0 FP
CWE-8431 TP · 0 FP
CWE-9171 TP · 0 FP
CWE-9231 TP · 0 FP
CWE-9261 TP · 0 FP
Data-flow shapes the engine follows
The same taint confirmation applies across every one of these flows, and across files.
Scope: the deterministic native and semantic engine, no external scanners, no LLM, the proprietary engine that ships identically everywhere. Want to see it on your code? Book a Demo.
See Cybseco reason about a real-world system
Watch the interactive demo, then request a guided trial for your team.