What we measure, what we do not, and why
The benchmark exists to be checked, not admired. This page states its exact perimeter: what the published figures cover, what they deliberately leave out, and which of our capabilities cannot honestly be reduced to a percentage.
What the benchmark measures
One question, asked over a versioned corpus of hand-judged projects: given a file we have already decided about, does the engine report the vulnerability that is there, on the right line, without reporting things that are not?
What the benchmark does not measure
Cybseco ships these. No figure on the benchmark page describes them, and each one carries its reason. A capability left out silently is a capability a reader assumes the numbers covered.
- Dependency vulnerabilities (SCA) — Produced by external scanners that read a live vulnerability database. Their output changes when the database changes, with no change to Cybseco, so a precision figure measured today would describe the database rather than the engine and would not reproduce tomorrow.
- Standalone secret scanning — Gitleaks is an external binary and is not installed for this benchmark. Hardcoded credentials found by Cybseco's own analyzers are measured, inside the code and CI/CD layers, under CWE-798 and CWE-532.
- Third-party SAST (Semgrep, Bandit) — An external ruleset Cybseco orchestrates but does not author. Measuring it would report Semgrep's accuracy, not Cybseco's.
- Licence compliance policy — A policy decision about a project's licences, not a judgement about its code. The corpus labels vulnerabilities, so a licence finding has nothing to be right or wrong against here.
- LLM-assisted analysis — Its output depends on a model and a prompt rather than on the engine, and it is not part of the deterministic result this benchmark measures.
- Graph, attack paths, decision engine — A different kind of claim: these produce paths and plans, not findings, so precision and recall over labelled lines cannot express them. The methodology page explains what could be measured, and which of these cannot honestly be reduced to a percentage at all. Publishing a number before the method is settled is how a benchmark stops being evidence.
Why some capabilities are not a percentage
Detection has a right answer: the vulnerability is on that line or it is not. Some of what Cybseco does makes a different kind of claim, and precision and recall cannot express it. Publishing a number anyway would be the easiest thing on this page to do, and the fastest way to lose the argument with someone who reads carefully.
How to read the figures
Four things worth knowing before you draw a conclusion from any single number on the benchmark page.
The rules we hold ourselves to
These are constraints on us, not claims about the product. They are what make the figures worth reading.
- Anything Cybseco announces is either measured or named on this page as out of scope. There is no third state, and an automated gate fails the build when an analyzer ships without one or the other.
- Ground truth describes the code, decided before the engine runs. A label written to match engine output is not evidence.
- A figure is published with the sample it was computed from.
- A defect we know about is published, not excluded from the metric. A gate also fails when a published defect quietly stops reproducing, so the list can only shrink honestly.
- No comparison against another product on capabilities where we would be choosing their configuration, their corpus and the scenario. Nobody should believe such a comparison, including us.
- No figure before its methodology is public and its ground truth inspectable.
See Cybseco reason about a real-world system
Watch the interactive demo, then request a guided trial for your team.