Benchmark methodology

What we measure, what we do not, and why

The benchmark exists to be checked, not admired. This page states its exact perimeter: what the published figures cover, what they deliberately leave out, and which of our capabilities cannot honestly be reduced to a percentage.

What the benchmark measures

One question, asked over a versioned corpus of hand-judged projects: given a file we have already decided about, does the engine report the vulnerability that is there, on the right line, without reporting things that are not?

A versioned corpus
Each case is a small project holding the smallest code that exhibits, or safely avoids, one vulnerability. Each carries a comment saying what is wrong, what is expected and why. The corpus is versioned: a change to it is a change to the published figures.
Ground truth written about the code
The expected result is a judgement about the code, decided before the engine runs, never a description of what the engine produced. When the two disagree we investigate the engine first. A label written to match engine output measures nothing, and we have had to correct labels that had drifted that way.
The pipeline that ships
Every case is analysed the way a customer scan runs: the full orchestrator at maximum depth, no external scanners, no model in the loop. The figures describe the engine that ships identically everywhere, not a laboratory configuration assembled for the occasion.
Correct code, deliberately
A third of the corpus is code that is not vulnerable, most of it the corrected form of a case that is. A corpus of vulnerabilities alone cannot detect a false positive, and false positives are what make a security tool stop being read.

What the benchmark does not measure

Cybseco ships these. No figure on the benchmark page describes them, and each one carries its reason. A capability left out silently is a capability a reader assumes the numbers covered.

  • Dependency vulnerabilities (SCA) — Produced by external scanners that read a live vulnerability database. Their output changes when the database changes, with no change to Cybseco, so a precision figure measured today would describe the database rather than the engine and would not reproduce tomorrow.
  • Standalone secret scanning — Gitleaks is an external binary and is not installed for this benchmark. Hardcoded credentials found by Cybseco's own analyzers are measured, inside the code and CI/CD layers, under CWE-798 and CWE-532.
  • Third-party SAST (Semgrep, Bandit) — An external ruleset Cybseco orchestrates but does not author. Measuring it would report Semgrep's accuracy, not Cybseco's.
  • Licence compliance policy — A policy decision about a project's licences, not a judgement about its code. The corpus labels vulnerabilities, so a licence finding has nothing to be right or wrong against here.
  • LLM-assisted analysis — Its output depends on a model and a prompt rather than on the engine, and it is not part of the deterministic result this benchmark measures.
  • Graph, attack paths, decision engine — A different kind of claim: these produce paths and plans, not findings, so precision and recall over labelled lines cannot express them. The methodology page explains what could be measured, and which of these cannot honestly be reduced to a percentage at all. Publishing a number before the method is settled is how a benchmark stops being evidence.

Why some capabilities are not a percentage

Detection has a right answer: the vulnerability is on that line or it is not. Some of what Cybseco does makes a different kind of claim, and precision and recall cannot express it. Publishing a number anyway would be the easiest thing on this page to do, and the fastest way to lose the argument with someone who reads carefully.

The architecture graph
A factual claim: these assets exist and these relations hold between them. Measurable, and the first thing we intend to measure. It needs a corpus of complete repositories annotated by hand, two independent reviewers, and the agreement between them published next to the score. Anything less measures the annotation process rather than the extractor.
Attack paths
Factual, but only once someone states what the attacker starts with. “An attacker can reach the database” is neither true nor false until you say whether they begin as an anonymous visitor, an authenticated user, or a compromised build runner. The same engine is right under one assumption and wrong under another, so a figure without that assumption printed beside it is not reproducible.
Contextual prioritisation
A claim about what matters more, and there is no fact of the matter independent of an objective. A pre-revenue startup and a regulated bank read the same repository and are both right to rank it differently. Any correct order we published would be our own product opinion, validated against a ground truth we also wrote.
The decision engine
Act now, plan, verify: a recommendation against a policy, and the policy is a product decision. We can test that it follows its own policy, that no finding is silently dropped, and that every recommendation cites evidence that exists. That is a conformance test. Calling it accuracy would invite exactly the misreading this page is written to prevent.

How to read the figures

Four things worth knowing before you draw a conclusion from any single number on the benchmark page.

Sample size before percentage
Each layer shows how many findings its precision was computed from. The layers with small samples are the recently added ones, and their figures will move as their corpus grows. This cuts against the flattering rows as much as the awkward one: a perfect score over nine findings is a weak claim, not a strong one.
Every false positive is published
Findings the engine produces on code that is not vulnerable are listed with the analyzer at fault and the reason the finding is wrong. They are counted against precision, never excluded from it. Three of them fire on the mitigation the rule itself recommends, which is precisely the kind of defect a benchmark exists to surface rather than absorb.
The LLM layer analyses your code
It reads an application's own integration code for LLM and agent mistakes: model output reaching a shell, a credential placed in a prompt, an agent tool that can run anything. It is not a measure of an AI inside Cybseco. This benchmark runs the deterministic engine with no model in the loop.
Known gaps are named
A vulnerability that is real and outside the engine's current reach is recorded as an accepted gap, excluded from recall, and listed. Recording a limit is how a future improvement shows up as a measured gain instead of being claimed today.

The rules we hold ourselves to

These are constraints on us, not claims about the product. They are what make the figures worth reading.

  • Anything Cybseco announces is either measured or named on this page as out of scope. There is no third state, and an automated gate fails the build when an analyzer ships without one or the other.
  • Ground truth describes the code, decided before the engine runs. A label written to match engine output is not evidence.
  • A figure is published with the sample it was computed from.
  • A defect we know about is published, not excluded from the metric. A gate also fails when a published defect quietly stops reproducing, so the list can only shrink honestly.
  • No comparison against another product on capabilities where we would be choosing their configuration, their corpus and the scenario. Nobody should believe such a comparison, including us.
  • No figure before its methodology is public and its ground truth inspectable.

See Cybseco reason about a real-world system

Watch the interactive demo, then request a guided trial for your team.