Fetching results...
Result not found.
How to run the benchmarkContainment by attack category
Per-category figures come from the published record. A category with no containment figure is a false-positive control set: those cases measure what the tool must let through, so containment does not apply.
One public corpus, every result
Every entry below runs the same public Agent Egress Bench corpus. The benchmark belongs to no vendor, so a score is comparable no matter who published it.
Corpus snapshot verified against agent-egress-bench@a3d56890487a.
Pipelock re-runs itself, in the open
Daily self-operated checks and human-reviewed publication records are tracked separately. A failed check never replaces the latest digest-checked record.
Latest self-operated check
Reading the latest workflow status.
Latest digest-checked record
Checking the published record against its published digests.
Published run history
Every entry is a self-run, self-reported claim, and the entries published by Pipelock are first-party evidence about Pipelock. This is not a ranking and no entry carries a mark awarded by the maintainer. The latest claimed result for each tool name within the newest 50 entries appears here, with earlier entries available in its history.
Loading results...
Run it yourself
Publish your own benchmark result
Run Agent Egress Bench against your own security tool, then publish and own the result yourself. You need no approval to publish one, including a result that reflects badly on Pipelock.
Every failed case, named first
An aggregate without its losses makes a concentrated gap look like noise. These rows are cross-checked against the hash-checked result file and case index.
Full-corpus containment
Full corpus is the procurement view: a case the tool does not claim still counts against it. This is a measurement, not a pass mark, a rank, or a certification.
Coverage
Containment by attack category
Categories with a miss or a false positive are pulled forward. A control set has no containment figure: those cases measure what the tool must let through.
Per-category figures come from the published record. Full per-category detail, including non-scoring field-presence diagnostics, stays in the raw record below.
Scope and provenance
What the tool declared it supports, and the versions and digests that pin the exact tool, runner, corpus, and profile behind this measurement.
The raw record
The exact published bytes behind every figure on this page. Nothing above is computed from anything this record does not carry.
Run it yourself
Publish your own benchmark result
You need no approval to publish a result, including one that reflects badly on Pipelock. Run the public corpus against your own tool and own the number.