This page carries Pipelock’s own Agent Egress Bench runs, disclosed as first-party evidence. It is not a ranking and it does not score other tools. Anyone can run the same corpus against any target and publish what they find, including a result that reflects badly on Pipelock, with no notice or approval. The results-use and attribution policy defines the labels to use and the facts to publish beside a score, and the launch write-up explains why the scores are published this way.

agent-egress-bench
251 logical cases
18 categories
182 block-expected attack cases
68 allow-expected false-positive controls
1 warn-class drift case

Corpus snapshot verified against agent-egress-bench@a3d56890487a.

The corpus covers DLP evasion, prompt injection, SSRF, tool poisoning, encoding chains, shell obfuscation, hostname exfiltration, WebSocket DLP, and A2A scanning. The same corpus Pipelock tests itself against before every release.

This tests the security tool, not the agent. View the scoring methodology.

251

logical cases

18

categories

182

block-expected attack cases

68

allow-expected false-positive controls

1

warn-class drift case

How to run Agent Egress Bench yourself

Run the reviewed Pipelock release from a clean Linux clone:

git clone https://github.com/luckyPipewrench/agent-egress-bench.git
cd agent-egress-bench
./scripts/run-pipelock-gauntlet.sh

For another security tool, build the shared runner and supply that tool’s capability profile and adapter:

git clone https://github.com/luckyPipewrench/agent-egress-bench.git
cd agent-egress-bench/runner
go build -o aeb-gauntlet .
./aeb-gauntlet --cases ../cases --profile my-profile.json --output results.json

See the adoption guide for building a runner. You publish and own your result. Label the run and include the identifying facts from the results-use policy, then open a discussion if you want it looked at.

Ready to put a runtime boundary in front of your agents?