Open benchmark · maintained by PipeLab
Agent Egress Bench
A tool-neutral benchmark for testing agent egress controls against public attack cases with reproducible, offline-checkable results.
- Cases
- 251
- Scope
- Tool neutral
- Results
- Reproducible
- License
- Apache 2.0
What it measures
One case, one observed decision, evidence you can keep
- Expected block, observed block
- Pass
- Expected allow, observed allow
- Pass
- Wrong decision
- Fail
- No route, delivery, or decision
- Not scored
Corpus snapshot verified against agent-egress-bench@a3d56890487a.
01 · What a result adds
Attack cases and false-positive controls belong in the same run
Attack cases check whether the control stops an unsafe request and proves the decision it made. False-positive controls check whether ordinary traffic still works. A useful result needs both.
A missing route, delivery record, or decision isn’t counted as a pass. The result stays bound to the exact product, version, configuration, adapter, capability profile, corpus, and scoring version used for that run.
02 · Neutrality
PipeLab maintains it. Pipelock doesn't control it.
PipeLab also builds Pipelock, so the boundary matters. The benchmark publishes no ranking and awards no mark. Anyone can test another product or publish a result that reflects badly on Pipelock without asking PipeLab.
03 · Run it yourself