Open benchmark · maintained by PipeLab

Agent Egress Bench

A tool-neutral benchmark for testing agent egress controls against public attack cases with reproducible, offline-checkable results.

Cases
251
Scope
Tool neutral
Results
Reproducible
License
Apache 2.0

What it measures

One case, one observed decision, evidence you can keep

A benchmark case passes through an egress control and produces a scored, recorded outcome.
One case goes through the security control. The runner records delivery and the control's decision before assigning a score.
Expected block, observed block
Pass
Expected allow, observed allow
Pass
Wrong decision
Fail
No route, delivery, or decision
Not scored
251logical cases
18categories
182block-expected cases
68false-positive controls

Corpus snapshot verified against agent-egress-bench@a3d56890487a.

01 · What a result adds

Attack cases and false-positive controls belong in the same run

Attack cases check whether the control stops an unsafe request and proves the decision it made. False-positive controls check whether ordinary traffic still works. A useful result needs both.

A missing route, delivery record, or decision isn’t counted as a pass. The result stays bound to the exact product, version, configuration, adapter, capability profile, corpus, and scoring version used for that run.

02 · Neutrality

PipeLab maintains it. Pipelock doesn't control it.

PipeLab also builds Pipelock, so the boundary matters. The benchmark publishes no ranking and awards no mark. Anyone can test another product or publish a result that reflects badly on Pipelock without asking PipeLab.

03 · Run it yourself

The corpus and method are public