Version A · 14,401-tool catalog

Find any tool. Call only what was approved.

An agent can search every tool, but it gets authority for exactly one call: one tool, one argument set, once. Anything that can change state waits for a person first.

Allowed or executedWaiting for a personRefused

Workbench

Run a call through the boundary yourself. Every button talks to the running server, and the diagram above follows along.

  1. 1Search
  2. 2Choose a tool
  3. 3Bind arguments
  4. 4Authorize and run

1 · Search

Results appear here. Only tools the policy allows are ever returned.

2 · The call

Choose a search result to see what the agent may send.

Waiting for approval

Decision log

Verified

go run ./cmd/verify replays the whole boundary and writes a receipt. The last full run passed every gate.

14/14
retrieval cases
16/16
grant scenarios
15/15
MCP protocol checks
9/9
product flows

Search time as the catalog grows

Median of 200 queries, policy allowing every tool. Filtered HNSW only overtakes exact BM25 at 14,401 tools, and not once encoding the query (29.6 ms on CPU) is counted.

Show as a table
ToolsExact BM25Filtered HNSW
5120.38 ms5.47 ms
2,0482.88 ms10.59 ms
8,1927.01 ms7.29 ms
14,40112.81 ms7.41 ms

Limits

Local, single process
Grants live in one process and approval is a local action. There is no identity provider, tenant model or shared grant store.
Admitted once, not exactly-once
A call is admitted once. If the downstream tool times out, whether its effect happened is unknown.
Risk comes from tags
Catalog tags (version A) or the classification file (version B) decide risk. A wrong tag gives a wrong decision.
Top-level argument checks
Required fields, types, enums, lengths and closed schemas are checked; nested rules are not.
Loopback only
The server binds to 127.0.0.1 and checks the request origin. It is not an authenticated network service.
Measured on one machine
Scale timings are single-process and sequential, with no independent relevance labels.