Can an agent that knows nothing about your system use it correctly on the first try? This tool finds out by running one.
An MCP server ships with documentation, and almost none of it reaches the agents that connect to it. Skills, READMEs, and contracts live in a repository. A remote client sees the protocol and nothing else: an initialize payload, a list of tools, whatever resources the server chooses to serve. So the better the documentation, the more confident the author, and the worse the surprise when a headless agent misuses a tool that was, to the author, clearly explained.
The author cannot see this failure, because the author cannot un-know the domain. The docstring reads fine to them. A linter cannot see it either. It can confirm that a tool has a description; it cannot tell you that the description's first line is the entire onboarding a remote caller will ever receive, or that one boolean parameter silently makes a write unreadable.
So this tool does not grade prose. It puts a real agent with no context in front of the server, gives it a task, and reads what the agent did. The number that comes out is called Prior: how much the agent had to already know to succeed. The optimum is zero.
Five axes, each scored 0 to 4. Four are levers. The fifth is the outcome, and it runs the other way: lower is better.
| Axis | The question | Optimum |
|---|---|---|
| Connect | What does a cold client receive before it acts? | An instructions block with an entry point |
| Errors | Do failures teach the next call, or drop the stream? | Typed, actionable, self-correcting |
| Workflow | Can the steps be run out of order? | Order is enforced, not merely documented |
| Disclosure | Is there a thin front door that points deeper? | Progressive: not a wall, not a void |
| Prior | How much must the agent already know? | Zero |
A Prior of 4 means the agent completed the task wrongly while reporting success. That is the worst outcome in the rubric, worse than an outright failure, because nothing downstream knows it happened.
1. Connect snapshot. Reads the server source and reports what a client receives on initialize and tools/list: the instructions block or its absence, the tools in registration order, resources, providers, and how many lines of explanation sit in a module docstring that never crosses the wire. Static, free, no credentials, no running server.
2. Repo inventory. Lists the skills and guidance the repository holds, the workflow order they claim in prose, and how much of it a remote caller can reach. A workflow that lives only in a skill file is not a workflow for a caller who never reads skill files, which is every remote client.
3. Headless trial. Spawns claude -p with only the target server attached, file and shell tools disallowed, and a strict MCP config, so the agent cannot read the repository. That isolation is the experiment: a remote caller cannot read your skills either, and the gap between the two situations is what the audit measures. The trial runs against a copy of any state in a temp directory. It spends money, because a real agent runs, and nothing in the tool runs it silently.
The transcript is then read for five tells: the agent guessed a value the server never offered; it stalled on something it could not discover; it succeeded by accident; it invented a fact and reported it confidently; it hit an error and could not recover. Scoring is a judgement against the rubric's described levels, not arithmetic. There is deliberately no script that averages five judgement calls into a number.
The connect snapshot of this tool's own server, run on the repository as published.
=== what a cold client receives from mcp/server.py ===
initialize:
serverInfo.name : 'legibility'
version : '0.1.0'
instructions : 'Audits an MCP server and its repository for legibility on
connect ... Start with audit_repo(path) for everything at
once ... This server does NOT run the headless trial ...'
tools/list (4, in registration order):
connect_snapshot What a cold MCP client receives from the server in `repo`
repo_inventory Skills, supporting files, cross-links and stated workflow ordering
audit_repo Both static halves at once, with the findings merged
trial_command The exact command that runs the headless trial. Returned, never executed.
resources: 0 prompts: 0
providers: ["SkillsDirectoryProvider(roots=ROOT / 'skills')"]
=== findings (connect axis) ===
1. OK: a skills provider is declared, so this repo's skills ARE reachable over the wire.
2. FIRST TOOL IN REGISTRATION ORDER is `connect_snapshot` ...
A real audit from 4 September 2026, run with the headless trial, against an internal recruiting-match server. Same day, the same agent was put in front of this tool's own server. Both reports are in the repository under skills/legibility-audit/examples/; the first is rendered here.
Trial run: yes. claude -p, sonnet, 11 turns.
Prior 3
The trial agent said it, unprompted, about its own work:
I got there via a noise-level exploratory ranking that happened to agree with the real answer. That's a lucky corroboration on a 3-role pool, not a demonstrated method.
That is the third tell, succeeded by accident, self-reported. Not a 4 only because it did not report success falsely: it caught itself. What let it catch itself is under "what already works".
| Axis | Score | The one change that raises it |
|---|---|---|
| Connect | 0 | Add instructions=. One parameter. |
| Errors | 3 | Refusals already name the side, the id, the directory and the count, and the agent acted on them. |
| Workflow | 2 | Order is discoverable only by reading all nine descriptions first. |
| Disclosure | 0 | 32 files of guidance in the repo; a remote caller reaches none. |
initialize: serverInfo.name = 'recruiting-match' version = None instructions = None tools/list: 9 tools. resources: 0 prompts: 0 providers: none not on the wire: a 28-line module docstring
The first thing an agent ever reads about this system is the first tool in registration order: "Candidate to roles: rank the JD pool against a candidate profile (full JSON as a string)." Read that knowing nothing. It does not say the pool may be fixtures, that the profile must be supplied because nothing will give it to you, or that one write has a trap.
Five skills, 32 files of guidance, none reachable. Including a skill written specifically for headless agents driving this service, visible to everyone except headless agents driving this service.
Task, with no domain context: work out who is on the bench, find the best-matching role, record the float, and report how much to trust it.
The trial is as strong on this side, and these are the parts a refactor would quietly remove. The warnings list in the payload was read and passed on: the agent told its reader that all three pool entries were test fixtures and not to act as if a real requisition existed. The calibration tool refused to let a bad ranking stand, and the agent downgraded its own claim on the strength of it. The listed write path recomputed its own score rather than trusting the caller's number.
The lesson across all three: every save came from a fact in the response body. Nothing was saved by documentation. That is the whole case for the instructions block and for serving the skills: the same channel, one step earlier.
The same harness, the same model, the same task shape, against two servers. The only variable is the surface.
| Recruiting-match server | This tool's own server | |
|---|---|---|
| Prior | 3 | 0 |
| Turns | 11 | 5 |
| Instructions block | absent | present, with an entry point |
| Skills over the wire | 0 of 32 files | all, served as resources |
| Outcome | right answer, wrong method, self-caught | right answer, calibrated confidence |
The legible surface took less than half the turns. Legibility is not only a correctness property; the agent spent six extra turns on the illegible server working out what it was looking at. On the legible one, it separated inspected fact from its own judgement without being asked, and surfaced the caveat that makes the method honest, quoting a source it could only have got over the wire:
I did not run the headless trial, so I have no evidence of what an agent actually does when it hits this gap. The server's own instructions are explicit that a report without a trial is a lint, not evidence.
That sentence is in the instructions block. It crossed the wire, the agent read it, and it changed what the agent claimed.
As a Claude Code plugin, which gives you the /legibility-audit skill and an MCP server exposing the static half as tools:
/plugin marketplace add DMG-Venture-Studio/mcp-legibility /plugin install mcp-legibility
Then, inside Claude Code:
/legibility-audit # audit the repo you are in /legibility-audit --target <path> # audit another repo /legibility-audit --no-headless # static only, spends nothing
Or clone the repository and run the scripts directly. Each declares its own dependencies inline, so uv run --script is the whole setup. Python 3.11 or later.
uv run --script skills/legibility-audit/scripts/probe_connect.py --repo <path> uv run --script skills/legibility-audit/scripts/inventory_repo.py --repo <path> uv run --script skills/legibility-audit/scripts/headless_trial.py --repo <path> --task "<what to do>" uv run --script mcp/server.py --selftest # the server audits its own repository
A report without the trial is a lint, and the report says so. The trial is the only step that produces evidence rather than opinion.
The fixes the audit recommends are all things FastMCP already provides: an instructions= parameter that rides the initialize response, a skills provider that serves a repository's skills as resources with progressive disclosure by default, typed errors and an error-handling middleware so failures reach the agent as sentences rather than dropped streams. The reference in the repository lists each mechanism and what it buys a cold caller.