mcp-legibility

Can an agent that knows nothing about your system use it correctly on the first try? This tool finds out by running one.

Why inspection is not enough

An MCP server ships with documentation, and almost none of it reaches the agents that connect to it. Skills, READMEs, and contracts live in a repository. A remote client sees the protocol and nothing else: an initialize payload, a list of tools, whatever resources the server chooses to serve. So the better the documentation, the more confident the author, and the worse the surprise when a headless agent misuses a tool that was, to the author, clearly explained.

The author cannot see this failure, because the author cannot un-know the domain. The docstring reads fine to them. A linter cannot see it either. It can confirm that a tool has a description; it cannot tell you that the description's first line is the entire onboarding a remote caller will ever receive, or that one boolean parameter silently makes a write unreadable.

So this tool does not grade prose. It puts a real agent with no context in front of the server, gives it a task, and reads what the agent did. The number that comes out is called Prior: how much the agent had to already know to succeed. The optimum is zero.

The method

Five axes, each scored 0 to 4. Four are levers. The fifth is the outcome, and it runs the other way: lower is better.

AxisThe questionOptimum
ConnectWhat does a cold client receive before it acts?An instructions block with an entry point
ErrorsDo failures teach the next call, or drop the stream?Typed, actionable, self-correcting
WorkflowCan the steps be run out of order?Order is enforced, not merely documented
DisclosureIs there a thin front door that points deeper?Progressive: not a wall, not a void
PriorHow much must the agent already know?Zero

A Prior of 4 means the agent completed the task wrongly while reporting success. That is the worst outcome in the rubric, worse than an outright failure, because nothing downstream knows it happened.

Three steps

1. Connect snapshot. Reads the server source and reports what a client receives on initialize and tools/list: the instructions block or its absence, the tools in registration order, resources, providers, and how many lines of explanation sit in a module docstring that never crosses the wire. Static, free, no credentials, no running server.

2. Repo inventory. Lists the skills and guidance the repository holds, the workflow order they claim in prose, and how much of it a remote caller can reach. A workflow that lives only in a skill file is not a workflow for a caller who never reads skill files, which is every remote client.

3. Headless trial. Spawns claude -p with only the target server attached, file and shell tools disallowed, and a strict MCP config, so the agent cannot read the repository. That isolation is the experiment: a remote caller cannot read your skills either, and the gap between the two situations is what the audit measures. The trial runs against a copy of any state in a temp directory. It spends money, because a real agent runs, and nothing in the tool runs it silently.

The transcript is then read for five tells: the agent guessed a value the server never offered; it stalled on something it could not discover; it succeeded by accident; it invented a fact and reported it confidently; it hit an error and could not recover. Scoring is a judgement against the rubric's described levels, not arithmetic. There is deliberately no script that averages five judgement calls into a number.

What the static half sees

The connect snapshot of this tool's own server, run on the repository as published.

=== what a cold client receives from mcp/server.py ===

initialize:
  serverInfo.name : 'legibility'
  version         : '0.1.0'
  instructions    : 'Audits an MCP server and its repository for legibility on
                     connect ... Start with audit_repo(path) for everything at
                     once ... This server does NOT run the headless trial ...'

tools/list (4, in registration order):
  connect_snapshot   What a cold MCP client receives from the server in `repo`
  repo_inventory     Skills, supporting files, cross-links and stated workflow ordering
  audit_repo         Both static halves at once, with the findings merged
  trial_command      The exact command that runs the headless trial. Returned, never executed.

resources: 0   prompts: 0
providers:  ["SkillsDirectoryProvider(roots=ROOT / 'skills')"]

=== findings (connect axis) ===
  1. OK: a skills provider is declared, so this repo's skills ARE reachable over the wire.
  2. FIRST TOOL IN REGISTRATION ORDER is `connect_snapshot` ...

Example report

A real audit from 4 September 2026, run with the headless trial, against an internal recruiting-match server. Same day, the same agent was put in front of this tool's own server. Both reports are in the repository under skills/legibility-audit/examples/; the first is rendered here.

Legibility audit: recruiting-match server

Trial run: yes. claude -p, sonnet, 11 turns.

Prior 3

The trial agent said it, unprompted, about its own work:

I got there via a noise-level exploratory ranking that happened to agree with the real answer. That's a lucky corroboration on a 3-role pool, not a demonstrated method.

That is the third tell, succeeded by accident, self-reported. Not a 4 only because it did not report success falsely: it caught itself. What let it catch itself is under "what already works".

AxisScoreThe one change that raises it
Connect0Add instructions=. One parameter.
Errors3Refusals already name the side, the id, the directory and the count, and the agent acted on them.
Workflow2Order is discoverable only by reading all nine descriptions first.
Disclosure032 files of guidance in the repo; a remote caller reaches none.

What the cold client receives

initialize:  serverInfo.name = 'recruiting-match'   version = None   instructions = None
tools/list:  9 tools.  resources: 0   prompts: 0   providers: none
not on the wire: a 28-line module docstring

The first thing an agent ever reads about this system is the first tool in registration order: "Candidate to roles: rank the JD pool against a candidate profile (full JSON as a string)." Read that knowing nothing. It does not say the pool may be fixtures, that the profile must be supplied because nothing will give it to you, or that one write has a trap.

What the repo holds that the wire does not carry

Five skills, 32 files of guidance, none reachable. Including a skill written specifically for headless agents driving this service, visible to everyone except headless agents driving this service.

The trial

Task, with no domain context: work out who is on the bench, find the best-matching role, record the float, and report how much to trust it.

  • Guessed a value the server never offered. No tool returns record text, so it ranked on a placeholder query and got a score of 0.195, which is noise.
  • Stalled on something it could not discover. "There's no tool that lists bench records directly." It enumerated the bench by ranking against a throwaway job description and reading the ids out of the result. Discovery by side effect.
  • Succeeded by accident. See Prior.
  • Invented a fact and reported it confidently. Did not happen, and that is a finding below.
  • Could not recover from an error. Recovered from every refusal it hit.

Findings, worst first

  1. There is no way to read a record, and the primary tool requires one. The ranking tool takes a whole profile as JSON; nothing returns one. A remote caller cannot rank honestly. This is a missing capability, not a missing explanation, and no instructions block fixes it. Fix: a getter and a lister.
  2. No instructions block. The entire onboarding is nine tool descriptions in registration order. Fix: the template in the repository, and it must state finding 1 until finding 1 is fixed.
  3. 32 files of guidance, none reachable. Fix: serve the skills as resources. The default mode preserves the front-door-then-depth shape over the wire.
  4. Bench enumeration only via side effect. Closed by finding 1.

What already works

The trial is as strong on this side, and these are the parts a refactor would quietly remove. The warnings list in the payload was read and passed on: the agent told its reader that all three pool entries were test fixtures and not to act as if a real requisition existed. The calibration tool refused to let a bad ranking stand, and the agent downgraded its own claim on the strength of it. The listed write path recomputed its own score rather than trusting the caller's number.

The lesson across all three: every save came from a fact in the response body. Nothing was saved by documentation. That is the whole case for the instructions block and for serving the skills: the same channel, one step earlier.

The A/B, same day

The same harness, the same model, the same task shape, against two servers. The only variable is the surface.

Recruiting-match serverThis tool's own server
Prior30
Turns115
Instructions blockabsentpresent, with an entry point
Skills over the wire0 of 32 filesall, served as resources
Outcomeright answer, wrong method, self-caughtright answer, calibrated confidence

The legible surface took less than half the turns. Legibility is not only a correctness property; the agent spent six extra turns on the illegible server working out what it was looking at. On the legible one, it separated inspected fact from its own judgement without being asked, and surfaced the caveat that makes the method honest, quoting a source it could only have got over the wire:

I did not run the headless trial, so I have no evidence of what an agent actually does when it hits this gap. The server's own instructions are explicit that a report without a trial is a lint, not evidence.

That sentence is in the instructions block. It crossed the wire, the agent read it, and it changed what the agent claimed.

Install and run

As a Claude Code plugin, which gives you the /legibility-audit skill and an MCP server exposing the static half as tools:

/plugin marketplace add DMG-Venture-Studio/mcp-legibility
/plugin install mcp-legibility

Then, inside Claude Code:

/legibility-audit                      # audit the repo you are in
/legibility-audit --target <path>      # audit another repo
/legibility-audit --no-headless        # static only, spends nothing

Or clone the repository and run the scripts directly. Each declares its own dependencies inline, so uv run --script is the whole setup. Python 3.11 or later.

uv run --script skills/legibility-audit/scripts/probe_connect.py --repo <path>
uv run --script skills/legibility-audit/scripts/inventory_repo.py --repo <path>
uv run --script skills/legibility-audit/scripts/headless_trial.py --repo <path> --task "<what to do>"
uv run --script mcp/server.py --selftest       # the server audits its own repository

A report without the trial is a lint, and the report says so. The trial is the only step that produces evidence rather than opinion.

Built on FastMCP's own mechanisms

The fixes the audit recommends are all things FastMCP already provides: an instructions= parameter that rides the initialize response, a skills provider that serves a repository's skills as resources with progressive disclosure by default, typed errors and an error-handling middleware so failures reach the agent as sentences rather than dropped streams. The reference in the repository lists each mechanism and what it buys a cold caller.