AGENTS.md doesn't outperform skills in my agent evals, but agent DX does
In the 1980s, an interdisciplinary team of scientists, linguists, and artists was given an impossible job: design permanent warnings for a nuclear waste repository that had to remain understandable for the next 10,000 years.
They couldn’t assume future humans would speak our languages, share our symbols, or think the way we do. Every decision, what to say first, how to combine text and images and physical architecture, what to make salient, had to account for readers whose cognition was unknown.
They eventually coined the term nuclear semiotics for the discipline.

We are now facing the same problem on a much shorter timeline. Coding agents are reading our docs, and they “read” nothing like us. The discipline doesn’t have an agreed name yet. Modal is calling it Agent DX.
Recently, Modal posted a new role: Member of Technical Staff, Agent DX Researcher. From the listing:
“Most of the code deployed on Modal will be written by agents in the near future.”
They want someone to define quantitative objectives for agent productivity, design measurement systems, run evaluations, and feed the results back into the platform.
Meanwhile, engineers have spent weeks arguing about AGENTS.md versus Skills. Vercel shipped a post favoring AGENTS.md with some clever invocation tricks. Some vocal devs have been calling Skills the winner. It’s the wrong debate. The question it pretends to answer, where do I put instructions for the agent, is the least interesting question in agent DX. I can prove it with more than 1,500 trials.
I built a benchmark called ModalBench. The tasks were deliberately post-cutoff, version-sensitive Modal APIs, the exact kind of tasks where the model’s pre-training priors are actively wrong and the docs are the only signal that can save the agent from generating deprecated code. Claude Code in a sandbox, same model across seven conditions.
I started this project planning to replicate Vercel’s “AGENTS.md outperforms skills in our agent evals” post on Modal. Five conditions, the same spectrum Vercel ran on Next.js 16, and a clean story about which file format wins. By the third scale-up the top three conditions were statistically tied on pass rate and the story fell apart.
Baseline run, no docs, 4% pass. Best run, 92%. Between those numbers sit six conditions that each isolate one variable: whether the skill gets loaded, what it says on line one, whether Modal’s shipped skill or mine is doing the work, whether I route into the docs conditionally, whether I dispatch research to a sub-agent, whether one sentence of prose inside the skill is reordered. I’ll introduce each one at the point in the argument where it starts mattering, with a snippet of what the agent actually sees. The short version: most of the delta lives in a property of the prose the CLAUDE.md-vs-Skills debate never mentions.
The filesystem is a paradigm built for organization. What agents need is declaration: prose written knowing exactly which distribution it is going to be sampled from.
The three regimes
In my runs, alongside verifiers in code, I decided to bring in rubrics to comb through the deeper nuances. The rubric judge was Sonnet 4.6. Reading the trajectories from 1,599 trials, I started seeing the same patterns repeat enough times broke them into three regimes:
- Invocation. Did the agent load your docs at all? Pass rates stay at 4 to 10% when this fails, regardless of doc quality.
- Instruction-shape. Given the agent loaded the docs, do the docs point it at the current API? Pass rates stay at 57 to 80% when docs exist but don’t foreground the specific deprecation.
- Composition. Given the agent has the right APIs, can it wire them together? Pass rates plateau around 88% on multi-step tasks.
The regimes have an ordering: the ceiling of one is the floor of the next. If your skill isn’t being invoked, no amount of beautiful migration documentation inside it matters. If your skill is invoked but line 9 of a reference file recommends a deprecated API, no amount of <important if> routing saves you. If the agent has the right APIs but doesn’t understand that pip_install_from_requirements runs at build time, you’ve hit the composition ceiling and more docs won’t help.
Regime 1: Invocation is a cliff
The control first. The baseline condition is the empty workspace. No .claude/skills/, no CLAUDE.md, the agent works from priors and the task file alone. Pass rate, 4%.
On 2026-04-14, Modal shipped an official Claude Agent Skill as part of modal-auto-research-skills. I pinned it at commit e8ecdf47 and dropped it into .claude/skills/modal/ with no CLAUDE.md nudge. Here is what the agent is supposed to consult and mostly doesn’t:
---
name: modal
description: >
This skill provides guidance for building and deploying applications on the
Modal cloud platform. Use this skill whenever the user mentions Modal or has
Python code that imports the `modal` SDK...
---
Pass rate, 8%. Discoverability 2.1/10. The agent knows the file exists and ignores it.
But, if you were to add two lines of CLAUDE.md naming the skill:
Before writing code, first explore the project structure, then use the
`modal` skill for Modal platform work.
Pass rate, 57%. Vercel ran the same comparison on their Next.js evals and saw 53 move to 79 from the same intervention. This simple nudge is worth tens of percentage points on its own.
Before Modal shipped that skill, there was no first-party option, so I wrote my own. It pins at .claude/skills/modal-docs/. The base skill is deterministic templating, the body is Modal’s own docs snapshotted at generation time from modal.com/llms-full.txt via a plain httpx.get(). The distinctive content is a 404-line migration doc that foregrounds every post-cutoff API change next to its replacement:
## Deprecating Mount as part of the public API
# Old way (deprecated)
mount = modal.Mount.from_local_dir("data").add_local_file("config.yaml")
@app.function(image=image, mount=mount)
# New way
image = image.add_local_dir("data", "/root/data").add_local_file("config.yaml", "/root/config.yaml")
@app.function(image=image)
With the same kind of two-line CLAUDE.md nudge applied, this condition pins at 89.5%. Discoverability 9.98/10. Same nudge, different skill, same tasks, the ceiling moves from 57 to 89.5.
Invocation is worth 80 percentage points. It is the single cheapest intervention in agent DX, in both time and context. Modal’s official skill and my pinned skill behave identically on this axis. Passive invocation pins at 8 to 10%. Instructed invocation goes to roughly 100%.
It follows that every skill needs a machine-legible CLAUDE.md nudge. A specific sentence that names the skill and tells the agent when to load it. Two lines is enough. Without it, models are surprisingly hesitant to invoke anything.
Regime 2: The menu problem
Once invocation is solved, the next bottleneck is whether your docs actually tell the agent to do the right thing.
Your docs are a menu. A restaurant menu isn’t a random assortment. The first item sells the most. The dish with the photo outsells the text-only listing by an order of magnitude. The item the menu doesn’t mention, even if the kitchen can make it, doesn’t get ordered.
Coding agents read docs the way customers read menus. What’s first gets chosen. What has a worked example gets chosen. What isn’t listed, effectively does not exist, even when the kitchen can cook it.
Here is the test case. Modal shipped an official Claude Agent Skill on 2026-04-14, it was from their Autoresearch experiment’s GitHub repo. Its references/app-structure.md is 13 lines long. Line 9 reads:
Additional local dependencies can be included in the modal.Image definition, e.g. using modal.Image.add_local_file or modal.Image.add_local_directory.
That line is a two-item menu. What’s not on it: modal.Image.add_local_python_source, the correct Modal 1.x API for packaging importable Python modules into an image. The kitchen makes it. The menu doesn’t list it.
On the two cohort tasks that require add_local_python_source, the instructed upstream skill passes at 27.5% and 25%. 27 of 40 mount-removal failures write literally add_local_file("helpers.py", remote_path="/root/helpers.py"), the exact call the menu describes, rejected by the verifier because the task cohort was designed around the API the skill omits.
I couldn’t believe it, so I walked 15 failing trajectories by hand. 13 of them saw the correct API later in their own context. The agent read app-structure.md, then fetched modal.com/docs/guide/images, which explicitly names add_local_python_source alongside add_local_file. The agent then wrote add_local_file anyway.
One trajectory went further. Step 56 of trial mount-removal__5pinFiR:
AGENT RESPONSE: “looking at the documentation again, add_local_python_source is for modules that are importable locally. Since helpers.py is in the workspace root and is importable locally, I could use that instead. Let me revise the implementation to be cleaner.”
The next edit kept add_local_file and merely reordered the code. The agent explicitly reasoned toward the correct API and then did not use it. This is not a retrieval failure. It is not a comprehension failure. It is first-reading stickiness: the earliest coherent information the agent encountered won, even after the agent later read the correct version.
I didn’t believe it at first, thought I might’ve missed a bug. Which is why I ran an ablation.
The one-line ablation
I wanted to falsify three simpler explanations of the pattern above:
- Skill content weighs more than results fetched from the web, regardless of reading order.
- Short canonical text dominates longer ambient text on any axis.
- The model’s pre-cutoff prior for add_local_file was always going to win, and the skill is incidental.
Each of these predicts zero effect if you change only the order of mention inside the skill. So I did exactly that.
I cloned Modal’s official skill at commit e8ecdf47 to a byte-identical sibling under assets/source/official-skills/modal-reordered/. Same SKILL.md, same three reference files, same CLAUDE.md nudge, same 98-line total. One line of references/app-structure.md rewritten:
I removed the following:
- Additional local dependencies can be included in the modal.Image definition,
- e.g. using modal.Image.add_local_file or modal.Image.add_local_directory.
And replaced it with this:
+ Additional local dependencies can be included in the modal.Image definition.
+ Use modal.Image.add_local_python_source for importable Python modules
+ (e.g. modal.Image.add_local_python_source("helpers")), and
+ modal.Image.add_local_file or modal.Image.add_local_directory for data
+ files and directories.

Nothing else changed. I wired the reordered sibling as a new benchmark condition and ran it same-job against the original instructed condition, 400 trials.
Targeted-task pass rate: 21 out of 80 trials passed → 41 out of 80 trials passed. +25 percentage points. p ≈ 0.001.
API choice on those same 80 trials: add_local_python_source writes went 28 → 50 (+27.5pp). add_local_file writes went 44 → 19 (−31.25pp). The three alternative hypotheses all fall. You cannot move API choice by 27.5 percentage points by changing only the order of mention inside the skill, unless order of mention inside the skill is causally load-bearing.
This is the single most useful thing I learned from shipping this year, and it generalizes: the first coherent instruction an agent encounters has disproportionate weight on what it writes, even when later retrieved docs name a better API. If your docs bury the canonical API under a legacy menu and trust the agent to find the right thing later, it won’t. Not reliably. Not at N=200.
Why the first line wins
The mechanism is worth slowing down on. It goes way further than Modal.
I think two effects compound. The first is the primacy effect baked into the architecture. Early tokens participate in every subsequent attention computation across every layer. The information at the start of a reference file is doing work on every later token; the information at line 50 only conditions whatever comes after it. Influence isn’t uniform, it’s stacked.
The second is what I’d call demonstration placement bias. When a worked example sits inline next to one item, that item becomes the salient demonstration in the agent’s working representation. Nearby content without an example is effectively down-weighted in the same attention graph. The exampled item gets read carefully. Its neighbors get skimmed.
This shows up in the ablation as a second-order finding I did not predict. On the one task where the reorder regressed (local-file-copy, −17.5pp), I traced the failures to a specific sub-bucket: agents that wrote the simplest two-argument form of add_local_file and missed the copy=True kwarg. The reordered line gave add_local_python_source an inline example but left add_local_file as a plain mention in the same sentence. 5 of 7 of the new failures wrote add_local_file(“requirements.txt”, “/root/requirements.txt”), the shortest syntactically valid call, written by agents who never bothered to consult the parameter list. The API with the example got attention. The API sharing a sentence with it lost attention.
So “first coherent recipe wins” morphs sharply into “the recipe that gets an example wins, and the recipe sharing a sentence with an exampled recipe loses.”
Practically, this changes how you write reference files. Line 1 of every reference file is doing roughly 10x the work of line 50. Worked examples aren’t decoration, they’re attention allocators. The API you want the agent to pick should get the example. The API you want the agent to avoid shouldn’t share a sentence with the API that got the example. There’s an interesting open experiment here: a length-preserving reorder with no inline example, to isolate which lever (order OR example) is doing more of the work.
Regime 3: The ceiling of your information, and a platform-level fix
Even with perfect invocation and perfect instruction-shape, the best condition I ran pins around 92%. The remaining 8 points are where docs stop helping.
I walked every failed trial across 800 trials marked with high discoverability (judge score ≥ 7.5/10, verifier rejected). 74 trials clustered into three patterns:
- 58% are build-versus-runtime semantics misses. Agent wrote add_local_file without copy=True, so the file exists at container runtime but not at image-build time. pip_install_from_requirements runs during build, fails with FileNotFoundError. The agent understood the API. It didn’t understand when Modal runs things.
- 23% are multi-decorator wiring errors on build-decorator-removal. Agent knew run_function,
@modal.enter, and@app.cls, but missed wiring@app.cls(image=…)to the image containing run_function. Three APIs composed correctly, the fourth reference missed. - 14% are miscellaneous build errors and residual content-shape slips.
The first bucket is the interesting one, because it stops being a docs problem in the usual sense. It is a platform error-messaging problem. A FileNotFoundError at build time that said “this file was added with copy=False, did you mean copy=True?” would close most of the 58% for free, and would help human developers identically. The cheapest remaining points in agent DX on Modal are in the build-time error path, not in the docs. If Modal ships a platform-level “did-you-mean” layer on Image.pip_install_from_requirements, I’d expect the composition ceiling to move from 92% to about 96%.
There is a boundary where your docs stop and your runtime begins. On the docs side, the lever is prose. On the runtime side, the lever is error messages designed to be read by something that will sample its next edit from your error string. Both sides are agent DX. Most teams only invest in one.
Sub-agents are context firewalls and context bottlenecks
Discover with sub-agents. Compose in the parent.
Two harness patterns, both sitting on top of the same pinned .modal-docs/ tree as my prescriptive skill. The first is conditional routing: a CLAUDE.md fragment of <important if="…"> blocks, each pointing at one reference file and naming the canonical API alongside the deprecated alternatives. HumanLayer’s post on getting Claude to actually read your CLAUDE.md argued that <important if> wrappers override the default “may or may not be relevant” framing the harness stamps on plain CLAUDE.md content. I generate the blocks with Claude Opus 4.6 from the docs snapshot and check them into the repo. One example:
<important if="you are adding local files or Python source to a Modal image,
staging local data into a container, or migrating from Mount-based patterns
to image-based add_local_* methods">
Read .modal-docs/01-guide/10-data-sharing-and-storage/01-passing-local-data.md
Use modal.Image.add_local_file, modal.Image.add_local_dir, modal.Image.add_local_python_source.
Do NOT use modal.Mount.from_local_dir, modal.Mount.from_local_file, removed.
</important>
The second is sub-agent dispatch. HumanLayer’s harness-engineering post proposed sub-agents as “context firewalls”: isolate research-heavy work in a separate context window, return a condensed result, avoid context rot. My implementation hands the agent a docs root and a research sub-agent prompt template (also Opus-4.6-generated) and forbids the parent from reading .modal-docs/ directly:
Sub-agent prompt template:
SEARCH .modal-docs/ recursively for files relevant to {{QUERY}}
RETURN exactly these sections:
Current Recommended API, Code Example,
Deprecated / Superseded Alternatives, Source Files.
When I tested both, I found some unexpected results. The firewall works exactly where you’d predict. On build-decorator-removal, a multi-page cross-reference task where the agent must locate Image.run_function, @modal.enter, and with_options across three docs pages, sub-agent dispatch passed 36/40. Conditional routing passed 31/40. +12.5pp, clean and reproducible.
The firewall costs you, in exactly the place the firewall framing wouldn’t predict. On local-file-copy, a compositional task where the agent has to wire Image.add_local_file(copy=True) together with run_commands and a build-time invocation, sub-agent dispatch passed 28/40. Conditional routing passed 36/40. That is −20pp. The sub-agent returned a condensed summary, and what got condensed away was the exact cross-referent information the compose step needed.

Sub-agents firewall context. Compose steps need context. Split your harness by cognitive phase, not by knowledge domain. If your workflow has a “research this” stage and a “wire the APIs together” stage, the first should go through a sub-agent and the second should not.
One more thing: combining levers is anti-additive on this cohort. Running <important if> routing and sub-agent dispatch together pins at 87.5%, below either alone. When people talk about “stacking” harness patterns, additivity is not the default. Some patterns interfere.
The missing metric
I asked the LLM judge to score each trial’s discoverability on seven binary criteria, including applied_correct_api and avoided_deprecated. The judge read the skill content and the agent’s transcript.
On the best pinned-skill instructed condition, the judge thought the agent used the correct API on 19 of 21 failing trials. Those trials still failed the verifier, the actual deploy-and-execute check against Modal. About 9.5% of all trials were silently wrong by this measure: the docs agreed with the agent, the agent still produced broken code.
On Modal’s instructed official Skill, the judge thought the agent used the correct API on 64 of 86 failing trials. 32% of all trials fall in this silent-wrong bucket.
Under the hood, the judge is reading the same skill as the agent. If the skill doesn’t mention add_local_python_source, the judge doesn’t know to penalize its absence. The docs are scoring themselves. The only signal that catches the drift is the AST contract and live runtime check, the verifier that actually tests the output.
I’m proposing this as a new metric: Docs-Consistent Failure Rate (DCFR). The fraction of failing trials where an LLM judge, reading your docs and your agent’s transcript, thinks the agent did the right thing. It measures how much of your documentation is a bad oracle, prose that confidently endorses the wrong answer. I’ll leave the naming part to Mintlify.
The implication is sharper than it looks. If you’re evaluating agent performance with an LLM judge and your docs are in the judge’s context, your eval does not measure correctness. It measures docs-coherence-with-agent-output, which is trivially high when the docs are wrong in the same direction the agent is wrong. Every serious agent eval needs a verifiable oracle: AST check, live deploy, numerical equivalence, something the docs cannot lie to. Rubrics will keep being the de-facto for fuzzy tasks for good reasons, but the limitations are larger than most teams think. Spend on determinism before you lean on rubrics.
Agent DX is RLVR-shaped
A developer ergonomics discipline that treats docs as a prompt, trials as rollouts, verifiers as reward models, and ablations as gradients is functionally identical to RLVR. The loop is the same. The only difference is that the update target isn’t the model’s weights, it’s the docs. And because the update target is human-readable prose, the loop closes by humans writing better prose, not by SGD on a billion parameters.
This is what Modal’s job listing actually describes when it says “quantitative objectives, measurement, product improvements.” It is the RLVR loop, applied to developer surface area, with a human in the gradient step.
Three implications fall out once you take the framing seriously.
Your agent’s docs should be focused on the out-of-distribution sections.
A human reader skims, cross-references, fills gaps with common sense. A coding agent samples from a distribution conditioned on what it just read. If your API changed post-cutoff, the model’s priors have mass in the wrong place. Your job as a docs author is counter-conditioning: writing prose that moves the sample’s mass to where you want it, with the same attention to order, framing, and example placement you’d use in a careful prompt.
Every skill needs a verifier. An example says “here’s a correct usage.” A verifier says “here’s the AST contract the agent’s output must satisfy, here’s the runtime check it must pass.” If your docs ship without a verifier, you can’t close the loop. You cannot know which prose changes helped and which hurt. For tasks where verification is genuinely fuzzy, ship a mini-rubric or a checklist the agent can self-check against.
Ablations are the unit of progress. In ML, you don’t ship a model change without an eval. Why are we shipping docs changes without one? The cost of a ModalBench-style ablation is dropping fast enough that “every docs PR ships with its N=200 pass-rate delta” is going to be the norm within a year on serious platforms. The teams that get there first will compound. Their docs will improve along a measurable gradient while everyone else is arguing about bullet formatting.
This is also where the CLAUDE.md versus Skills debate dissolves. Skills are a file format. AGENTS.md is a file format. Neither matters independent of whether the prose inside is designed for the generation distribution and wired to a verifier. You can ship great agent DX in either format. You can ship bad agent DX in both. Format can only take you so far, point me to someone still using TOON.
Why the loop only closes on certain platforms
Everything in this post is RL infra in a trench coat: Trials are rollouts, conditions are policy priors, AST checks are verifiers, pass rates are rewards, one-line ablations are the gradient. The structure is identical to RLVR as practiced on code generation today.
What changes between platforms is the cost of running the loop. Modal already runs it. Their sandboxes are the rollout environment. Their platform provides the verifier substrate. Their customers, the people doing RL training and batch inference and autoresearch, already operate in a world where “run the agent at N=200, grade each run, iterate” is the normal mode of production.
This is why the Agent DX Researcher role makes structural sense to me. The job description (“define quantitative objectives, design systems to measure performance, translate results into product improvements”) only works out economically on a platform where the measurement loop is cheap. Good developer experience sets you up to build good agent DX, but the former doesn’t guaranteee the latter
Honest caveat: My 5-task cohort is deliberately a worst case for pointer-style skills. Every task hits a post-cutoff API change. On the stable-API tasks that dominate real agent work, Modal’s official skill almost certainly saturates with a strong harness. The benchmark does not say the official skill is bad. It says the official skill is pointer-shaped, most skills are, and on the cohort that punishes pointer-shaped skills most, the design choice costs about 33 percentage points. The interesting question is what fraction of your users’ tasks look like my cohort.
How to improve your Agent DX
- Every skill you ship needs a CLAUDE.md nudge that names the skill and tells the agent when to load it. If you have a skill without a nudge, you shipped an artifact nobody is reading.
- Read line 1 of every reference file. Is it the canonical API for the task, or a legacy mention with the canonical API buried later? Rewrite. The first line is doing 10x the work.
- If any of your APIs changed post-cutoff, ship counter-conditioning. The model’s priors are actively wrong. A 404-line migration doc that foregrounds deprecated APIs next to their replacements was the single biggest source of lift in my pinned skill.
- Track Docs-Consistent Failure Rate alongside pass rate. If your LLM judge agrees with your agent more often than your verifier agrees with either, your docs are the bad oracle.
- Discover with sub-agents. Compose in the parent. Don’t stack levers without measuring; harness patterns interfere.
- Stop arguing about CLAUDE.md versus Skills. The file format is noise. Order of mention, worked-example placement, counter-conditioning against priors. Agent DX research is extremely immature, and there is a lot of ground to cover.
Readers that don’t exist yet
The CLAUDE.md versus Skills debate is a proxy for the real question: what does it mean to design for readers that don’t exist yet?
The readers exist already. They read in fundamentally different ways than humans. The patterns are measurable. We just have to start measuring.
As a scientist on that nuclear waste team saw the same limit we now face: no static text survives minds you cannot predict.
The only solution was continuous attention, a tradition that would renew the warning for every new generation of readers.
Originally published on X.