← Back to Home
Agentic Engineering

AGENTS.md Outperforms SkillsInside Vercel's agent evals

Vercel hardened their agent evals and found an 8KB compressed docs index inside AGENTS.md hit a 100% pass rate—beating skills-based retrieval that stalled at 79% even with explicit trigger instructions. Here’s the setup, why passive context wins right now, and how to apply it to your own projects.

01
The problem
Outdated training data vs. fast-moving stacks
Platform drift

New APIs and runtime behaviors land faster than model training cycles; older projects risk the inverse (models proposing capabilities they don’t ship). Agents hallucinate or miss config changes.

Two teaching models

Skills package tools + docs for on-demand retrieval. AGENTS.md injects project instructions on every turn. Vercel built both: a reusable skill and a compressed docs index stored in AGENTS.md.

02
Eval setup
Hardened, behavior-first tests
Scope
Fresh platform APIs

Coverage included new caching primitives, request/response helpers, authorization branches, async request state, and proxy routing—APIs unlikely to be in model training data.

Method
Behavior over shape

Removed leaky tests and asserted observable outcomes (build, lint, test) with retries to smooth model variance.

Variants
Four configurations

Baseline (no docs), skill (default), skill with explicit invocation instructions, and compressed docs index injected into AGENTS.md.

03
Results
Passive context wins decisively
Pass rates
Baseline
53% overall; 84/95/63 across build/lint/test.
Skill
53% without instructions; 79% with “explore project, then invoke skill” guidance (95/100/84 breakdown).
AGENTS.md
100% across build, lint, and test using an ~8KB compressed docs index plus a “prefer retrieval-led reasoning” note.
Why skills underperformed

Skills weren’t invoked in 56% of cases. Trigger wording was brittle: “invoke skill first” fixed some tasks but missed config edits; “explore then invoke” worked better yet stayed below parity with passive context.

04
Why passive context helps
No decision points, consistent availability
Zero trigger risk

Instructions live in the system prompt every turn—no tool invocation required.

Stable ordering

Avoids sequencing pitfalls (docs first vs project first) that change behavior.

Context discipline

Compression to ~8KB preserved a 100% pass rate while keeping a minimal footprint.

05
Apply it to your repo
One command setup
Command
# add compressed docs index to AGENTS.md
echo "[Docs Index]|root: ./framework-docs|IMPORTANT: Prefer retrieval-led reasoning|...compressed index..." >> AGENTS.md
Detect
Read your runtime/framework version.
Download
Store matching docs locally (e.g., ./framework-docs/).
Inject
Adds the compressed docs index + “prefer retrieval-led reasoning” line to AGENTS.md.
Example
AGENTS.md (repo root)
# Repo: ./AGENTS.md
[Docs Index]|root: ./framework-docs|IMPORTANT: Prefer retrieval-led reasoning over pre-training-led reasoning
01-guides:{00-intro.md,01-routing.md,02-data.md}
02-apis:{cache.md,cookies.md,headers.md,auth.md}

Place AGENTS.md at the repository root so agents read it every turn; keep compressed entries pointing to locally checked-in docs under ./framework-docs/.

Example
Skill package (./skills/framework-docs/skill.json)
{
  "name": "framework-docs",
  "description": "Framework docs with caching/auth examples",
  "version": "1.0.0",
  "docs": "./framework-docs",
  "tools": []
}

Place skills under ./skills/<skill-name>/skill.json; the agent must choose to invoke this skill to read the same docs your AGENTS.md index references.

06
Practical guidance
What to do today
Checklist
1Run the codemod and commit AGENTS.md + .next-docs/ (or your framework equivalent).
2Keep the index compressed; avoid pasting full docs unless your agent supports long contexts.
3Retain a single instruction: “Prefer retrieval-led reasoning over pre-training-led reasoning.”
4Use skills for explicit, vertical workflows (upgrades, migrations) rather than general framework literacy.
For framework authors

Ship a ready-to-paste AGENTS.md snippet and a compressed index for your docs. Users get correct, version-matched hints without relying on tool invocation reliability. Skills can stay as opt-in, workflow-specific accelerators.

Takeaway
Reduce decisions, raise pass rates

Until agents reliably trigger tools, keep the critical context in AGENTS.md. Compress aggressively, point to local docs, and reserve skills for tasks that need explicit user intent.