[PLAN-0035] Compono Agent Skill Pack¶
Status: Done
Implements: ADR-0035
Goal¶
A skills/compono/ skill, installable via npx skills add <owner>/compono, that makes an AI coding agent noticeably better at writing, modifying, reviewing, and troubleshooting Compono-based unit tests than one relying only on pretrained knowledge — verified by evals that the skill correctly activates on genuine Compono work, stays silent on ordinary non-Compono .NET test work, and every API/example it cites is real and current.
Scope¶
Per ADR-0035's Decision Outcome: one skill (skills/compono/SKILL.md) with package-conditional references/. In scope:
skills/root structure installable vianpx skillsSKILL.md— detection, routing, default workflow, guardrailsreferences/— composition model, registrations/profiles/scopes, diagnostics, xunit-v3, nsubstitute, bogus, patterns-and-antipatterns (file boundaries may be renamed/consolidated during Group 1 based on actual content density, per ADR-0035's explicit non-freeze on the list)skills/compono-evals/(wasskills/compono/evals/— moved out so it never ships vianpx skills add, see Notes) — positive/negative activation + correct-behavior scenarios- Root
README.mdupdate (Compono packages table area) documenting the skill's existence and install command - A
docs/*.mdpage (or section) explaining what the skill is, how to install/update it, and that it's agent guidance, not runtime behavior
Explicitly deferred (not this plan):
- Cross-agent packaging beyond
npx skillscompatibility (Copilot/Codex/ Cursor-specific marketplace files) — ADR-0035 doesn't require it; revisit only if it becomes low-cost and clearly wanted - Any Compono runtime/API change — if implementation surfaces a real doc or API defect, it's called out and scoped as its own fix, not folded in here
- A second skill for any future integration package — the escape-hatch principle in ADR-0035, not work to do now
One atomic PR, not phase-per-PR: the sections below are grouped by concern (scaffold, reference content, evals, docs, verification) for readability, not as independent phase boundaries each shipping its own PR. design-decisions.md's "each phase ships as its own PR" rule applies to a large or multi-milestone effort where one PR for the whole thing would be unreviewable — this plan's total diff (one new skill directory plus a handful of doc-nav updates) doesn't meet that bar, and splitting it into five artificially-sequenced PRs would have been fragmentation for its own sake, not genuine independent reviewability. Below, "Task groups" replaces the earlier "Phases" framing to avoid implying a shipping promise this plan never intended to keep.
Task groups¶
Group 0 — Skill scaffold and detection/routing¶
-
skills/compono/SKILL.mdfrontmatter (name, pushydescriptionwithUSE FOR/DO NOT USE FOR/SCOPES TO), Detection table (package refs, attribute/API grep signals, confidence), default workflow (recognize → inspect → decide → act → validate), hard guardrail section (no reflection fallback, noActivator .CreateInstance, no silent AutoFixture substitution) - Skeleton
references/files created (empty sections, filled in Group 1)
Group 1 — Reference content¶
-
references/composition-model.md—Composer,Create<T>()/CreateMany<T>(),[Composable], discovery, determinism/seeding -
references/registrations-profiles-and-scopes.md—Register<T>(),For<T>().Use()/.Member(),ICompositionProfile,[Shared], recursion -
references/diagnostics.md— CMP0001–CMP0012 table, runtimeCompositionExceptiontree-path/seed format, reproduce-a-failure workflow -
references/xunit-v3.md,references/nsubstitute.md,references/bogus.md— package-conditional integration guidance -
references/patterns-and-antipatterns.md— guardrail catalog + AutoFixture concept-mapping table - Consolidate/rename any reference file whose content turned out too thin to justify a standalone file (per ADR-0035's non-freeze note) — all 7 files carry enough distinct content to stand alone; no further consolidation needed
Group 2 — Evals¶
Evals must prove three independent things, not just "does it trigger": activation (fires on genuine Compono work, stays silent otherwise), routing/reference selection (loads only the reference files the detected packages warrant), and behavioral correctness (the guidance it gives is actually right). Each scenario in skills/compono-evals/evals.json is tagged with which of the three it targets.
- Activation scenarios — agent activates for genuine Compono work; agent does not activate for ordinary xUnit/NSubstitute/Bogus usage with no Compono involvement; agent does not unilaterally introduce Compono into a project that doesn't reference it
- Routing scenarios — agent only recommends
Compono.NSubstituteguidance when that package is referenced; agent only recommendsCompono.Bogusguidance when that package is referenced - Behavioral-correctness scenarios — agent never invents a Compono API; agent does not introduce AutoFixture as a substitute when Compono is already in use; agent does not "fix" a composition failure with reflection or
Activator.CreateInstance; agent respects registration/rule precedence (duplicateRegister<T>()is a conflict, not an override); agent understands[Shared]correctly (type-keyed,Compono.XunitV3-only, resolves first); agent knows when not to use Compono (a hand-built value is clearer than composing one, even in a Compono-using project) - 18 scenarios total in
skills/compono-evals/evals.json, each taggedactivation/routing/behavioral-correctness - Manual spot-check pass — 6 of 18 scenarios (covering all three categories, including the AutoFixture-introduction, reflection-workaround, and when-not-to-use-Compono scenarios) run as one-off subagent prompts, self-graded by the subagent against the eval's
expectations, with no independent grader pass. All 6 read as passing on inspection. Real signal, but explicitly not/skill-creator's documented eval workflow — no<skill-name>-workspace/run directories, no with-skill/baseline pairing, nograding.json/timing.jsonartifacts, no aggregatedbenchmark.json. - Run a
/skill-creator-style eval workflow across all 18 scenarios — with-skill + baseline pairs (1 run each, not/skill-creator's default 3), independent grading (separate grader subagent per scenario, not self-graded),benchmark.json/benchmark.mdaggregated viascripts.aggregate_benchmark. Precisely what is and isn't retained, so this isn't overclaimed:eval_metadata.jsonand each run'soutputs/response.mdwere generated during the run, but only in ephemeral scratch space — not committed to the repo, and not durable evidence a future reader can inspect.timing.json/metrics.jsonwere never captured at all, not merely omitted from commit —/skill-creator's workflow calls for per-run timing/token data captured live from subagent completion notifications, and this run didn't do that step, sobenchmark.md's Time/Tokens columns are genuinely empty, not just unpublished. What is retained and durable:grading.jsonper scenario per variant (the actual pass/fail/evidence record) and the aggregatedbenchmark.json/benchmark.md, committed atskills/compono-evals/benchmarks/2026-08-07/. Result: 97.4% pass rate with the skill (38/39 assertions) vs. 56.4% without it (22/39) — real, evidence-backed in the grading.json sense, but a genuinely lighter-weight run than/skill-creator's full documented workflow, not that workflow itself. See that directory's README for the full limitations list (single run per config, baseline wasn't repo-isolated, no timing data) and the eval-quality feedback the graders surfaced for a futureevals.jsonrevision.
Group 3 — Installation UX and docs¶
- Layout matches the convention microsoft/aspire-skills uses successfully (a top-level
skills/<name>/SKILL.md, no separate manifest file required — see Notes). This is evidence the shape is right, not evidence the install path actually works end to end. - Real
npx skills addrun, actually executed (not inferred from layout convention). GitHub's URL parsing can't disambiguate a branch name containing slashes (feat/skills-add-...) from a subpath, so a.../tree/<branch>/skillsURL against this PR's branch isn't a valid target string — ran against the local checkout instead (npx skills add /Users/.../compono/skills -a claude-code -y), which exercises the identical discovery/copyDirectorycode path, just skipping the git-clone step. Real output:Found 1 skill→compono, installed to.claude/skills/compono/containing exactlySKILL.md+ the 7references/*.mdfiles — nothing fromcompono-evals/, confirming the move in the previous round actually keeps it out of the install payload.skills-lock.jsonrecorded the local source and a content hash. A second run against the real pushed remote branch (e.g.owner/repo#branch-namesyntax if the CLI supports it, or simply againstmainonce merged) would still be worth doing as a final sanity check, but the mechanism itself is now verified end-to-end, not assumed. - Update root
README.md - Add/update a
docs/*.mdpage: what the skill is, install/update instructions, supported agents, relationship to the NuGet packages —docs/getting-started/ai-agent-skill.md, linked from nav, Next Steps, and README - Cross-link from this plan's ADR and from the doc page back to each other
Group 4 — Verification and closeout¶
- Every API/attribute/type named in the skill grepped against
src/to confirm it's real and current — full sweep of every code example and every named symbol acrossSKILL.mdand all 7references/*.mdfiles (144 unique backtick-quoted identifiers enumerated and checked), not a spot-check - Every code example verified against current public API signatures (parameter order, overloads, defaults) — found and fixed one real defect:
xunit-v3.mdcited a non-existentBindingPlan .ValidateSignaturemethod (the actual type isinternal sealed class BindingPlanwith aSignatureErrorproperty, no such method) — rewritten to describe the observable behavior without naming the internal type or an invented member - Confirm the skill never references internal implementation types, generator internals, test-only helpers, or any API that's visible in the repository but not intended for consumers — swept for this specifically;
PlanCache<T>,NSubstituteProvider,BogusMemberNameProvider,ProfileCycle,UniqueValueResolver,ICompositionContext.Resolve<T>()(descriptor-less overload) are all confirmedpublicand already part of the published API reference site, so describing them is fine; added one clarifying note incomposition-model.mdthat the descriptor-takingResolve<T>(...)overload is generated-code-only, not something to hand-write; confirmedCMP0003's "historical/rare, not reached via ordinary composition" claim againstLeafTypeClassifier .IsProviderResolved(interfaces/abstract/delegate types are classified provider-resolved before ever reachingConstructorSelector, so itsCMP0003checks for those shapes are unreachable via the normal discovery path) - Links resolve —
mkdocs build --strictclean, no warnings/errors - Confirm ordinary non-Compono test work doesn't trigger the skill — eval scenarios 8/9/10/14 (activation category), all spot-checked clean
- Confirm optional-integration guidance only fires when that package is referenced — eval scenarios ⅗/18 (routing category); 3 and 5 spot-checked clean, 18 documented not run live (same pattern as 3)
-
dotnet build/dotnet test—dotnet build Compono.slnxclean (0 warnings, 0 errors).dotnet test Compono.slnx(both Debug and the documented-c Release) fails with a Microsoft Testing Platform handshake error across every test project, including ones this plan never touches — confirmed to be a localdotnet testCLI orchestration issue, not a real test failure, by running each compiled test executable directly instead of through thedotnet testdriver:Compono.Tests(213/213),Compono.Generators.Tests(84/84),Compono.XunitV3.Tests(47/47),Compono.NSubstitute.Tests(23/23),Compono.Bogus.Tests(63/63) — 430/430 passing. No.cs/test/files are touched by this plan, consistent with a pre-existing local-environment issue rather than a regression from this change; worth a separate look (CI almost certainly isn't affected, since it presumably isn't hitting this handshake failure on every PR, but that's an assumption, not verified here). - Set
Status: Done, closeout note. Both items that kept this atIn Progressare now genuinely resolved: the eval run (honestly scoped as/skill-creator-style, not its full workflow — see Group 2 and the benchmark README for exactly what is/isn't retained) and a realnpx skills addinstall verification (see Group 3). Two PR review rounds (Copilot, then Jonas/j-d-hatwice) each found real, confirmed issues — every one fixed, not disputed or downplayed. Closing this plan doesn't mean no further feedback is possible, only that every currently-known finding has been addressed and every checklist item reflects what's actually true, not what would be convenient to claim.
Critical Files¶
skills/compono/SKILL.md— newskills/compono/references/*.md— new (7 files, subject to renaming)skills/compono-evals/*— newREADME.md— updated (skill install mention)docs/*.md— new or updated page documenting the skill packdocs/adr/0035-compono-agent-skill-pack.md,docs/adr/README.md,docs/plans/README.md— already updated
Test Plan¶
No .cs/runtime test changes expected — this is documentation/tooling content, not code. Verification is: skill-creator eval scenarios tagged activation/routing/behavioral-correctness (Group 2), a full (not spot-checked) manual API-signature and public-vs-internal accuracy sweep of every code example (Group 4), link resolution, and confirming the existing dotnet build/dotnet test suite is unaffected (sanity check only, no new automated coverage needed since nothing in src//test/ changes).
Notes¶
Design-review round (before implementation proceeded far): the user reviewed the ADR/plan and asked for five refinements, all incorporated before/during implementation:
- Evals must prove activation, routing, and behavioral correctness independently, not just "does it trigger" —
skills/compono-evals/evals.json's 18 scenarios are now tagged by category, with explicit coverage for registration precedence,[Shared]semantics, never inventing an API, never introducing AutoFixture as a silent substitute, and never "fixing" a failure with reflection/Activator.CreateInstance. - Every code example verified against current public API, not spot-checked — done in Group 4; found and fixed one real defect (see Group 4).
- ADR-0035's escape-hatch principle reworded so a new integration package alone is explicitly not sufficient reason to split into a second skill — the test is whether it changes how an agent works, not just what API surface it adds.
- Added an explicit Group 4 verification step confirming the skill never teaches internal/generator-internal/non-consumer-facing API as something to use.
- Added eval scenario 15 (age-boundary test) proving the skill recommends literal values over composition when that's genuinely clearer, even in a Compono-adopting project — directly exercises the "When not to use Compono" section.
Eval execution: 18 scenarios authored across all three categories, 6 spot-checked live via subagents (one per category from the original set, plus all three of the new critical guardrail scenarios — reflection refusal, AutoFixture-swap refusal, when-not-to-use-Compono). All 6 passed clean on first run — no skill revision needed. Full with/without-skill benchmark matrix (all 18 × 2 configurations × N runs, per /skill-creator's complete workflow) deliberately deferred as disproportionate for a v0.1 skill pack; revisit if real-world usage surfaces triggering or accuracy problems the spot-checks didn't catch.
PR #63 Copilot review (post-merge-request): 5 inline findings, all confirmed real and fixed (commit f0a368b). Four were the same class of defect — Composer.Create<T>() written as if Create<T>()/CreateMany<T>() were static generics on Composer, when they're instance methods on the Composer the static, non-generic Composer.Create(...) returns (SKILL.md, composition-model.md, registrations-profiles-and-scopes.md, skills/compono-evals/evals.json) — notable for landing in a skill whose explicit point is teaching agents not to invent Compono APIs. The fifth was a real seed-type gap in diagnostics.md's reproduce-a-failure step: CompositionDiagnostic.Seed is ulong (an unseeded composer draws a full random 64-bit value) and doesn't always fit the int-typed WithSeed(int)/[Compose(Seed = ...)] reproduction APIs the way a Compono.XunitV3 row failure's seed always does.
Real defect found and fixed during Group 4: references/xunit-v3.md originally cited BindingPlan.ValidateSignature as the mechanism behind a runtime CompositionException for stacked Compose-family attributes. BindingPlan is internal sealed class BindingPlan with a SignatureError property — no ValidateSignature method exists at all. Rewritten to describe the observable behavior (fails at data-binding time, not compile time) without naming the internal type.
PR #63 human review (Jonas / j-d-ha, 🛑 Request changes): 11 inline findings (5 🐛, 6 ⚠️ per this repo's review-emoji convention), all confirmed real against source and fixed: - SKILL.md's frontmatter description was 1523 chars folded, over skill-creator's 1024-char validator limit — trimmed to 914. - diagnostics.md's seed-reproduction step (already partly rewritten for the Copilot round above) still implied a "supported reproduction path" for an out-of-int-range ulong diagnostic seed that doesn't actually exist — rewritten to say so plainly instead of hand-waving an alternative. - This plan's own "Phases" framing implied phase-per-PR shipping per design-decisions.md's rule, but all five landed in one PR — reframed as "Task groups" with an explicit note on why one atomic PR was the right call here (small, tightly-coupled scope, not a large/ multi-milestone effort). - The eval-workflow-completion and npx skills install-verification claims both overstated what was actually done — split each into a checked item for the real, narrower thing that happened and an unchecked item for the genuinely outstanding work; Status reverted from Done to In Progress accordingly. - The dotnet test claim was self-contradictory (claimed green, then admitted not independently run) — actually run; dotnet test's CLI driver hits a local Microsoft Testing Platform handshake error on every project (including ones this plan never touches), but every compiled test executable run directly passes clean (430/430 across the 5 core test projects) — recorded as a local-environment issue to look at separately, not a regression from this change. - SKILL.md's hardcoded 0.x.y-preview.N/--prerelease version claim was stale — the repo's actual published version policy has moved on; removed the hardcoded claim and pointed at installation.md instead of duplicating a fact that changes independently of this skill. - The no-retry CompositionException guardrail over-generalized — scoped to Compono's own deterministic generated/built-in path, with an explicit call-out that consumer-supplied factories/providers/ IServiceProvider can be genuinely non-deterministic. - composition-model.md's "rebuild throws away seed/config" rationale was inaccurate — corrected to distinguish a seeded rebuild (stays reproducible) from an unseeded one (draws a fresh random seed each time). - docs/getting-started/ai-agent-skill.md's Update section claimed re-running add overwrites an install, unverified — replaced with the real, documented npx skills update compono command. - docs/documentation-architecture.md still declared 5 Getting Started pages and omitted the new ai-agent-skill.md from its canonical tree — added an entry with audience/purpose/handoff, consistent with every other page's treatment.
Full /skill-creator benchmark run (2026-08-07, closing the eval-workflow gap Jonas flagged): ran the real workflow — 36 subagent runs (18 evals × with-skill/baseline), 18 independent grader subagents (one per eval, grading both variants against the eval's own expectations), aggregated via scripts.aggregate_benchmark. 97.4% pass rate with the skill (38/39) vs. 56.4% without (22/39) — a real, evidence-backed gap. Artifacts committed at skills/compono-evals/benchmarks/2026-08-07/ (summary + per-scenario grading, not raw transcripts, per the chosen scope). Honest limitations recorded in that directory's own README: one run per configuration rather than three, no timing/token capture, and a methodology gap multiple graders independently flagged — the baseline subagents kept full repo filesystem access even though told not to read the skill, and at least one (eval 9) still produced accurate Compono-specific terminology, likely by exploring the repo directly. That means the true skill-driven gap is probably larger than 97.4/56.4 against a genuinely repo-isolated baseline, not smaller. Graders also surfaced concrete eval-quality feedback (several assertions pass regardless of skill use) — recorded as a follow-up, not acted on in this pass. The one remaining Group 3 item (a real npx skills add run against a merge-ready ref) is still outstanding, so Status stays In Progress.
Real defect: evals/ was inside the installable skill directory (caught by the user after reading benchmark.md, not by any review round). Checked npx skills' actual install behavior against its real source (vercel-labs/skills, src/add.ts): a disk-based install does a recursive copyDirectory of the whole skill folder, excluding only .git — no .skillignore/manifest mechanism exists to exclude files. skills/compono/evals/ (18KB evals.json plus the 40-file benchmarks/2026-08-07/ directory) would therefore have shipped into every consumer's .claude/skills/compono/evals/ on npx skills add — pure internal-QA dead weight with no value to a consumer. Fixed by moving the whole directory to skills/compono-evals/ (sibling to skills/compono/, no SKILL.md so npx skills' discovery never offers it as an installable skill, and it sits outside skills/compono/'s own copy scope). All path references in this plan updated accordingly. This is exactly the kind of installation-payload question Group 3's still-open real npx skills add run (above) would also have needed to catch — another reason that item stays open rather than being treated as optional polish.
PR #63 second human review round (Jonas / j-d-ha, second 🛑 Request changes): 6 more inline findings (4 🐛, 2 ⚠️), all confirmed real and fixed: - diagnostics.md's troubleshooting step 6 still had the exact same unscoped "CompositionException is deterministic, don't retry" claim the first review round already fixed in SKILL.md's guardrail — I fixed one location and missed the duplicate. Scoped identically this time (consumer factories/providers/IServiceProvider can be non-deterministic, check those first). - evals.json's eval 4 prompt never established that Compono/Compono.XunitV3/Compono.NSubstitute were installed or that it's a [Compose] row, yet its assertions required a [Shared] recommendation — meaning a correctly package-gated decline would have failed the assertion. Added the missing context to the prompt. - composition-model.md's seed rationale was still wrong in a subtler way than the first round's fix: re-verified Composer.cs directly and found the fresh-random-seed behavior isn't about rebuilding via Composer.Create(...) at all — _configuration.Seed ?? CompositionSeed.Generate() is evaluated inside Create<T>()/ CreateMany<T>() themselves, so an unseeded composer draws a fresh seed on every individual call, whether or not the Composer instance is reused. Rewrote to say this precisely — reuse alone never makes unseeded calls correlated; only WithSeed(...) does. - skills/compono-evals/benchmarks/2026-08-07/README.md linked ../evals.json, which resolves to a nonexistent skills/compono-evals/benchmarks/evals.json — should be ../../evals.json (two directories up from benchmarks/2026-08-07/, not one). Fixed. - The eval-workflow completion claim was still imprecise about exactly what's retained: reworded to state plainly that eval_metadata.json/ per-run outputs/ existed only in ephemeral scratch (never committed), and timing.json/metrics.json were never captured at all, not merely left out of the commit — only grading.json + the aggregated benchmark.json/benchmark.md are the actual durable record. - Ran the real, outstanding npx skills add verification (the one Group 3 item genuinely left open through both prior rounds): GitHub's URL parsing can't disambiguate a slash-containing branch name from a subpath, so tested against the local checkout instead (same discovery/copyDirectory code path, npx skills add /Users/.../compono/skills -a claude-code -y) — real output: Found 1 skill → compono, installed to .claude/skills/compono/ containing exactly SKILL.md + the 7 references/*.md files, confirming compono-evals/ genuinely stays out of the install payload. This closes the last item that had kept Status at In Progress through both review rounds.