Benchmark · September 2026
Tested three ways.
Different questions. Different methods. Mixed results.
- 01Advice qualityDoes installing the skills change the advice?35 → 45 / 51Better, but activation was inconsistent
- 02Exhaustive orchestrationDoes using every skill produce a better result?Not necessarilyEvery skill read, no rendered check
- 03Rendered design systemsDo authored systems stay distinct when rendered?17 renderedMost stayed distinct, some converged
Limits
- Small samples.
- LLM-based judging.
- Each evaluation asked a different question.
- Historical revisions are not directly comparable.
01 Advice quality · 25 Sep 2026
Does installing the skills change the advice?
- Agent
- Claude Code
- Judge
- LLM, blind to condition
- Scenarios
- 13 reviews
- Conditions
- With / without skills
- Limit
- Small sample
+7 to +10 principles against the 35 baseline, in a small sample.
Skills improved several scenarios, but activation remained inconsistent.
Worked
Join screen 0/4 to 3/4 in every with-skills run, and it stopped decorating the code input.
Banking transfer improved in every with-skills run.
Editorial homepage improved in every with-skills run.
Didn't
Four scenarios unchanged: blank dashboard, justified gradient, fragile layout, fashionable redesign.
Setup-flow error remained: admin-setup-flow still recommended dropping the final review in two of three with-skills runs.
Specialist loading incomplete: in the strongest run, expected specialists loaded 13 of 44 times.
Any skill loaded
Expected specialists loaded
What we changed
Description rewrites and a router handoff rule were adopted. They were not retested with fresh conversations.
View methodology
- 25 September 2026. Historical 0.2-era evidence. Later enforcement changes were not resampled.
- Claude Code CLI 2.1.280, default model, was both the agent and the judge. The judge was not told which condition produced an answer.
- Thirteen review scenarios. The agent could read, list, invoke skills, and run node. It could not edit files.
- Without skills: 35/51 twice, with 2 unacceptable recommendations in each sample. With skills: 42/51, 45/51, 45/51. Unacceptable recommendations: 1, then 0, then 1.
- The two baselines both total 35/51 and differ by scenario: government-form 5 vs 3, ecommerce-checkout 2 vs 4.
- No movement in blank-saas-dashboard (3/4), justified-gradient (4/4), fragile-layout-real-content (3/4), and fashionable-redesign (4/4).
- Sycophancy did not separate the conditions. Praise openings were rare on both sides.
- Evidence wording appeared in at most 3 of 13 answers. Only 1 of 23 skill-loading runs in runs 2–3 opened the shared core file. Only run 3's fragile-layout scenario executed measurement scripts.
- Invocation depended on how descriptions were written. The router often answered from its own file until a handoff rule raised specialist loads from 7 to 13 of 44. The write-up does not treat the totals as an effect size.
View per-scenario results
Showing 13 of 13 scenarios. Principles met. Higher with-skills cells are bold, lower ones are marked with ▼.
| Scenario | Without A | Without B | With 1 | With 2 | With 3 |
|---|---|---|---|---|---|
| blank-saas-dashboard | 3/4 | 3/4 | 3/4 | 3/4 | 3/4 |
| playful-game-joinBoth without-skills runs also made an unacceptable recommendation. | 0/4 | 0/4 | 3/4 | 3/4 | 3/4 |
| government-form | 5/5 | 3/5 | 5/5 | 5/5 | 5/5 |
| ecommerce-checkout | 2/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| mobile-banking-transfer | 2/3 | 2/3 | 3/3 | 3/3 | 3/3 |
| CLI-project-initThe first with-skills run scored below both baselines. | 3/4 | 3/4 | 2/4 ▼ | 4/4 | 4/4 |
| editorial-homepage | 2/4 | 2/4 | 3/4 | 3/4 | 3/4 |
| long-running-ai-jobThe second with-skills run scored below both baselines. | 3/4 | 3/4 | 3/4 | 2/4 ▼ | 4/4 |
| data-dashboard | 2/3 | 2/3 | 2/3 | 3/3 | 2/3 |
| justified-gradient | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| fragile-layout-real-content | 3/4 | 3/4 | 3/4 | 3/4 | 3/4 |
| admin-setup-flowUnacceptable recommendation in both baselines and in with-skills runs 1 and 3. | 2/4 | 2/4 | 3/4 | 4/4 | 3/4 |
| fashionable-redesign | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
02 Exhaustive orchestration · 26 Sep 2026
Does using every skill produce a better result?
- Agent
- Codex CLI
- Runs
- 2 edit runs
- Fixture
- Invented starter product
- Condition
- All skills, no baseline
- Limit
- Not scored out of 51
Not necessarily.
Process
Verified product
More loaded is not more verified. The report does not claim that reading more files caused a better product.
What we changed
Version 0.6.0 made use-all-skills selective: consider every specialist, activate only those with material value. No external A/B was run for that release.
View methodology
- 26 September 2026. This predates the final v0.5.0 shape. anti-ai-slop and the Niche Design Atlas were not part of the conductor under test. It is not evidence that v0.5.0 improves results.
- Two Codex CLI edit runs on an invented starter-product fixture. No without-skills comparison was run.
- The second run read 32 distinct references, replaced the fixture with a functional repair flow, and compared three design directions before choosing one.
- The second answer marked responsive, keyboard, clipboard, and storage behavior unverified. The report calls these two behavioral observations, not a success-rate estimate.
- The same 0.6.0 release, on 27 September 2026, also pushed visual expression: energy, color behavior, typography, and a check against a repeated underdesigned house style. Those are contract changes, not a new score.
03 Rendered design systems · 27 Sep 2026
Do authored design systems remain distinct when rendered?
- Rendered
- 17 systems
- Viewport
- 1440×900, desktop
- Tool
- playwright-core
- Scoring
- None, not an agent score
- Limit
- No mobile renders
Most stayed distinct. Some structures converged.
Strong differentiation
- No pair collapsed into the same generic dashboard with different labels.
- Three teaching layouts for one lesson stayed unambiguous.
- Several dark systems and several stage systems stayed distinct.
Ledger DeskBroadsheetAdmin table versus newspaper.
Weaker differentiation
- With composition and type pinned, systems differed by container and chrome only. Still sortable, but the second-weakest margin in the sample.
Case DeskThreat BoardSimilar list-plus-detail panes. Severity color, type and a decision column remained.
Unverified
- Mobile: no renders.
- Two motion claims were not caught on camera.
- 11 of 15 niche groups covered. A stress test, not a census.
What we changed
No design-intelligence source file was edited from this sample. Versions 0.7.0 and 0.7.1 added journey guidance and a rendered before/after gate. Not a new agent score.
View methodology
- 27 September 2026. Seventeen systems were hand-built from their own layer tables and rendered with playwright-core at 1440×900, then screenshotted again after the system's own interaction where one existed.
- Content was held constant inside controlled groups. No file under the design-intelligence source was changed because of this run.
- Ledger Desk and Broadsheet shared an abstract index label and rendered as an admin table versus a newspaper.
- The built-in browser could not open file URLs, so the evidence depends on a one-off script.
- The report's conclusion is narrow: authored differences survived rendering often enough that the systems did not fall into a handful of shells. It does not say agents produce better products when they use the library.
What changed
Test. Learn. Change.
Each test changed the project. None of these changes has its own new score.
After test 01
v0.2.0
Descriptions rewritten, router hands off
Skill descriptions were rewritten and a router handoff rule was added.
Not retested with fresh conversations.
After test 02
v0.6.0
Selective, not exhaustive
use-all-skills now considers every specialist and activates only those with material value. Visual-expression rules were strengthened.
No external A/B was run.
After test 03
v0.7.0 – 0.7.1
Journey guidance and a render gate
Added journey guidance and a rendered before/after gate.
Not a new agent score.
Now
v0.7.1
Current release
The version the site was written against.
Install it
Methodology and limits
Full limits
- The only scored comparison is 13 advisory scenarios: two samples without skills, three with-skills runs of different revisions, one run per cell.
- The judge was an LLM in the same model family as the agent, grading principles written by the skills' authors.
- The agent's own user-level skills were present in both conditions. Only Claude Code was tested. The scenarios were reviews, not builds.
- A one-point swing on one scenario is not treated as a finding.
- Later records do not repeat that comparison. A changelog entry is not a new score.
- The records say "without skills" and "with skills." They do not score baseline, implicit, and explicit as three conditions.
Evaluation 01 recordEvaluation 02 recordEvaluation 03 recordChangelog
The benchmarks changed the project. Now try the current version.
Install Experience Skills
npx skills add moh-obaida/experience-skillsOr let your coding agent install it: paste the prompt into it.
Manual installInstallation guide
Loading GitHub counts