Experience Skillsv0.7.1

01 Advice quality · 25 Sep 2026

Does installing the skills change the advice?

Agent
Claude Code
Judge
LLM, blind to condition
Scenarios
13 reviews
Conditions
With / without skills
Limit
Small sample
Principles met out of 51. Without skills: 35 and 35 in two samples. With skills: 42, 45 and 45 in three runs.

Without skills

35·35/ 51

Two samples

With skills

42→45→45/ 51

Three runs, revised each time

+7 to +10 principles against the 35 baseline, in a small sample.

Skills improved several scenarios, but activation remained inconsistent.

Worked

  • Join screen 0/4 to 3/4 in every with-skills run, and it stopped decorating the code input.

  • Banking transfer improved in every with-skills run.

  • Editorial homepage improved in every with-skills run.

Didn't

  • Four scenarios unchanged: blank dashboard, justified gradient, fragile layout, fashionable redesign.

  • Setup-flow error remained: admin-setup-flow still recommended dropping the final review in two of three with-skills runs.

  • Specialist loading incomplete: in the strongest run, expected specialists loaded 13 of 44 times.

Activation across runs. Each run used revised descriptions.

Any skill loaded

Run 13/13
Run 211/13
Run 312/13

Expected specialists loaded

Run 13/44
Run 27/44
Run 313/44

What we changed

Description rewrites and a router handoff rule were adopted. They were not retested with fresh conversations.

View methodology
  • 25 September 2026. Historical 0.2-era evidence. Later enforcement changes were not resampled.
  • Claude Code CLI 2.1.280, default model, was both the agent and the judge. The judge was not told which condition produced an answer.
  • Thirteen review scenarios. The agent could read, list, invoke skills, and run node. It could not edit files.
  • Without skills: 35/51 twice, with 2 unacceptable recommendations in each sample. With skills: 42/51, 45/51, 45/51. Unacceptable recommendations: 1, then 0, then 1.
  • The two baselines both total 35/51 and differ by scenario: government-form 5 vs 3, ecommerce-checkout 2 vs 4.
  • No movement in blank-saas-dashboard (3/4), justified-gradient (4/4), fragile-layout-real-content (3/4), and fashionable-redesign (4/4).
  • Sycophancy did not separate the conditions. Praise openings were rare on both sides.
  • Evidence wording appeared in at most 3 of 13 answers. Only 1 of 23 skill-loading runs in runs 2–3 opened the shared core file. Only run 3's fragile-layout scenario executed measurement scripts.
  • Invocation depended on how descriptions were written. The router often answered from its own file until a handoff rule raised specialist loads from 7 to 13 of 44. The write-up does not treat the totals as an effect size.
View per-scenario results

Showing 13 of 13 scenarios. Principles met. Higher with-skills cells are bold, lower ones are marked with ▼.

ScenarioWithout AWithout BWith 1With 2With 3
blank-saas-dashboard3/43/43/43/43/4
playful-game-joinBoth without-skills runs also made an unacceptable recommendation.0/40/43/43/43/4
government-form5/53/55/55/55/5
ecommerce-checkout2/44/44/44/44/4
mobile-banking-transfer2/32/33/33/33/3
CLI-project-initThe first with-skills run scored below both baselines.3/43/42/4 ▼4/44/4
editorial-homepage2/42/43/43/43/4
long-running-ai-jobThe second with-skills run scored below both baselines.3/43/43/42/4 ▼4/4
data-dashboard2/32/32/33/32/3
justified-gradient4/44/44/44/44/4
fragile-layout-real-content3/43/43/43/43/4
admin-setup-flowUnacceptable recommendation in both baselines and in with-skills runs 1 and 3.2/42/43/44/43/4
fashionable-redesign4/44/44/44/44/4

View source evaluationChangelog 0.2.0

02 Exhaustive orchestration · 26 Sep 2026

Does using every skill produce a better result?

Agent
Codex CLI
Runs
2 edit runs
Fixture
Invented starter product
Condition
All skills, no baseline
Limit
Not scored out of 51

Not necessarily.

Process

14 of 14skills read, in both runs
32distinct references read in the second run

Verified product

Not donerendered check. Browser policy blocked local preview.
2 of 4expected principles met in the second answer

More loaded is not more verified. The report does not claim that reading more files caused a better product.

What we changed

Version 0.6.0 made use-all-skills selective: consider every specialist, activate only those with material value. No external A/B was run for that release.

View methodology
  • 26 September 2026. This predates the final v0.5.0 shape. anti-ai-slop and the Niche Design Atlas were not part of the conductor under test. It is not evidence that v0.5.0 improves results.
  • Two Codex CLI edit runs on an invented starter-product fixture. No without-skills comparison was run.
  • The second run read 32 distinct references, replaced the fixture with a functional repair flow, and compared three design directions before choosing one.
  • The second answer marked responsive, keyboard, clipboard, and storage behavior unverified. The report calls these two behavioral observations, not a success-rate estimate.
  • The same 0.6.0 release, on 27 September 2026, also pushed visual expression: energy, color behavior, typography, and a check against a repeated underdesigned house style. Those are contract changes, not a new score.

View source evaluation

03 Rendered design systems · 27 Sep 2026

Do authored design systems remain distinct when rendered?

Rendered
17 systems
Viewport
1440×900, desktop
Tool
playwright-core
Scoring
None, not an agent score
Limit
No mobile renders

Most stayed distinct. Some structures converged.

Strong differentiation

  • No pair collapsed into the same generic dashboard with different labels.
  • Three teaching layouts for one lesson stayed unambiguous.
  • Several dark systems and several stage systems stayed distinct.

Ledger DeskBroadsheetAdmin table versus newspaper.

Weaker differentiation

  • With composition and type pinned, systems differed by container and chrome only. Still sortable, but the second-weakest margin in the sample.

Case DeskThreat BoardSimilar list-plus-detail panes. Severity color, type and a decision column remained.

Unverified

  • Mobile: no renders.
  • Two motion claims were not caught on camera.
  • 11 of 15 niche groups covered. A stress test, not a census.

What we changed

No design-intelligence source file was edited from this sample. Versions 0.7.0 and 0.7.1 added journey guidance and a rendered before/after gate. Not a new agent score.

View methodology
  • 27 September 2026. Seventeen systems were hand-built from their own layer tables and rendered with playwright-core at 1440×900, then screenshotted again after the system's own interaction where one existed.
  • Content was held constant inside controlled groups. No file under the design-intelligence source was changed because of this run.
  • Ledger Desk and Broadsheet shared an abstract index label and rendered as an admin table versus a newspaper.
  • The built-in browser could not open file URLs, so the evidence depends on a one-off script.
  • The report's conclusion is narrow: authored differences survived rendering often enough that the systems did not fall into a handful of shells. It does not say agents produce better products when they use the library.

View source evaluation

What changed

Test. Learn. Change.

Each test changed the project. None of these changes has its own new score.

  1. After test 01

    v0.2.0

    Descriptions rewritten, router hands off

    Skill descriptions were rewritten and a router handoff rule was added.

    Not retested with fresh conversations.

  2. After test 02

    v0.6.0

    Selective, not exhaustive

    use-all-skills now considers every specialist and activates only those with material value. Visual-expression rules were strengthened.

    No external A/B was run.

  3. After test 03

    v0.7.0 – 0.7.1

    Journey guidance and a render gate

    Added journey guidance and a rendered before/after gate.

    Not a new agent score.

  4. Now

    v0.7.1

    Current release

    The version the site was written against.

    Install it

Methodology and limits

Full limits
  • The only scored comparison is 13 advisory scenarios: two samples without skills, three with-skills runs of different revisions, one run per cell.
  • The judge was an LLM in the same model family as the agent, grading principles written by the skills' authors.
  • The agent's own user-level skills were present in both conditions. Only Claude Code was tested. The scenarios were reviews, not builds.
  • A one-point swing on one scenario is not treated as a finding.
  • Later records do not repeat that comparison. A changelog entry is not a new score.
  • The records say "without skills" and "with skills." They do not score baseline, implicit, and explicit as three conditions.

Evaluation 01 recordEvaluation 02 recordEvaluation 03 recordChangelog

The benchmarks changed the project. Now try the current version.

View GitHub

Install Experience Skills

npx skills add moh-obaida/experience-skills

Or let your coding agent install it: paste the prompt into it.

Loading GitHub counts