1What each system is
The static scores.
Three of six properties can be read from a repository. Each is 0 to 10. Not measured means the instrument could not read the evidence, and says why.
Shapes before ranks. A React library, a Sass framework and a component tree inside a product are different materials. Adherence was measured against the product for three systems and the docs and examples for the rest; the column does not rank across them. Open a row for the detail.
| System |
What an agent finds to read |
Single source |
Legibility |
Adherence |
Mean |
Trial · informed / cold |
|
| Ghost (Shade + Admin X Design System)React, inside the product |
AGENTS.md · 19 skills · Storybookadherence: the product |
10 |
9 |
9 |
9.3of 3 |
not trialledno installable package |
→ |
| PostHog Lemon UIReact, inside the product |
django-python.mdc, react-typescript.mdc, rust.mdc +2 · 93 skills · Storybook · MCPadherence: the product |
8 |
10 |
10 |
9.3of 3 |
not trialledno installable package |
→ |
| GOV.UK FrontendSass and Nunjucks, framework-agnostic |
no context file · no skillsadherence: the product |
5 |
9 |
—not measured |
7of 2 |
4 informed4 cold, completeness |
→ |
| Mozilla ProtocolSass and HTML, framework-agnostic |
AGENTS.md, CLAUDE.md · no skillsadherence: docs and examples |
6 |
8 |
—not measured |
7of 2 |
4 informed0 cold, completeness |
→ |
| IBM CarbonReact library |
AGENTS.md · no skills · Storybookadherence: docs and examples |
3 |
10 |
5 |
6of 3 |
9 informed4 cold, completeness |
→ |
| Siemens iXStencil web components |
copilot-instructions.md, AGENTS.md · 4 skills · Storybookadherence: docs and examples |
5 |
10 |
3 |
6of 3 |
7 informed7 cold, completeness |
→ |
| Bootstrap ItaliaSass and vanilla JS, framework-agnostic |
no context file · no skillsadherence: docs and examples |
3 |
7 |
—not measured |
5of 2 |
7 informed2 cold, completeness |
→ |
| Home Assistant frontendLit web components, inside the product |
copilot-instructions.md, AGENTS.md · 12 skillsadherence: the product |
7 |
3 |
—not measured |
5of 2 |
not trialledno installable package |
→ |
| Microsoft Fluent (v9)React library |
AGENTS.md · 13 skills · Storybookadherence: docs and examples |
5 |
7 |
3 |
5of 3 |
7 informed7 cold, completeness |
→ |
| Primer ReactReact library |
copilot-instructions.md · 15 skills · Storybook · MCPadherence: docs and examples |
3 |
8 |
4 |
5of 3 |
7 informed7 cold, completeness |
→ |
Context files: AGENTS.md, CLAUDE.md, copilot instructions, cursor rules, llms.txt, at the root or beside the system.
Skills: SKILL.md files under a conventional skills directory.
Mean: of the measured properties only.
Calibration reference. Astryx, the system the instrument was developed against, is never a ranked row. Under the same conditions: single source 9, legibility 10, adherence 4, mean 7.7. The prediction made before any run held on the first two. Detail · Page.
2What an agent did with them
Reading the docs helped. It did not fix how the screen was put together.
Draw 1 of 3 · 16 of 16 screens judged · second read pending
One agent, claude-opus-4-8 pinned. One frozen task: a notifications settings screen. Every system built twice, cold (told nothing) and informed (pointed at its docs). The distance between the arms is the arm gap.
0/8cold runs that opened the documentation unprompted. All 8 read the package's source or types instead.
3/7systems that did better informed at the two planted gaps. None did worse.
At the boundary (16 runs, two gaps each): anticipated by the system 10, named and composed around 1, composed silently 12, hand-rolled silently 2. The silent middle is the habit.
Composition: 6 of 16 screens fail at least one of six questions, and reading the documentation did not help: 2 of 8 cold against 4 of 8 informed. The calibration pair reproduced the earlier read: informed Astryx used the layout primitives correctly and lost on containment and the viewport edge. A finding about defaults, not documentation.
Per system
Cold band, informed beneath, and under the constraint band the number of violations counted in each arm. Arm gap is informed minus cold on completeness. Each cell is a single run. This is draw 1 of 3, so there is no median and no range yet; the number is that one run, and draws 2 and 3 will show how far it moves.
| System | Constraint | Legibility | Completeness | Interaction | Arm gap | Looked unprompted |
| IBM Carbon |
12 informed41 → 24 violations |
28 informed |
49 informed |
87 informed |
+5on the gaps |
0/1 |
| Microsoft Fluent (v9) |
43 informed8 → 16 violations |
28 informed |
77 informed |
98 informed |
0on the gaps |
0/1 |
| Primer React |
13 informed69 → 19 violations |
22 informed |
77 informed |
88 informed |
0on the gaps |
0/1 |
| Siemens iX |
13 informed72 → 11 violations |
25 informed |
77 informed |
58 informed |
0on the gaps |
0/1 |
| Bootstrap Italia |
11 informed143 → 73 violations |
27 informed |
27 informed |
66 informed |
+5on the gaps |
0/1 |
| GOV.UK Frontend |
24 informed22 → 9 violations |
210 informed |
44 informed |
810 informed |
0on the gaps |
0/1 |
| Mozilla Protocol |
12 informed147 → 20 violations |
24 informed |
04 informed |
78 informed |
+4on the gaps |
0/1 |
| Astryx calibration reference, not ranked | 28 informed29 → 1 violation | 210 informed | 1010 informed | 86 informed | 0on the gaps | 0/1 |
Reading the constraint column. A run is banded twice, on how many kinds of the system's rules it broke and on how many times it broke them, and the worse of the two is the score. Breaking one rule ninety times and breaking three rules once each are both violations, and neither can be traded against the other. Of the 16 runs, 10 were held down by volume, 1 by kinds, and 5 scored the same on both. The counts sit under each band, and each run's page shows both readings. This is a published deviation from the rubric frozen 2026-07-22, made 2026-09-10 after the first draw was scored: as pre-registered the band read kinds alone, so an arm could cut its violations sharply and still band lower. Taking the worse of two readings can only lower a score and never raise one, and over these 16 runs no band rose, 10 fell and 6 were unchanged. Both the old and the new bands are in the results file.
Shapes before ranks, here too. Of the 7 systems, 3 are React libraries, 3 are framework-agnostic and 1 is a web-component library. A JSX screen built against a typed React library and an HTML screen built against a Sass framework are not made of the same materials, so a column compares more safely within a shape than across one. Each system was scaffolded in its own shape rather than forced into a React wrapper, which is what makes the runs fair to each system and the ranking weaker between them.
What the counts never saw. Constraint counts five kinds of violation. Across the 14 ranked runs, hardcoded dimensions in 14, hardcoded colours in 8, raw elements where a component exists in 4, and invented names and foreign UI imports in none. Not one agent invented a component name from a system's own namespace, and not one reached for a UI library other than the system under test. The two failures the rubric treats as most serious did not happen.
Judged by an agent under authorisation; a second read is pending. Rules: Methods. Every count, answer and screenshot is on the system pages.
3How to read a score
Eleven bands, five words.
Each property reduces its evidence to an index from 0 to 1; the index maps onto a band. Each word covers two bands.
Absent
Incidental
Partial
Substantial
Strong
Systematic
0 to 5 until 2026-09-08; re-banded from the same indexes, nothing re-measured. The trial's ladders followed on 2026-09-09.
4Methods
Pre-registered, pinned, reproducible.
4.1Protocol
Frozen 2026-07-22, before any run.The git history is the proof. Every change since is a published deviation.
4.2Static run
One instrument, one day, pinned commits.No configuration, no cooperation. Each system's directory pinned in the manifest. Re-running is one command.
4.3Readers
What it cannot read is part of the result.Tokens: CSS, DTCG, StyleX, Sass, CSS-in-JS. Components: TS, JS, Vue, Svelte, Astro, template directories. Adherence: TS and JS only. A Lit template or a Nunjucks view is outside it today, and the rows say so.
4.4The instrument on trial
Five runs, not one.The first run scored adherence 5 of 5 for two systems with nothing to read and found GOV.UK's 39 components as two. Each defect was fixed with a test that fails against the old code, and the field re-run.
4.5Deviations
Six, all published.Interaction property added (2026-07-23). Usage-meter pretext reframed (2026-07-23). Static scale to 0 to 10 (2026-09-08). Trial ladders followed (2026-09-09). Interaction checklist frozen after generation, before scoring (2026-09-09). First judge is an agent, second read pending (2026-09-09).
5Limits
What this page cannot say.
5.1One draw
The protocol asks for three draws and a person. This is draw 1 and an agent. Both replace without changing anything else.
5.2Three of six
Constraint, fidelity and completeness are not read from source. The mean averages what was measured.
5.3Seven of ten
PostHog, Ghost and Home Assistant cannot be installed from a package, so they are not trialled. One screen, one model, one task.
5.4Surfaces
Adherence was measured against the code in the repository. GitHub.com is not in Primer's; www.gov.uk is not in GOV.UK Frontend's.
6Submit a system
Public systems only, the same protocol.
An installable package and a repository that can be pinned. Nothing scored privately, nothing for a fee, and a low score is published the same way as a high one.
Propose a system →
7Licence
CC BY 4.0. The instrument is described in the thesis and the repository it reads in Structure. Learey, C. (2026). Correct by Design: the index. correctby.design/index.
Ghost (Shade + Admin X Design System)
Ghost Foundation (non-profit; HQ jurisdiction unverified) · React, inside the product · measured at apps/shade
- Context files
- AGENTS.md, apps/shade/AGENTS.md
- Skills
- 19 SKILL.md files under a conventional skills directory
- Markdown docs
- 44 doc artifacts in the repository; 91 of 108 components documented, 82 colocated
- Other surfaces
- Storybook (3 configs, in the repository); docs directory: docs
- Tokens
- 693 across 33 authored sources; dark mode generated
The strongest single-source column in the field: 693 tokens, 91% of references resolving to one, dark mode generated. Adherence is first-party, measured against the admin apps that consume Shade: 87.5% of declarations resolve to a token, and 255 of 906 replaceable elements are raw HTML. 19 skills, one of which routes an agent to a page template rather than letting it invent chrome. Not trialled: the package was never published.
PostHog Lemon UI
PostHog Inc., US · React, inside the product · measured at frontend/src/lib/lemon-ui
- Context files
- .cursor/rules/django-python.mdc, .cursor/rules/react-typescript.mdc, .cursor/rules/rust.mdc, .cursorrules, AGENTS.md
- Skills
- 93 SKILL.md files under a conventional skills directory
- Markdown docs
- 82 doc artifacts in the repository; 65 of 76 components documented, 55 colocated
- Other surfaces
- Storybook (1 config, in the repository); MCP: .mcp.json; docs directory: docs
- Tokens
- 320 across 25 authored sources; dark mode generated
The product is the consuming surface, and it is the largest in the field: 12,808 application files, 3,698 of them importing lemon-ui, 92% of style declarations resolving to a token. That is a first-party adherence measurement and it is not comparable with the docs-site rows above. Legibility is carried by 5 agent context files and 93 skills, more than any other system, alongside 65 of 76 documented components. Not trialled: the published package is a placeholder.
GOV.UK Frontend
Government Digital Service (Cabinet Office), UK · Sass and Nunjucks, framework-agnostic · measured at packages/govuk-frontend/src/govuk/components
- Context files
- none at the root or beside the system
- Skills
- none
- Markdown docs
- 63 doc artifacts in the repository; 38 of 39 components documented, 38 colocated
- Other surfaces
- docs directory: docs
- Tokens
- 100 across 22 authored sources; dark mode none
Almost every component carries its own documentation (38 of 39, a README and a machine-readable YAML options file in each directory), which puts it in the top band for legibility with no agent context file and no skills at all. The tokens are Sass variables across 22 settings files with no dark mode. Adherence could not be measured: the in-repo consumer is a Nunjucks review app the instrument cannot read, and the JavaScript it can read never touches a style. The instrument found 201 application source files, 15 of them readable, and none import the design system, use one of its components, or carry a style declaration it can read. The trial recipe records the defining fact: there is no component layer in code. Components are HTML copied from the GOV.UK Design System website, so every structural decision on a screen is markup the agent writes itself.
Mozilla Protocol
Mozilla Foundation / MZLA, US · Sass and HTML, framework-agnostic · measured at components
- Context files
- AGENTS.md, CLAUDE.md
- Skills
- none
- Markdown docs
- 87 doc artifacts in the repository; 48 of 64 components documented, 37 colocated
- Other surfaces
- docs directory: docs
- Tokens
- 46 across 9 authored sources; dark mode unobservable
Measured as its Fractal component library: 64 template directories, 48 of them with a readme, plus an AGENTS.md and a CLAUDE.md at the root. The tokens are Sass settings, 46 of them across 9 files, and the HTML templates reference none by name, so the reference ratio is unmeasurable rather than zero. Adherence is not scored: the only in-repo consumer is the docs build. The instrument found 17 application source files, all 17 readable, and none import the design system, use one of its components, or carry a style declaration it can read. The trial recipe's finding is the sharpest in the field: the published package's README documents installation and nothing else, so the Sass entry points had to be found by reading the file tree.
IBM Carbon
IBM, US · React library · measured at packages/react
- Context files
- AGENTS.md, packages/react/AGENTS.md
- Skills
- none
- Markdown docs
- 290 doc artifacts in the repository; 344 of 350 components documented, 226 colocated
- Other surfaces
- Storybook (2 configs, in the repository); docs directory: docs
- Tokens
- 268 across 70 authored sources; dark mode unobservable
The most thoroughly documented library in the field by count: 344 of 350 components, 226 colocated doc artifacts, two agent context files. Single source scores low because the React source references no tokens by name (0) while carrying 412 raw values, almost all of them hex colours inside bundled SVG illustrations, and because the tokens live in 70 separate Sass files across the monorepo. Adherence was measured against Carbon's own examples and docs, where 38 of 110 replaceable elements are raw HTML.
Siemens iX
Siemens AG, Germany · Stencil web components · measured at packages/core
- Context files
- .github/copilot-instructions.md, AGENTS.md
- Skills
- 4 SKILL.md files under a conventional skills directory
- Markdown docs
- 3 doc artifacts in the repository; 197 of 200 components documented, 1 colocated
- Other surfaces
- Storybook (1 config, in the repository); docs directory: packages/documentation
- Tokens
- 332 across 35 authored sources; dark mode generated
197 of 200 components documented, two agent context files, 4 skills: legibility in the top band. Single source is middling: 332 tokens with a generated dark theme, but only 19% of the colour and size references in the Stencil source resolve to one. Adherence was measured against the framework test apps, where 60 preview examples re-declare a system component name and every replaceable element is raw HTML. The trial recipe records a four-line package README with no setup instructions.
Bootstrap Italia
Developers Italia / AgID + Dipartimento per la trasformazione digitale, Italy · Sass and vanilla JS, framework-agnostic · measured at src/js
- Context files
- none at the root or beside the system
- Skills
- none
- Markdown docs
- 92 doc artifacts in the repository; 62 of 64 components documented, 0 colocated
- Other surfaces
- docs directory: docs
- Tokens
- 466 across 7 authored sources; dark mode unobservable
Measured as its JavaScript behaviour layer, one file per component (64 found, 62 with a head comment), because the SCSS components are style-only partials the instrument does not treat as components. The tokens are real and numerous (466 Sass variables, 1 of them in one _variables.scss), but the JavaScript cannot reference a Sass variable, so the ratio is 0 references to 18 raw values. No agent context file, no skills. Adherence is not scored: nothing in the repo consumes the system. The instrument found 79 application source files, 52 of them readable, and none import the design system, use one of its components, or carry a style declaration it can read. The trial recipe records that the published Sass entry point does not compile against its own dependency.
Home Assistant frontend
Open Home Foundation / Home Assistant (community, non-profit; HQ jurisdiction unverified) · Lit web components, inside the product · measured at src/components
- Context files
- .github/copilot-instructions.md, AGENTS.md
- Skills
- 12 SKILL.md files under a conventional skills directory
- Markdown docs
- 0 doc artifacts in the repository; 103 of 401 components documented, 0 colocated
- Other surfaces
- docs directory: docs
- Tokens
- 684 across 10 authored sources; dark mode unobservable
No published design system: the 401 ha-* components are the product. Tokens are declared as custom properties inside Lit css templates in the theme files, 684 of them, and 59% of the component source's references resolve to one. Legibility is low: 103 of 401 components carry any documentation, and no doc artifact is colocated, though an AGENTS.md, a CLAUDE.md and 12 skills exist. Adherence is not scored because the instrument cannot read style declarations inside Lit templates; the 723 hardcoded colours it did find are reported as a finding without a band. It also found 2431 application source files outside the system, 2407 of them readable, and none import the design system, use one of its components, or carry a style declaration it can read.
Microsoft Fluent (v9)
Microsoft Corporation, US · React library · measured at packages/react-components
- Context files
- AGENTS.md
- Skills
- 13 SKILL.md files under a conventional skills directory
- Markdown docs
- 1,049 doc artifacts in the repository; 261 of 527 components documented, 215 colocated
- Other surfaces
- Storybook (14 configs, in the repository); docs directory: docs
- Tokens
- 661 across 44 authored sources; dark mode generated
527 components found, 261 documented, 13 skills and an AGENTS.md: a large system with uneven coverage. Tokens are numerous (661) and dark mode is generated from the same source, but the component source carries 2,139 raw values against 10 token references, concentrated in the v8-to-v9 migration shims and the styles hooks. Adherence was measured against the in-repo apps and examples: 3,648 hardcoded colours, 230 raw elements where a component exists, and 85 local components that reuse a system name. The trial recipe records that the documented setup is light theme only.
Primer React
GitHub, Inc., US · React library · measured at packages/react
- Context files
- .github/copilot-instructions.md
- Skills
- 15 SKILL.md files under a conventional skills directory
- Markdown docs
- 5 doc artifacts in the repository; 148 of 223 components documented, 172 colocated
- Other surfaces
- Storybook (2 configs, in the repository); MCP: .vscode/mcp.json
- Tokens
- 139 across 30 authored sources; dark mode patched
The library documents 148 of 223 components and ships 15 skills under .github/skills, which is what carries legibility. Single source is the weak column: 3,551 hardcoded colour and size values against 4 token references in the component source, most of them the legacy theme's colour tables, and dark mode patched in a separate file from the light values. Adherence was measured against the in-repo docs and examples, not GitHub.com, and 7,063 of the hardcoded colours found there sit in the legacy theme package. The trial recipe notes the documented setup imports the light theme only.
Astryx
React (StyleX) · measured at packages/core
- Context files
- AGENTS.md, CLAUDE.md
- Skills
- none
- Markdown docs
- 356 doc artifacts in the repository; 270 of 276 components documented, 174 colocated
- Other surfaces
- Storybook (1 config, in the repository); docs directory: docs
- Tokens
- 237 across 11 authored sources; dark mode generated
The system the instrument was built against, scored here under the same conditions as the field and never ranked among it. 270 of 276 components documented with 174 colocated doc artifacts, an AGENTS.md, and a CLI that prints its own documentation; 237 tokens with 75% of the component source's references resolving to one. Adherence measured against its own docs and apps. The prediction stated before any run held on single source and legibility; the trial found the gap under live use, on the informed arm's composition.