AI in practice
AI helps write the code. I protect the rules it follows.
Useful output isn’t the only test of an AI workflow. The rules guiding the tool need to stay visible and open to review.
Where the work went
Implementation got faster. The judgment moved to either side of it.
An agent writes code quickly, so less of my time goes to typing it. More goes to setting up what the agent works within, and to proving what it produced is right.
Before AI
Specs
Working out what to build.
Implementation
Where most of the engineering time went.
Testing
Checking it once it was built.
With AI
Specs
Unchanged: working out what to build.
Context engineering
The rules the agent works within: CLAUDE.md, AGENTS.md, the hooks, and the decision log.
Implementation
Much faster to produce. Still reviewed, and still mine to answer for.
Testing and verification
Automated tests of every page in two browsers, at all 28 color settings. The counts are further down.
The repository behind this site.
This page describes the repository for this site, the code you're reading now. I build it with Claude Code, and I added hooks that treat changes to the agent's own rules differently from everyday edits.
Governance files the hooks protect
- .claude/settings.json
- .claude/hooks/*.sh
- CLAUDE.md
- AGENTS.md
- docs/agent-decision-log.md
How it works
Two paths to the same files
Claude Code can change a file with its Edit and Write tools or with a shell command. Each path gets its own check, and they share one record of what's already been reviewed.
Edit / Write
Direct edits
Step 1: Tool call
Claude Code tries to edit or write a governance file.
Step 2: PreToolUse
A hook pauses the edit and asks me to confirm it.
Step 3: PostToolUse
If I approve, the edit applies and a second hook marks the file as reviewed.
All protected files, including AGENTS.md
Checked before anything changes.
Bash
Shell commands
Step 1: Tool call
Claude Code runs a shell command, such as sed, a redirect, or cp.
Step 2: Command runs
Nothing checks the command first, so any file changes happen.
Step 3: PostToolUse
A hook compares governance file timestamps with the session baseline and review markers.
Step 4: Result
An unreviewed change blocks Claude Code from continuing until I review it.
All protected files except AGENTS.md
next dev rewrites AGENTS.md on its own, so its timestamp moves without any agent action.
Shared session state
A SessionStart hook records a baseline time. The Edit / Write path adds a review marker for each approved file, and the Bash check reads both, so it only flags changes made since the session began that nobody has reviewed.
Why governance files
Rules need a different check than code
A bug in a component shows up in tests, review, or on the page. A change to the agent's rules can quietly remove the check that would catch the next problem. So those changes stop and wait for me.
Why this subset
Only the checks this repo needs
These hooks are trimmed down from a larger set in another project that also covers branding, APIs, auth, quality gates, docs, and accessibility. This small site only needs the governance checks.
The decision log records what I kept, what I dropped, and why, so the setup can be questioned rather than taken on trust.
Scope
A focused safeguard, not a safety net
What it covers
What it doesn’t
Trade-offs
- The Bash check runs after the command. It stops work for review but doesn’t undo the change.
- Shell changes to AGENTS.md aren’t flagged, because next dev rewrites that file itself. Direct edits still ask first.
- On first install there’s no baseline yet, so the first Bash check may flag existing governance files. After that, the baseline is set.
How this site is built
Built the way I'd build it for a team
The same habits as the case study, at a smaller scale: a design system turned into tokens, Server Components by default, and accessibility checked by tests rather than by eye.
Every hue the color picker can make is tested.
The picker in the header can put the site on any color, gray, white, or black, so the contrast suite checks every page at each of them. The suites run in CI on every push, so a change that breaks contrast at any setting fails the build, whether I wrote it or an agent did. Only six components run in the browser; everything else renders on the server.
- Next.js 16 App Router
- React 19
- TypeScript
- Tailwind CSS v4
- Server Components
- Playwright
- axe-core
- Lighthouse CI
- Claude Code
- 303automated Playwright tests across every page, run in Chromium and WebKit
- 28color settings each page is checked at for WCAG 2.2 AA contrast
- 22keyboard tests for navigation, menus, and the color picker
The tools will keep changing. The responsibility won't.