The Agentic Framework
The tools, the rules, and the loop I use to ship apps with AI instead of chatting with it.
I make the calls. The AI does the typing.
- Claude Code
- VS Code
- Codex
Set up once, then loop, then ship
The whole framework in three lines.
- 1
Set up once
A rules file (CLAUDE.md), my notes and the three tools, wired before the first task.
- 2
Build in small loops
Frame, build, check, fold back. One small piece per run, repeated until the criteria pass.
- 3
Ship on a yes
Nothing reaches a repo, a client or a send button until I approve it.
Who does what
The split that keeps the loop honest.
Me
- Decide what done means
- Approve every write
- Take over when precision beats speed
- Turn mistakes into rules
The AI
- Read the files that matter
- Write the code and the tests
- Run it and grade itself against the criteria
- Report back with the diff
When the AI starts deciding scope, or I start typing boilerplate, the loop is broken.
Three tools, one job each
Each one has a job. I do not use three where one will do.
- 01
Claude Code
The builder. Long tasks, real files, subagents for anything that can run in parallel.
Opus for planning, Sonnet for volume - 02
VS Code
The editor. Where I read what was written and take over when precision beats speed.
Human in the loop - 03
Codex
The second opinion. A fresh model reviewing a diff catches what the author missed.
Review pass, not a build pass
Four rules I do not break
Non-negotiable, and they are what keep the output usable.
- R1
Write the spec first
If I cannot describe done in a paragraph, the agent will not find it in a thousand tokens.
Spec before prompt - R2
Every write is approval-gated
Agents draft. Nothing reaches a client account, a repo, or a send button unreviewed.
No autonomous writes - R3
Small units of work
One task, one context. A fresh session beats a long one every time.
Reset early, reset often - R4
The model grades itself
Acceptance criteria go in with the task, and the output is checked against them before I look.
Self-review built in
Frame, build, check, fold back
Four steps, repeated until the criteria pass. Click any step or pin.
- A task (One small piece), then Frame
- Frame (Claude Code · Opus), rule R1, then Build
- Build (Claude Code · Sonnet), rule R3, then Check
- Check (Codex · VS Code), rule R4, then Approval if passes; Fold back if fails
- Approval (Me), rule R2, then Shipped if approved; Build if send back
- Shipped (Live on the URL)
- Fold back (Into CLAUDE.md), then Frame if becomes a rule
One real feature, from reference to live
The End-to-End Workflow page, taken through the loop.
- 1Frame
Studied a reference page for its shape only, then answered the design questions one at a time.
- 2Frame
Wrote the spec and ticked every line of it before any code. Then a six-task plan, with the code for each task.
- 3Build
Built each task test-first: write the test, watch it fail, make it pass.
- 4Check
511 unit tests and 194 browser tests green. Then a fresh reviewer model read the whole branch and found four bugs.
- 5Fold back
Each bug got a test that failed first, then the fix. Two lessons became permanent tests: labels fit on one line, and the drag test waits for the page to settle.
- 6Ship
I checked it, said “push, PR, merge”, the checks passed, and it went live.
What I do every time
Small, boring, and the reason the loop works.
- 1
Fresh session per task
A long chat drifts. I close it and start clean with the spec.
- 2
Read every diff
I read what was written before it goes anywhere. No exceptions for small changes.
- 3
No new asks mid-run
A new idea waits for the next loop, so the current one can finish.
- 4
Check it on the real URL
Tests passing is not the end. I open the live page and use it.
Done is a URL, not a diff.
Every term on this page, in one line
Plain English for the technical words above.
- Claude Code
- Anthropic’s AI coding tool. It works in the terminal, on real files.
- VS Code
- Visual Studio Code, the code editor where I read what was written and edit by hand.
- Codex
- OpenAI’s coding agent, used here to review work it did not write.
- Opus
- Claude’s most capable model, used for planning.
- Sonnet
- Claude’s faster model, used for most of the building.
- Spec
- A short written description of what done looks like, agreed before any code.
- Context
- Everything the AI can see while it works. Less is better.
- Acceptance criteria
- The checks the work must pass to count as done.
- Subagent
- A helper AI the main one hands a side task to, so parts run at once.
- Diff
- The exact lines a change adds and removes.
- Second opinion
- A different model reading the work, to catch what the author missed.
- Approval gate
- A point where the work stops until a person says yes.
- CLAUDE.md
- The rules file Claude Code reads at the start of every session.
- Loop-back
- A cable that sends the work back to an earlier step.
- Fold back
- Turning a mistake into a rule or a test, so it does not happen again.
- Fresh session
- A new, empty AI chat, started clean for the next task.
- Test-first
- Writing the test before the code, and watching it fail first.
- Reviewer model
- A fresh AI model asked only to review finished work.