Recall artifacts · skill series · v2 · 2026-08-19

Building with Agents

The five skills the series teaches — re-cut from laws to practice after your altitude ruling — with the first one-pager drafted beside it.

For Robert’s red pen · drafts only · nothing ships to the team without his word

The five: S1 validation · S2 loops · S3 environment · S4 code & intelligence · S5 frame fitness  |  the S1 draft one-pager →

  1. The frame, v2 — skills, not laws

    Your altitude ruling, folded in: v1 mapped the laws our machinery earned — outputs of the practice. The series teaches the practice itself. A law files itself in the reader's head under "Robert's system"; a skill files itself under "me, tomorrow, when I sit down with Claude." Scars are demoted to evidence inside skill-first artifacts.

    Your rulings on the v1 six: validation and loops confirmed. "Demand receipts" demoted — real, not primary; it rides inside the others as a sub-move. "Add context, not rules" re-centered on legibility — the builder's job is the environment around the agent. "Code never decides" re-centered on the relationship — how code and intelligence work together. "Name the frame" re-cut as frame fitness at the start of a job, not mid-job debugging.

    Each artifact is one page: the skill (a verb) → when it fires → a scar as proof → how to run it tomorrow.
  2. S1 · The art of validation — drafted, red-pen it here

    Before you build, and before you accept "done": name what actually needs to be true and for whom — the gestalt, the real-world consumers — then give the agent a way to verify that. The agent is all too happy to report all checks passed; all checks passing is not the thing you care about. Aiming the verification is the art; adding more of it is not.

    ↘ sub-moves and evidence

    Sub-moves: make the agent enumerate the consumers and what each needs to be true; ask "what would make your green a lie?"; real data over self-invented fixtures, every disagreement listed; walk the result as the real user before "done."

    Evidence: 11,207 tests + 8 adversarial reviews green, first real-inbox run misfiled a vendor ticket as a sales lead (HANDOFF-2026-08-18-VALIDATION-GAP.md); the "all quiet" dashboard over a live board (experience-loop-doctrine-2026-06-26.md); three green releases, six defects found by the first real teammate walk.

  3. S2 · Build the loop

    Nothing grades itself — not a test, a spec, or a mind. Every consequential ask gets a closed loop: an independent leg genuinely trying to break the work, findings fed back, repeat until it can't. One-shot prompting is the amateur tell; the loop is where quality comes from, and the verifying leg is never the author.

    ↘ sub-moves and evidence

    Sub-moves: the adversary must be able to fail the work — a check that cannot fail proves nothing; independence binds who authored, not which model; make the worker produce evidence the verifier can check (the receipts move lives here).

    Evidence: builders rigged their own red proofs twice in one day — only review from outside the authorship chain caught it (constitution, amended 2026-08-16); four spec revisions each "fixed" a hole that the next independent re-grade found merely relocated (eve-llm-code-boundary-2026-06-17.md). Your Breeze conversation on loops belongs in this artifact when we draft it.

  4. S3 · Build the environment — legibility

    Your job is the environment around the agent. When it doesn't do the job well, the first question is not "what rule do I add" — it's "what context would it have needed to do this job well?" Rules constrain what intelligence may decide and pile up into a lobotomy; context and legibility compound.

    ↘ sub-moves and evidence

    Sub-moves: distinguish rules that constrain judgment from structures that prove a decision happened — the second kind is the trust floor; when a guard is proposed, weigh the failure asymmetry out loud: what does it prevent vs. what does it eat?

    Evidence: deepening the model's inputs took a golden run from 4/13 correct to 13/13 while adding rules to the same problem stalled (REPRESENTS-NOT-GOVERNS-GUARDIAN.md); the safety gate that silently ate a real warm introduction (judge-first-email-mirror-intent-2026-08-18.md); the "honesty" scanner that read the word Tuesday as a fabricated name (receipts-audit-spec-2026-06-16.md).

  5. S4 · Code and intelligence — the relationship

    Know how they work together: code holds facts, creates legibility, and enforces the few hard floors; intelligence reads, judges meaning, and decides. The recurring disease is code quietly doing intelligence's job — a matcher, a threshold, a filter deciding what something means. The recurring cure: code assembles the evidence, the model judges it, code applies the decision through a guarded door.

    ↘ sub-moves and evidence

    Sub-moves: on any design, ask "is this code trying to understand something?"; a verifier reads the writer's own receipt rather than re-deriving meaning from output; the few legitimate code floors — never claim an unperformed action, never corrupt the store.

    Evidence: the meaning-matcher refuted twice, standing verdict "kill, do not patch" (neuron STATE-OF-PLAY.md); the same shape re-earned in warp's courts within the month (warp docs/courts-0675/); ~nine independent mintings in two weeks across two codebases (~/.claude/archive/2026-07-30-code-never-decides/arc-evidence.md).

  6. S5 · Frame fitness — at the start of the job

    Before building, ask: what type of system am I actually building? How does that type of system get validated? What already exists at the frontier — the field's primitives and infrastructure — so you don't build it all yourself? Excellence inside a wrong frame is anti-evidence; the frame question comes first because no amount of in-frame rigor can see out of it.

    ↘ sub-moves and evidence

    Sub-moves: name the problem class abstractly and cite the field's default and strongest-alternative approach before writing code; the mid-job tripwire — a second failed fix on the same spot indicts the frame, not the fix, never a third cleverer patch; before any plan, ask in writing whether it's the surface of a deeper one.

    Evidence: ~17 patch cycles that never converged, dissolved by naming the problem class and adopting the field's standard technique (quanta CLAUDE.md, Frontier-first, 2026-06-21); three consecutive plans in one session each turned out to be the surface of a deeper one (constitution step-back law, 2026-07-03).

  7. Parked — real, not in this series

    Domain insights about memory systems — "memory is testimony, never verdict" (the three-shards incident), "store what produced the memory, not a summary" (the 93.5%→1.8% regression), staleness labeling, identity-as-judgment. Product-design gifts, not building skills; possibly a later "memory field notes" one-off if you want it.

    The v1 law map — fifteen law-first candidates with full receipts, preserved in this file's git history if any needs promoting.

  8. The pipeline

    One artifact a week. Every artifact one page: the skill → the scar → how to run it tomorrow. New laws ratified in our work enter the series as evidence under a skill, not as artifacts of their own.

    Team-facing words stay yours: every artifact ships through your red pen. Drafts only from me.
  9. Your calls

    1 · Red-pen the S1 draft

    It's written as team-facing words in your voice — the draft one-pager. Annotate it there or dictate changes; nothing goes to Alie until you say so.

    2 · Does S1 open the series?

    My lean stands — it's the one you named unprompted. Say the word to reorder.

    3 · Anything missing from the five?

    The next dig can aim at your practice in the transcripts — the moves that never got written down as doctrine — if you feel a sixth hiding.