Field notes on running AI coding agents: gates, context economics, verification, and the craft of making agent work repeatable.
Task, constraints, output format, temperature: the four-part prompt contract that makes 3B to 13B local models reliable in agent loops, and the order to tune in when output is bad.
Client names, rates, hours, and invoice history are confidential business data. Why billing records belong in files you own, shared one document at a time.
A one person software shop building AI agent tools, field manuals, and services in the open. The belief system behind it, from the founder who still has the problem.
Agents that wake up briefed, a board that limits work in progress, criteria-gated loops, and offline local models that ship real work.
The Plan → Act → Observe → Reflect pattern that turns a 7B model on your laptop into a worker that finishes jobs — and why local models need loops the most.
What local-first means for your income, client, and invoice records — and the sixty-day math that ends the SaaS-dashboard discussion.
Why benchmark results rot in folders, and the three artifacts a workflow should produce instead: comparisons, trends, verdicts.
Three interactive field manuals for operating AI coding agents: the free doctrine manual, a harness generator you own, and an operator course. Now live.
AUTO, NOTIFY, APPROVE, FORBID — how to decide what your agent may do on its own, and why one table prevents most agent disasters.
Prompt caching is the single biggest cost lever in agent loops. Stable prefixes, just-in-time retrieval, and hard budget stops.
The generator–verifier gap, the test-editing incident every operator eventually meets, and machine-checkable definitions of done.