Reproducible public work
Fresh local tests
A reference implementation you can inspect and run.
I published a local-first reference shell so another person can inspect the action boundaries and see how provider selection behaves.
01 THE PROBLEM
What needed to change
A description of safe agent behaviour is difficult to trust without runnable code, failure checks and explicit limits that another person can inspect.
02 THE IMPLEMENTATION
What I built and directed
I defined an architecture with replaceable model providers, local file memory, explicit human review for consequential actions and append-only audit records. The public shell implements selected mechanisms with synthetic demonstration data. Its implementation and assessment were produced with coding agents under my direction. Review exposed a malformed risk-input boundary. I directed a fail-closed fix and regression checks before publishing it.
03 THE CHECK
What the evidence establishes
The public implementation passed 384 local tests on 9 October 2026, including 224 added numeric-domain and CLI regressions, plus eight verifier checks. Published CI also completed the Python suite, verifier and UI build. The tests use synthetic fixtures and mocked providers.
Scope of this example
A runnable local reference implementation using synthetic inputs and mocked providers, with owner-directed code review and published tests.
04 THE LESSON
The judgement behind the code
Make a useful claim small enough to reproduce. Simulations, inventory counts and model usage remain separate from observed operational value.