All posts
Company UpdatesSeptember 21, 2026·6 min read

Sharing Proof, Not Promises

By Dan McAulay

Sixth post in the series on how Shipwright actually came together — the first one covered the cloud review pipeline, the second covered OpenClaw and the original todo list, the third covered the merge into Vitals OS, the fourth covered the task store, the fifth covered the dispatcher that actually runs the pipeline today. That post ended with a promise: pulling Shipwright out of the Vitals OS monorepo and putting it in front of anyone who wants to run it themselves. This one’s that story — and the reason we actually did it wasn’t a roadmap milestone, it was a question we didn’t have a good answer to.

A client asked us, directly, on a call: how do we know Shipwright will still be maintained six months from now? Fair question. We didn’t have a good verbal answer for it — “trust us” isn’t an answer, it’s a request. We’d already been circling open-sourcing the harness before that call. No VC backing means no roadmap promises we can point to as collateral; an open-source product is one a developer can evaluate on its own evidence instead of taking our word for it. But circling an idea and doing something about it are different things, and that call is what moved it from the first column to the second.

Dave Didn’t Wait

The client’s question forced an answer, but it didn’t set the timing — Dave and I had already agreed the extraction needed to happen eventually. I just didn’t think it was time yet. Shipwright still had real room to get better, and pulling it into its own repo felt like it would slow that down, not speed it up. Dave disagreed, and instead of arguing it out, he just started — the first commit landed under his name, no more debate needed. Once it existed, waiting stopped being something I wanted either: running two codebases side by side, patching the old one while the new one caught up, was worse than just finishing. So I went heads-down and stayed there, and I wasn’t going to let token spend be the thing that stopped me from getting it done and back to actually making Shipwright better.

The Numbers

app-vitals/shipwright didn’t exist until June 6th. Warchild scaffolded the rest of that day — bun workspaces, CI, the metrics service ported over — not a gradual drift out of the monorepo, an actual extraction with a start date. The next two days went into building the harness itself — the entire agent runtime ported wholesale: crypto and config, Slack and cron handling, GitHub auth, the plugin registry, the entrypoint and Dockerfile — alongside a new admin API with session-cookie auth, and by June 7th a public marketing site scaffolded in the same repo. Every agent was still running on the old vitals-os runtime — none of it was live yet.

Migration tooling followed June 11th and 12th, verified read-only against all 30 of Bodhi’s live crons before anyone let it run for real. The first agent migrated June 12th to prove the path, canary-first; the rest followed over the next several days, each dropping the old runtime’s secrets once confirmed. June 17th, the last agent came online. June 18th, one commit deleted the entire legacy agent/ workspace — 17,165 lines across 107 files, the old Helm chart, the interceptor service, the database tables, the role — gone in a single shot.

Thirteen days, start to finish. Across both repos: 429 commits, roughly 400 merged PRs, one repository inflating from zero to 124,954 lines added while the other shed 105,487 — one side ballooning, the other hollowing out, at the same time. June 17th, the day before the cutover, was the spike: 46 commits in the old repo, 39 in the new one, in one day.

We queue tasks up, and any agent in the fleet with access to shipwright can pick one up and start executing — dev-task, review, patch, and deploy all draw from the same queue, the same pipeline dogfeeding its own extraction the way it works every other queue in the company. What Dave and I actually did for those thirteen days was plan, queue, unblock, and verify — the pipeline did the throughput. That’s where the token spend went: I blew well past my usage plan running this, landed around two thousand dollars over to Anthropic for the stretch, mostly in the June 16th–18th crunch.

The Four-Failure Night

Not every part of the rollout went this cleanly. Two days before the finish line, migrating shipwright-deploy itself to GKE — the service that deploys everything else — produced four separate failures in one night. Encryption keys reset on every upgrade, killing every live agent’s auth token. A missing health check caused 503s on the admin service. OAuth redirects silently pointed at localhost in production. A stuck deploy cascaded into blocking everything behind it.

All four hit the same night, and all four hit for the same underlying reason: the tool being migrated was the one doing the migrating. Every fix had to go out through the exact deploy path that was currently broken. That’s the real risk in an extraction like this — not that something breaks, but that the thing you’d normally reach for to fix it is the thing that’s down. All four got written up afterward as a named failure catalog in the migration runbook, not quietly patched and forgotten.

The Part You Can Now Go Read Yourself

That night used to be the kind of thing that stays inside the company — a Slack thread, a line in an internal runbook, gone from memory within a quarter. Once the repo went public, that stopped being true. It’s sitting in git history now, for anyone deciding whether to trust this tool to go read directly, not take from us secondhand.

Last chapter’s “Blast Radius, Not Luck” section covered why the ship-and-patch week that ran the same month never touched a live client — Keanu was on a completely separate path from the pipeline that kept breaking, and there was no mechanism yet that pushed code onto a client’s deployment automatically. Same fact applies here, and it’s doing different work this time. It’s not defending what happened — it’s the reason we could afford to let it become visible at all. The chaos was real, but it was contained to our own fleet the whole time, so opening the door on it costs us nothing with the people running Shipwright and buys something with everyone deciding whether to.

That’s the actual answer to the question this post opened on. Not a promise that nothing will break — a public record of what already did, and what we did about it.

Why This Shape

Every post in this series comes back to the same test: does this piece earn its place by pointing at a specific problem it solved for a specific person. Open-sourcing Shipwright is the one place that test points outward instead of in. We didn’t answer “will you maintain this” with a better sentence. We gave away the repo, mess included, and let anyone check the work themselves. That’s a stranger, harder-to-fake kind of trust than a roadmap slide — and it’s the same instinct this whole series runs on, just aimed at readers instead of at code.

Next up: the metrics journey — PostHog first, then a backend-agnostic pivot the moment this thing had to work for people without our accounts, and why we ended up querying the task store directly instead of a separate analytics platform at all.

Want to accelerate your engineering team?

Book a 30-minute discovery call to discuss your team's AI adoption strategy.

Get in Touch