AI Platform Engineer · Michigan · remote
Infrastructure evolved with AI. Now run by it.
I run a production private AI cloud on hardware I own — 19 hosts, a six-node Kubernetes cluster, HA secrets and GPU inference — operated day to day by local agents I direct. It survived a total power loss with zero data loss, and a deliberate chaos test before that. One person covers what a team used to, because the operational knowledge lives in a repo the agents read rather than in somebody's head. Twenty years leading operations in automotive manufacturing came first, and it is the half of this that can't be crammed for.
Every company is about to run agents against its infrastructure. Most will find out the expensive way that the hard part isn't the model.
19 hosts, a six-node cluster, 104 pods, HA secrets and GPU inference — operated continuously by one architect and a fleet of agents.
The runbook, the incident records and the decision log are the agents' training surface. When someone leaves, the operating knowledge doesn't.
The DR plan of record is an agent with a recovery skill, not a document. Proven live against a total power loss.
They confidently report work they didn't do. I've published the taxonomy of how they fail and the controls that catch it — including my own errors.
Everyone will have agents within a year. Almost nobody will have an estate an agent can safely operate. That gap is the work.
A 793-line runbook of traps that already cost real time. 50 dated, immutable incident records. 39 proven primitives. Written to be read by an agent at 3am, not by a person who already knows.
~2,900 files of operational contextKAT-Coder on a GPU in the rack, behind a deny-all firewall, routed through my own gateway. It reads the repo to learn how this estate works, then changes it.
zero cloud API callsMCP tools return what is actually running. A document describes what was built; when the two disagree the tool wins, and the stale document is itself a finding.
the rule that prevents driftChanges land as pull requests. Secrets never leave me. The agent proposes and executes; the decision about what is acceptable stays human — and so does catching it when it reports work it did not do.
the part that does not delegateNot a preference — a working condition. Most engineers have used one harness and one model family; the skill that transfers is the one worth hiring.
Seven harnesses, one of them mine, and 100B-class models served from hardware in the room. I'm effectively both model-agnostic and harness-agnostic — I pick what fits the task, swap when something better lands, and built my own when nothing fit. Which means none of this is a bet on a vendor: what transfers between them is the direction and the verification, and that is the only part that was ever scarce.
Nobody scheduled it. A whole private cloud cold-started in dependency order — storage, database quorum, hypervisors, then Kubernetes — with zero data loss and no split-brain. The disaster-recovery plan of record here is a local agent with a recovery skill, deliberately not a human runbook, because the runbook is useless if the person holding it is asleep. This outage is what proved it.
Two steps needed me, and both were things an agent physically cannot do: repairing a lost BIOS boot entry at the console, and unsealing Vault with shamir keys that by design never leave my hands. Everything else — the resume waves, the retry pass, the gateway diagnosis, the image steering — was executed under direction and written up the same evening, defects included. The full record →
Newest first, so this reads from most AI-directed down to hand-soldered. Durability runs the same direction — which is not the one anyone expects.
| Platform | How it was built | Where it is now |
|---|---|---|
| KWSbare metal → K8s, GitOps, SSO, Vault, GPU inference | AI-built under my direction | production · survived full power loss and a deliberate chaos test |
| Hirdforgemulti-agent engineering platform | directed build | deployed to production K8s · 712 merge commits |
| AsgardProxmox, Talos, full rack, 10G | AI-collaborated | completed · runs Hirdforge in production |
| Valhallak3s on mini PCs | by hand | retired |
Valhalla was three mini PCs I built by hand to teach myself Kubernetes. Those same machines now run the KWS control plane. The discipline is the deliverable, not the keystrokes.
All of it on hardware in the room, behind a deny-all firewall. Each choice below is one I made and have had to live with.
Every number verified against live state. Every claim one click from its own evidence.
Published · read this one first
Thirteen claims from three AI systems on the live estate, audited against running state. Eight contained a material error — three of them mine. Not hallucination: seven were a real measurement taken correctly against the wrong surface.
Read the taxonomy →The agent reported a 80% reduction. In consistent units it was 28%, and 28% is what got published. The gap between those two numbers is the entire job — and an estate that can't call out is one whose security work has to run inside it.
Flagship · production
Chaos-tested control-plane loss — hard-killed a host running two of three masters. Failed safe, quorum held, and it exposed an anti-affinity gap I fixed. Airgapped by default: deny-all egress, digest-verified image mirror.
Architecture, decisions, incidents →Product · in validation
Fleets of AI coding agents under human supervision. Routing and completion are deterministic and mechanical, never LLM-judged; changes land as PRs behind human approval. The decision log keeps rejected alternatives and corrections visible.
One doctrine, learned expensively. A status doc, a dashboard, or an AI's "done" is a claim. Only a query against the running system is a fact.
Vault's transit secrets had been sealed for over a day. Nothing alerted. Every dashboard was green.
An agent reported critical CVEs down 210 → 41. Both numbers were real; they measured different things. In consistent units it was 28% — and 28% is what got published.
Including the CNI. The comparison matched mirror image names against upstream names — two spellings of the same image.
Anyone can get an AI to emit code. Getting durable engineering out of it, and knowing which of its claims to distrust, is what the industry is currently failing at. I work across whichever harness fits the task — the skill is the direction and the verification, not the button I press.
Twenty years of operations leadership, bending toward systems over the exact years the platform work matured. Not a career-changer reinventing himself — an operator whose two tracks were always the same discipline.
I was leading teams at a low-volume OEM automotive finisher when GPT-3 landed. I had the certificates — TryHackMe, Google Cybersecurity — but certificates aren't systems, and I needed real ones to learn on. So I built them, broke them, and wrote down what broke. Root-cause discipline and verifying a fix instead of accepting a report is the judgment baseline, not a gap to explain away.
Open to platform, forward-deployed and AI-reliability roles. Remote preferred. US citizen — no sponsorship required.
Résumé · Sovereignty Labs · Michigan, Eastern Time