Kristopher L. Porath
AI Platform Engineer
Remote (US) · Michigan, Eastern Time · US citizen, no sponsorship required
kitporath@gmail.com ·
kitporath.com ·
github.com/kitporath ·
linkedin.com/in/kitporath
Summary
AI Platform Engineer who builds and operates production private AI
infrastructure — and extracts dependable engineering from unreliable models. Architected a
production private AI cloud (Kubernetes, GitOps, HA Vault, SSO, internal PKI, self-hosted
GPU inference) as a safe substrate for a self-built multi-agent platform, and recovered it
from a total power loss with zero data loss. The estate is operated day to day by
local agents under my direction. Two decades of operations leadership behind the rigour:
up to 35 direct reports, program management, and IT systems administration. Published the
failure taxonomy for agent-directed infrastructure work, including my own errors.
Target roles: AI Platform Engineer · AI Infrastructure / LLMOps · Platform
Engineer (AI/ML) · Forward-Deployed Engineer
Selected results
- Recovered a full private cloud from total power loss with zero data loss and no
split-brain — storage, database quorum, hypervisors and Kubernetes cold-started in
dependency order. Executed by an agent under my direction; the only steps requiring me
were a BIOS repair at the console and a shamir unseal. Converted into a verified runbook
and permanent fixes the same evening.
- Chaos-tested control-plane loss by hard-killing a host running two of three
masters. Cluster failed safe, quorum held, and it exposed an anti-affinity gap I then
closed.
- Ran a supply-chain campaign with a local model on owned GPUs — 210 → 151 critical
CVEs, measured rather than reported. Caught three of the agent's own reporting errors,
including 52 images retired as "unused" of which 15 were running the cluster.
- Upgraded RKE2 and Cilium (1.18 → 1.19.4) across six nodes, eliminating 24 critical
CVEs with zero workload outage.
- Published an error analysis of agent-directed infrastructure work
— 13 audited claims, 8 material errors, and the countermeasures now enforced in CI.
Core skills
- Platform & orchestration: Kubernetes (RKE2), ArgoCD app-of-apps GitOps,
OpenNebula/KVM, OneKE, Proxmox, Gitea Actions CI/CD, per-PR environments
- Infrastructure as code: Terraform, Ansible, declarative catalogs exported from
live state, drift detection wired to alerts
- Security & identity: Keycloak OIDC/SSO, HashiCorp Vault (3-node HA, transit
auto-unseal), cert-manager and internal PKI, Cilium network policy, Falco runtime
detection, airgap and egress control, digest-verified image mirroring
- AI / LLM systems: self-hosted serving (llama.cpp), OpenAI-compatible gateways
with routing and cost attribution (LiteLLM), GPU and model fleet lifecycle, MCP tooling,
multi-agent orchestration with mechanical verification
- Observability: Prometheus, Grafana, Loki/Promtail, blackbox and textfile
exporters, alert rules verified firable before shipping
- Storage & networking: LINSTOR/DRBD, MinIO (S3), ZFS, 3-2-1 backups,
DNS-as-code
- Languages: Go (AI-directed), Python, Bash, YAML/HCL
AI harnesses & models
Effectively model-agnostic and harness-agnostic — the skill that transfers
is the direction and the verification, not familiarity with one vendor's tooling.
- Harnesses: Claude Code, Codex, opencode (cloud and local), Grok, Antigravity,
Hermes, and Hirdforge — which I built.
- Frontier models: Claude (Opus 5, 4.8–4.6, Sonnet 5, Fable 5), GPT 5.1–5.5,
Gemini, Grok 4.5, Kimi 2.5/2.7, MiniMax 2.5–3, Nemotron Ultra 3.
- Served on my own GPUs: Qwen 3.5/3.6 (27B, 35B-A3B, 122B), Gemma (12B, 31B,
26B-A4B), Nemotron 3 Plus, Step 2.7 and DeepSeek V4 Flash on mixed CPU/GPU, and Laguna S 2.1 at 75GB RAM-resident.
Experience
Sovereignty Labs — Founder & AI Platform Architect
2025 – Present · Remote
- Architected and operate KWS, a production private AI cloud: 19 monitored hosts,
6-node RKE2 Kubernetes (~104 pods), 8 auto-synced ArgoCD applications, 3-node HA Vault
with transit auto-unseal, Keycloak SSO, internal PKI, and a digest-verified image mirror
for airgapped operation behind deny-all egress.
- Directed AI implementation of Hirdforge, a multi-agent engineering platform:
~67,000 lines of Go across 225 files and 712 merge commits, deployed to production
Kubernetes. Built deterministic routing and mechanical CI gates so agent output is
verified by tests rather than model self-assessment.
- Built and operate a self-hosted GPU inference layer serving OpenAI-compatible APIs
across a multi-GPU fleet, with a gateway providing per-role routing, quotas and cost
attribution.
- Made the estate operable by agents: the runbook, 50 dated incident records and 39
proven primitives are the corpus agents read to learn how the system works, with live
state always retrieved by tool call rather than from documentation.
Prefix Corporation — Manufacturing Operations & IT
2010 – Present · Michigan
- Automation Logistics (2024–Present): material logistics for an automated
industrial paint line; management-track role.
- IT System Administrator (2023): administered company IT systems — the point where
self-taught infrastructure work became the day job.
- Program Manager (2020–2021): cross-functional program and delivery ownership.
- Team Lead → Supervisor (2013–2022): led up to 35 employees; ran production
operations for quality, throughput and on-time delivery against fixed schedules.
Saleen Inc. — Team Lead
to 2008 · Michigan
High-performance automotive finishing.
Certifications & learning
- Google Cybersecurity Certificate; TryHackMe.
- Self-directed and hands-on, evidenced by four generations of production systems —
Valhalla → Asgard → KWS → Hirdforge — rather than coursework.