Systems & AI Engineer
I build and run AI systems // the whole stack, GPU up
scroll ↓
// featured work
Home AI orchestrator
One local door to the whole fleet: per-user RAG, citations through tools, capability-matched scheduling, nine safety invariants.
RAG · agents · scheduling →
Self-engineering loop
The AI proposes changes to its own orchestrator; every change clears a 2,424-test gate and a human promotion.
agents · CI/CD →
32B QLoRA on 24 GB
Fine-tuned a 32B model on a single consumer Radeon under Windows ROCm, no published recipe at this size.
training · ROCm →
AMD GPU wedge, root-caused
A silent compute stall every Windows health check called normal. Traced to MES firmware, cleared in ~30 seconds.
debugging · drivers →
Nursing & law pilots
Two real-user pilots on the orchestrator: hybrid retrieval, clinical-safety checks, defects fixed through use.
RAG · real users →
Cross-scale adapter transfer
A controlled negative result: can a 32B model's adapter cross scales onto its 4B sibling by linear algebra alone?
experiment · eval →
Distributed ARM compute
Eleven low-power nodes, seventy-six cores, for independent-task parallelism and model-scale evaluation.
distributed · fleet →
Operating doctrine
The rulebook for the agents in my lab: twenty-one failure classes, each anchored to a real incident, enforced in tooling.
practices · governance →
The Gauntlet
Over 1,000 model-by-device runs across the ARM cluster, producing verified worker tiers for local routing.
evaluation · benchmarking →
// more work