Architecture: AI and compute
One request, routed to the model that fits it, across three GPUs and an ARM cluster. The full local roster, and how work reaches each tier.
tiered routing capability-matched placement local-first full model roster
Click the map to enlarge.
Tap a tier to see its models.
quick chatBigHonker2›
mid · heavyMothership2›
deep-thinkMothership2›
vision · caption2›
medical · embeddings3›
the FellowshipARM cluster7›
image genComfyUI4›
nomic-embed-text runs under RAG grounding. Flux 2 is parked, impractical on the 7900 XTX until faster kernels.
An ordinary request does not choose a model; the scheduler does. Kirk grounds and pins it, then Spock places it on the tier whose capabilities fit, cheapest-energy hours first for work that can wait, and a local tier before the external API. I can force a specific tier when I want one, but the default path places the work for me. This map is the routing and the roster. How an order actually runs once it lands on a factory, the queues, lane claims, and heartbeats, is the fleet and factory scheduling map. The named components and the safety invariants are the Home AI page, the hosts are the lab overview, and the front door is the network and security map. Training on the 24 GB card is the 32B QLoRA work; the ARM side is the Fellowship.