Architecture: AI and compute

One request, routed to the model that fits it, across three GPUs and an ARM cluster. The full local roster, and how work reaches each tier.

tiered routing capability-matched placement local-first full model roster

Click the map to enlarge.

Tap a tier to see its models.

request
Kirk grounds, cites, pins
Spock capability-matched placement
quick chatBigHonker2
qwen3.5:9bqwen3:8b
mid · heavyMothership2
qwen3:14bqwen3.6:27b
deep-thinkMothership2
qwen3.6:35b-a3bqwen3-coder:30b
vision · caption2
qwen3-vl:8bJoyCaption
medical · embeddings3
medgemma:27bmedgemma:4bnomic-embed-text
the FellowshipARM cluster7
Qwen2.5-1.5BCoder-1.5BQwen2.5-0.5BCoder-0.5BQwen3.5-2BSmolVLM2-500MQwen2-VL-2B
image genComfyUI4
Flux.1 devFlux KontextRealVisXL V5SDXL

nomic-embed-text runs under RAG grounding. Flux 2 is parked, impractical on the 7900 XTX until faster kernels.

An ordinary request does not choose a model; the scheduler does. Kirk grounds and pins it, then Spock places it on the tier whose capabilities fit, cheapest-energy hours first for work that can wait, and a local tier before the external API. I can force a specific tier when I want one, but the default path places the work for me. This map is the routing and the roster. How an order actually runs once it lands on a factory, the queues, lane claims, and heartbeats, is the fleet and factory scheduling map. The named components and the safety invariants are the Home AI page, the hosts are the lab overview, and the front door is the network and security map. Training on the 24 GB card is the 32B QLoRA work; the ARM side is the Fellowship.