One local AI lab.
The flagship version uses a WRX90 platform with enough PCIe lanes, memory channels, cooling, and power headroom to make multi-GPU work possible without pretending that model parallelism is free.
A dream local-AI workstation running Omarchy Linux: enough VRAM for serious open-weight models, enough CPU and memory for research, and enough receipts to know when the expensive rectangle is actually helping.
The machine is designed around local model capacity, multi-GPU experimentation, fine-tuning, and concurrent agents. Omarchy supplies the human-facing layer; the hardware supplies the actual power.
The flagship version uses a WRX90 platform with enough PCIe lanes, memory channels, cooling, and power headroom to make multi-GPU work possible without pretending that model parallelism is free.
Omarchy is not a magic performance potion. It is a clean, opinionated interface to a machine whose critical boundaries remain visible.
Encrypted Btrfs and Limine for the workstation base. NVIDIA open kernel modules and pinned CUDA/PyTorch environments for fewer update-induced surprises.
GGUF inference, quantization experiments, GPU offload, CPU fallback, and the simplest path from a model file to a local OpenAI-compatible endpoint.
Continuous batching and tensor parallelism for serving; Transformers, PEFT, and TRL for fine-tuning. Containers keep experiments reproducible.
Encrypted local storage. No private model files sent away by default.
Measure weights, context, KV cache, quantization, and actual memory use.
Use local models for research, coding, agents, vision, and bounded experiments.
Record throughput, latency, energy, quality, and the cases where the monster was slower than the cloud.
These are engineering targets, not promises. Quantization, context length, architecture, interconnect topology, and framework support decide whether a model merely fits or is pleasant to use.
7B–32B models with large context windows, multiple resident services, vision tools, embeddings, and local agent loops.
70B-class models on one 96 GB GPU. 100B–120B-class models are plausible with suitable low-bit formats and a disciplined context budget.
200B–400B-class models through multi-GPU sharding. Capacity may work before latency does. The market has enough benchmark theatre already.
Contributions support the research and construction of the local AI workstation described here. The target configuration is a design, not a purchased asset. Hardware, power, taxes, availability, and prices can change. We will publish what was actually bought, what it cost, and what it can actually do.
Stripe checkout is an existing BIPU support link. It is not a hardware pre-order, equity offering, token sale, or promise of financial return. Stripe collects payment details; we do not put credentials or payment logic in this static page.
Then measure whether the second GPU earns its slot. The thesis means nothing without the artifact. Ship the pipeline. Measure the outcome. Publish the miss.