I have a server system available for running models locally. My first goal is to uncover unknown unknowns. Of course the industry is moving really fast, new techniques and tools are coming out every week it seems.
## learning journey
Not long after I hatched this plan, I spotted [this blog post](https://terminalbytes.com/run-qwen-3-8-27b-locally/), touting the model Qwen3.8 27B.
Then came a [youtube explainer](https://www.youtube.com/watch?v=mmntN7zIekU) sponsored by Framework itself (and hosted by Donato Capitella, who I hadn't heard of). This did a good job of translating what model names mean, what quantization is, how much memory is needed by different parts of a system, etc. Ultimately, it provided recommendations about which models to run on each configuration of the Framework Desktop.
- Donato really likes DeepSeek v4 Flash on the 128GB framework. Especially the version 2026-07-31 (as of the video taping in August, at least)
- he also mentions [Inkling-Small](https://thinkingmachines.ai). in a 2- or 3-bit quant.
- and Qwen3.7-27B
- and Qwen3.8-Flash-Next (2- or 3-bit quant)
- and GLM-5.3-Flash (2-bit quant) (though llama.cpp is still working on supporting this model)
As for what he says on memory requirements:
1. Linux uses some for Operating System tasks.
2. The desktop environment uses some (if the computer is not headless)
1. Any other applications open on the computer will use some (web browsers/sites are notoriously greedy)
3. the model weights (see below). Full-precision or quantized.
4. the context, stored in a KV cache
1. tokens already processed, documents and chats you give the model
2. the size required depends on the model's architecture
3. the inference engine can quantize the cache. Donato recommends an 8-bit precision for this.
Then I saw how one of the youtubers I subscribe to does it. He made [a whole video](https://www.youtube.com/watch?v=ZxjuEHTKXMw) showing his use case (which seems close to mine!).
- model provider: Ollama
- model: Qwen 3.5:9b
- prompt interface: Open WebUI
- AI gateway: LiteLLM
- Voice client (hardware): [Reachy Mini](https://pollen-robotics.com/reachy-mini/)
- speech-to-speech pipeline:
- Voice Activity Detection: ??
- Speech to Text: [Whisper](https://huggingface.co/Systran/faster-distil-whisper-large-v3)
- Model Prompt & Response: Qwen 3.5:9B
- Text to Speech: [Kokoro](https://huggingface.co/onnx-community/Kokoro-82M-v1.0-ONNX)
- (web) search provider: Tavily
- vector store: PostgreSQL pgvector
Somewhere along the way I found [this guide](https://github.com/hogeheer499-commits/strix-halo-guide/blob/main/STRIX_HALO_LOCAL_LLM_SETUP.md), which has promising words like "current measured known-good baseline". It recommends changing BIOS settings. And I'll likely want to install this headless OS on a partition.
I ran my own performance comparison on [local-llm-benchmarks.dev](https://local-llm-benchmarks.dev) and I want to try:
- **Model**: Qwen3.6-35B-A3B
- **Quantization**: UD-Q4_K_XL
- **Engine**: Llama CPP
- **Runtime**: Vulkan RADV Performance
- **Serving**: MTP - 3 draft tokens
- **Performance results**:
- **Prompt Processing**: 1230 tps / 614 tps / 405 tps (0B / 32K / 64K context)
- **Token Generation**: 92 tps / 84 tps / 65 tps
Linkdump:
- [lmstudio.ai](https://lmstudio.ai)
- [lemonade-server.ai](https://lemonade-server.ai)
- [huggingface.co](https://huggingface.co) (Nvidia just bought it)
- github.com/ggml-org/llama-cpp
- [strix-halo-toolboxes.com](https://strix-halo-toolboxes.com) (Donato has many dockerized environments there).
- [local-llm-benchmarks.dev](https://local-llm-benchmarks.dev)
- [thinkingmachines.ai](https://thinkingmachines.ai)
## hardware
I'm running a Framework Desktop:
- CPU: AMD Ryzen™ AI Max+ 395 ("Strix Halo")
- GPU: Radeon™ 8060S
- Memory: 128GB LPDDR5x-8000.
Theoretical memory bandwidth of 256GB/s.
It's in a rack in my server closet, so I have a JetKVM hooked up to it for connectivity.
## build log
- 2026-09-17: installed the mainboard and power supply on the rack, mounted the rack. Shuffled around some hard drives to free up a 2TB Samsung 970 EVO Plus
- 2026-09-21: got the empty hard drive into the machine, flashed a boot disk of Ubuntu Server 26.04.1 LTS, and booted the machine! Set up user and SSH keys.
## glossary
_NB: these definitions are mine. They're probably wrong, incomplete, or biased toward my own use-cases. Also, the markdown isn't rendering properly on the web version of this page. Sorry!_
AI Gateway
: An abstraction layer between the clients (chat windows, microphones, etc) and the models.
: Allows for configuring things like the System Prompt only once, to have all clients and all models use the same config.
Decode (aka Token Generation Decode)
: How fast a model can generate new tokens.
: For a coding agent use-case, target 15-25 tps.
Dense Model
: Uses all of its weights on every token. All weights must be in memory at once.
: See also _Mixture of Experts Model_.
Mixture of Experts Model
: Can select which weights it needs on each token. This reduces how many of its weights need to be in-memory at any given time. The primary effect this has is to increase the speed of response, since less data needs to be moved around in memory.
: Use this style when speed is important.
Prefill (aka Prompt Processing)
: How fast the model can ingest something and turn it into tokens.
: Half of the "performance" metric of a model.
: For a coding agent use-case, target 200-300 tps.
: See also _Decode_.
Quantization (Quant)
: Like compression. It's good for running larger models on memory-constrained systems.
: Usually come in 2-, 4-, or 8-bit formats (un-quantized model weights are usually in a 16-bit format).
: The tradeoff is that the fewer bits in the format, the lower quality results you get, but also the less memory you need to run the model to begin with.
: Many quants are published by Unsloth Dynamic, hence the "UD" prefix in the name of the model on huggingface.
Runtime
: The most common one for local inference is llama.cpp.
System Prompt
: Instructions that get sent into the model with every prompt you make.
: Can be formatted as markdown.
: Can include sections like Identity (who the AI is supposed to be), Voice/Format (what kind of language the AI is to use), Personality (how the AI behaves), Examples (some concrete prompts with their expected responses)
Vector Store
: Like a database for documents. Doesn't store the document, but stores references and locations within documents.
Weight
: When a model is "9 billion" or "27 billion", it's talking about how many weights it has.
: Usually in the BF16 format, which uses 16 bits per weight.
: You can usually do some light math to figure out how much memory a certain model and quantization will need. For example, a 9-billion-weight model in a 4-bit quant needs 9 billion ⨉ 4 bits = 36 billion bits = 4.5 GB.