---
title: "01 — Architecture"
description: "## The two-tier pattern"
section: ai-docs
raw: "01-architecture.md"
source: ai-generated
tags: ai, llama.cpp, docker
last-updated: 2026-09-15
---

# 01 — Architecture

> The overall shape of a self-hosted local AI stack: two tiers, GPU isolation, and how the pieces find each other.

## The two-tier pattern

A local AI machine splits cleanly into two tiers, and the split is worth keeping no matter how the details change:

- **Inference tier** — one or more `llama.cpp` servers running **natively on the host** (plain systemd services), each pinned to a specific GPU via `CUDA_VISIBLE_DEVICES`. They expose an OpenAI-compatible HTTP API.
- **Application tier** — Docker containers (chat UI, terminal, plugin orchestration) on a shared user-defined Docker network.

Why native for inference rather than Docker? A GPU workload wants the full device with no container indirection, and the model files live on the host filesystem. Keeping inference out of containers means the web tooling can be rebuilt, upgraded and restarted independently — a broken Open WebUI image never takes the models down with it.

The local machine follows this exactly:

| Tier | Components | Notes |
| --- | --- | --- |
| Inference (native) | `llama-primary` on the RTX 3090 (:8082), `llama-secondary` on the RTX 3070 (:8083) | One server per GPU, one role per server — see [Primary vs Secondary](/ai-docs/03-primary-vs-secondary/) |
| Application (Docker) | Open WebUI (:8081), Open Terminal (:8080), Open WebUI Pipelines (:9099) | All on one shared network — see [Open Terminal & Tooling](/ai-docs/05-open-terminal-and-tooling/) |

## One server per GPU

`CUDA_VISIBLE_DEVICES` is what gives you hardware isolation between servers: each `llama-server` only sees its own card, so a context-heavy request on the big card can never evict or starve the small one, and vice versa. This is the whole reason to run two servers on one box rather than one server on two GPUs — it is as good as two machines, except you can share models between them when convenient.

Each server is an ordinary, individually startable systemd unit (`Restart=on-failure` / `Restart=always`, `RestartSec=5`, logs appended to `/var/log/llama_*.log`). Ordinary units matter if the machine is dual-use: anything that needs the GPUs (a desktop, a render job) can simply `systemctl stop` the server, and the default state is whichever one you start at boot. This box additionally auto-switches between headless inference and an interactive desktop (or gaming machine) when a USB monitor switcher is plugged in or out — see [Dual-Use: Inference ⇄ Gaming/Desktop](/ai-docs/07-dual-use-gaming/).

## How the tiers talk to each other

Docker containers cannot see host services by name, so the standard trick is used: every container declares

```yaml
extra_hosts:
  - "host.docker.internal:host-gateway"
```

which maps the hostname `host.docker.internal` to the Docker host. Containers then reach the native servers with `http://host.docker.internal:8082/v1` and `:8083/v1`. Two practical notes:

- The shared network itself (`ai-shared-net`) exists mainly so the application containers can address *each other*; it is not how they reach the inference tier.
- The inference base URLs are **not** in any config file here — they are entered through the Open WebUI web UI and stored in its data volume. That means they survive image upgrades but not a wiped data volume, so keep the data volume backed up.

## Where data lives

A pattern that works well: application data is persisted to host paths under a single root (here `/var/local/docker-files/<service>/`), while model files live in one shared directory on the host (`/usr/local/llama/models/` for the `.gguf` files) with a preset file per server (`/etc/llama/models-*.ini`) deciding which models each server loads. The deployment was originally built around Ollama's per-server blob stores (`.ollama/` for the primary, `.ollama-secondary/` for the secondary) but has since migrated to a pure llama.cpp install — one directory is cleaner, and "which GPU does this model belong to" stays answerable from the preset it appears in. Databases that don't need to scale (Postgres for Open WebUI's data) can run on the host and be reached over a mounted Unix socket instead of being containerised — simpler lifecycle, lower latency, one fewer container.

## Typical request flow

```text
User ──HTTP──► Open WebUI (:8081)
                 │  base URL configured in the web UI:
                 │  http://host.docker.internal:8082/v1  (or :8083)
                 ▼
             llama-server (native, on the host)
                 │  runs inference on its own GPU
                 ▼
             token stream ──► Open WebUI ──► User
```

Tool-assisted chat adds a loop: the application tier calls a tool (e.g. Open Terminal's API, an MCP server), the tool does something, and its output goes back into the model's context for the next turn. Sub-agents are just a special case of that loop where the "tool" is another model endpoint — see [Sub-agents](/ai-docs/04-sub-agents/).

## Sizing the machine

The useful mental model when designing your own setup is **VRAM budget per card**, not raw model count:

```text
VRAM_needed ≈ weights(quant) + KV_cache(ctx × layers × quant) + ~0.5 GB overhead
```

Everything else — which models to run, how much context, whether a model is resident or swapped — falls out of that equation. The local setup works out the numbers for an 8 GB and a 24 GB card; see [Context & Quantisation](/ai-docs/02-context-and-quantisation/).

## Related

- [Context & Quantisation](/ai-docs/02-context-and-quantisation/)
- [Primary vs Secondary](/ai-docs/03-primary-vs-secondary/)
- [Open Terminal & Tooling](/ai-docs/05-open-terminal-and-tooling/)
- [Dual-Use: Inference ⇄ Gaming/Desktop](/ai-docs/07-dual-use-gaming/)

---

## Source Disclaimer

- [x] AI Generated
- [ ] Human Generated
- [ ] AI Edited
- [ ] Human Edited
