> For the complete documentation index, see [llms.txt](https://tinyhumans.gitbook.io/openhuman/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://tinyhumans.gitbook.io/openhuman/guides/guides/local-model.md).

# Use OpenHuman with a local model

Run OpenHuman's inference on your own machine with Ollama: detection, model selection, a real test, and every way it commonly breaks with the fix.

**Goal:** move some or all of OpenHuman's model work onto your own computer, so that data used for those workloads never leaves the machine.

Local AI is **opt-in** and ships **off**. Turning it on doesn't silently reroute everything. You choose which workloads go local.

For the config-level reference (every flag and provider field), see [Local AI (optional)](/openhuman/features/model-routing/local-ai.md). This guide is the task-oriented version: get it running and confirm it works.

***

## Prerequisites

* [**Ollama**](https://ollama.com) installed. OpenHuman talks to it at its default address `http://localhost:11434`. (LM Studio is also supported at `http://localhost:1234/v1`; see the [reference page](/openhuman/features/model-routing/local-ai.md#lm-studio-troubleshooting).)
* **8 GB+ RAM** to get real value. Machines with less than 8 GB fall back to cloud summarization by design, because a small local model won't have the headroom.
* Disk for the weights: a small chat model plus an embedding model is a few GB. OpenHuman does not ship weights; Ollama pulls them on demand.

## Privacy implications

* Workloads you route locally (embeddings, summary building, background loops, and, if you choose, chat/reasoning) run **entirely on-device**. Nothing about that work is sent out.
* Anything you leave on the default route still goes through the OpenHuman [model router](/openhuman/features/model-routing.md). Local AI is additive; it doesn't change what you didn't move.
* If the local provider becomes unreachable mid-session, requests **transparently fall back** to the remote provider. That means a crashed Ollama can send that data to the cloud path instead. If strict locality matters, watch the diagnostics (below).

***

## Steps

### 1. Start Ollama

Install and launch Ollama so its local server is running. You can confirm it's up from a terminal:

```bash
curl http://localhost:11434/api/tags
```

A JSON list of models (even an empty one) means the server is reachable. This is the exact probe OpenHuman uses to detect Ollama.

### 2. Turn on Local AI in OpenHuman

Open **Settings → AI & Skills → Local AI**. It's off until you opt in here. Pick the **model tier** that matches your machine's memory (for example the `ram_2_4gb` tier on a typical laptop). OpenHuman selects sensible on-device models for you, for example `gemma3:1b-it-qat` for chat and `bge-m3` for embeddings.

By default the tier keeps **embeddings and memory** on-device while chat and reasoning stay on the cloud route. To move those workloads local too, use **custom routing** to point the chat and reasoning workloads at the local provider, then test that a message stays on the machine.

### 3. Let it pull the models

When a workload needs a model that isn't installed yet, OpenHuman pulls it through Ollama and shows **download progress**. You can also pull manually:

```bash
ollama pull gemma3:1b-it-qat
ollama pull bge-m3
```

### 4. Test that it actually answers

Don't assume; verify. The Local AI settings expose a **test action** that sends a short prompt to the configured local provider and shows you the reply. If you get a coherent response back, the path is live end-to-end (detection → model loaded → inference).

***

## Success checks

Local AI is operational when:

* [ ] **Settings → AI & Skills → Local AI** shows Ollama as reachable and your selected models as available (not "downloading" or "missing").
* [ ] The test action returns a real reply from the local model.
* [ ] The inference status reads **`ready`** (not `degraded`, `downloading`, or `disabled`).
* [ ] For an embeddings-only setup: after the next memory sync, new summaries keep appearing in the **Memory** tab with Ollama running. That confirms embeddings are being produced locally.

{% hint style="info" %}
**Where to look under the hood.** OpenHuman surfaces a live diagnostics view for the local runtime (Ollama reachable? runner OK? which models are installed vs expected? what issues?). If a check fails, the diagnostics name the specific problem. Start there before changing config.
{% endhint %}

## Common failures

These are the actual failure states the runtime reports, and what each one means:

| What you see                                                                        | Meaning                                                                  | Fix                                                                                                                                      |
| ----------------------------------------------------------------------------------- | ------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------- |
| **"Ollama server is not running or not reachable"** (status `degraded`)             | The app can't reach Ollama at its base URL                               | Start Ollama; confirm `curl http://localhost:11434/api/tags` works; if you use a non-default port, set the base URL in Local AI settings |
| **"…reachable but cannot execute models. Restart the external runtime and retry."** | Ollama is answering but its model runner is broken (a fork/exec failure) | Quit and relaunch Ollama itself, then retry                                                                                              |
| **"Chat model '…' is not installed"**                                               | The configured model isn't pulled yet                                    | Let OpenHuman pull it, or run `ollama pull <model>`                                                                                      |
| **"…not reachable after fresh install. Start `ollama serve` manually and retry."**  | Ollama was just installed but the server isn't up                        | Run `ollama serve` (or launch the Ollama app) and retry                                                                                  |
| Embedding model rejected for **context window too small**                           | The chosen embedding model can't hold enough tokens for the memory layer | Choose a larger-context embedding model such as **`bge-m3`**                                                                             |
| Status stuck at **`downloading`**, then a retry message                             | A model pull stream was interrupted                                      | It retries automatically; if it keeps failing, check disk space and network, then pull manually                                          |
| It "works" but answers feel cloud-quality                                           | The local provider was unreachable and OpenHuman **fell back to remote** | Fix reachability above; strict-local users should confirm status is `ready` before relying on it                                         |

## Recovery

* **Back to cloud in one step:** turn Local AI back off in **Settings → AI & Skills → Local AI**. Workloads return to the default route immediately; no data is lost.
* **Free up a stuck runtime:** quit Ollama fully and relaunch it, then re-open the Local AI settings so OpenHuman re-probes.
* **Disk pressure:** interrupted pulls are usually low disk. Clear space and let the pull resume, or `ollama pull` the model by hand.

***

## Notes on what stays cloud anyway

Even with "everything local", some workloads are cloud by default unless you explicitly route them: speech-to-text, text-to-speech, and web search go through the backend proxy. See [what stays in the cloud](/openhuman/features/model-routing/local-ai.md#what-stays-in-the-cloud-by-default).

## See also

* [Local AI (optional)](/openhuman/features/model-routing/local-ai.md): the full config reference.
* [Keep sensitive data private](/openhuman/guides/guides/privacy-sensitive-data.md): the plain-language version of local vs external.
* [Automatic Model Routing](/openhuman/features/model-routing.md): how tasks get matched to models.
