> For the complete documentation index, see [llms.txt](https://tinyhumans.gitbook.io/openhuman/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://tinyhumans.gitbook.io/openhuman/developing/architecture/agent-harness.md).

# Agent Harness

How an agent turn runs: the tool-call loop, sub-agent dispatch, archetypes, triage, hooks, and the cost and budget machinery around them.

`crates/openhuman-core/src/agent/` is the host side of an agent turn: the tool-call loop and its dialects, sub-agent dispatch and archetypes, sessions and resume, triage, hooks, interrupts and cost accounting, all wrapped around the TinyAgents harness rather than reimplementing it.

## Embedding OpenHuman as a library

`openhuman_embed` exposes a two-step API. `Runtime::builder()` boots one in-process core per process, background services, registered domain families, backend URL and the TinyHumans API key, and `Runtime::agent(spec)` instantiates any number of agents on it. Each `AgentSpec` fully describes one agent: provider endpoint and model, access tier, `action_dir`, MCP servers, its own in-process tools, skill bundles, system prompt, tool scope, sandbox mode, allowlists, and a narrowed `DomainSet` / `ToolGroups`. A runtime identifies as `HostKind::Library`: inference does not depend on OpenHuman app login, including inference-readiness checks for workflow agent nodes.

### An embedder's own tools

`AgentSpec::tools` takes the host's own `Box<dyn Tool>` objects, which reach the model as real tools, their own schema on the wire, called by their own name. Before it existed the only road was `AgentSpec::mcp`, and the model paid for the indirection: a discovery call to learn what a server offers, and an `mcp_call_tool` envelope whose inner `arguments` object no provider can validate or constrain decoding against.

It takes a factory, not a belt. `Agent` is `Clone` and `Box<dyn Tool>` is not, and the session behind a spec is rebuilt from `Config` on every turn, so nothing holding a `dyn Tool` could survive in between. The closure therefore runs once per turn, which also means a host whose tools belong to something shorter-lived than the agent (one episode, one room, one assignment) can return a different belt each time instead of registering a second agent.

Implement `Tool` through `openhuman_embed::Tool`, not by depending on `tinytools` directly: a second path to that crate produces incompatible Rust types, and a tool built against it cannot be handed to a session at all.

One caveat on a varying belt. The prompt's tool catalogue is rendered from the same belt in the same build, so the two stay consistent on any turn that composes a prompt, but a resumed session reuses its persisted system messages, so a belt that moves under a long-lived thread is described by the prompt that thread opened with. Vary a belt only on turns that run on a session of their own.

Per-agent isolation is a context, not a second core. `Runtime::agent` clones the runtime's base `Config`, applies the spec, and derives a child `CoreContext` (`CoreContext::derive_with`) carrying that config, the agent's domain set, tool groups and skill-root policy. Every turn is dispatched under that context (`CoreRuntime::run_in` → `agent_chat_for` with an explicit `AgentDefinition`), so the config loader, the domain gate, the tool-group filter and skill discovery all read the agent's own settings. Transcripts are keyed by agent id and a turn resumes only its own thread.

Layout under a runtime-owned root: `<root>/config.toml` and the credential store; `<root>/workspace/` with the session database, `session_raw/` transcripts and each agent's `agents/<id>/skills/`; and `<root>/agents/<id>/action/` as each agent's default working root (a sibling of the workspace, never inside it).

`Harness` is the one-agent shorthand: a runtime plus one agent named `harness`. Build one runtime and issue concurrent `run` or `turn(...).send()` calls on its agents; do not build one core per agent. Each call owns a distinct session unless a prior session id is supplied.

Authentication in library mode is the TinyHumans API key (`RuntimeBuilder::api_key`), stored as an `api-key` auth profile beside the runtime's `config.toml`. Managed inference sends it as a bearer to the OpenAI-compatible endpoint; backend REST calls send `x-api-key`; there is no session JWT and nothing to expire. Agents that name their own `Provider` (BYOK) never touch it. The key covers every hosted feature, subject to its backend scopes: inference, embeddings, voice, search, media, Composio and the other integrations, channels and the socket. Only the session-bound `/auth/*` flows (OAuth connect, channel link tokens) still take `HarnessBuilder::session`.

`Workspace::Inherit` together with `Provider::inherit()` and no API key is deliberately not library-routed inference. It borrows the installed OpenHuman configuration and therefore keeps the installed application's session checks. Supply an explicit provider or an API key when embedding without app login.

Some settings remain runtime-wide for every agent (the live autonomy policy's `auto_approve*`, the approval gate switch, the sub-agent catalogue, the config sub-agents re-read); the crate README lists them under "Still runtime-wide".

Every agent turn runs through the published [`tinyagents`](https://crates.io/crates/tinyagents) harness. There are three entry points: `Agent::turn`, the channel and CLI bus path, and `run_subagent`. All of them go through the adapter seam in [`crates/openhuman-core/src/agent/tinyagents/`](https://github.com/tinyhumansai/openhuman/tree/main/crates/openhuman-core/src/agent/tinyagents/README.md) (`run_turn_via_tinyagents_shared`). Context trimming is `MessageTrimMiddleware`, and cancellation is the tinyagents steering channel. Policy stop hooks (budget, thread-goal and iteration caps) fire through a `StopHookMiddleware` ([`tinyagents/stop_hooks.rs`](https://github.com/tinyhumansai/openhuman/blob/main/crates/openhuman-core/src/agent/tinyagents/stop_hooks.rs)) that pauses the run on the first stop vote. The channel route forwards live `AgentProgress` like the chat route. The shared `TurnProgress` seam lives in `agent/session_host/tool_progress.rs`.

Multi-agent orchestration runs on the tinyagents graph layer, through `graph::parallel::map_reduce`, the `spawn_parallel_graph` scaffold and the shared `graph::orchestration` `TaskStore` lifecycle primitives re-exported from [`orchestration/mod.rs`](https://github.com/tinyhumansai/openhuman/blob/main/crates/openhuman-core/src/agent/orchestration/mod.rs):

* [`orchestration/delegation.rs`](https://github.com/tinyhumansai/openhuman/blob/main/crates/openhuman-core/src/agent/orchestration/delegation.rs) is a `plan → execute ⇄ review → finalize` `CompiledGraph`. It has conditional routing, `RecursionPolicy`, a durable `FileCheckpointer`, a `CancellationToken` and a `GraphTracingSink`.
* The workflow phase engine fans each phase's agents out on the graph (`with_max_concurrency`). The durable `WorkflowRun` ledger stays the source of truth for resume.
* `spawn_parallel_agents` runs its fan-out through `spawn_parallel_graph` and `graph::parallel::map_reduce`.
* The agent-teams member runtime is a conditional-routing graph (`execute → complete | fail → done`, [`agent_teams/runtime.rs`](https://github.com/tinyhumansai/openhuman/blob/main/crates/openhuman-core/src/agent/orchestration/agent_teams/runtime.rs)).
* The detached-sub-agent registry is backed by a typed `TaskStore` lifecycle ledger (Pending, Running, then Completed, Failed or Cancelled).

## TinyAgents crate: features and compatibility

OpenHuman depends on the split TinyAgents 2.1 crate family by path into the vendored git submodule `vendor/tinyagents`, `tinyagents-harness` (features `sqlite`, `multimodal`, `builtin-tools`), `tinyagents-definition`, `tinyagents-graph` (`sqlite`), `tinyagents-registry`, `tinyagents-session`, `tinyagents-runtime`, `tinyagents-orchestration` and the optional `tinyagents-live`, plus `tinyinference` from `vendor/tinyagents/vendor/tinyinference`, so SDK changes can be tested in-tree before they go upstream (see [`crates/openhuman-core/Cargo.toml`](https://github.com/tinyhumansai/openhuman/blob/main/crates/openhuman-core/Cargo.toml)). The reasons, so future upgrades do not regress it:

* Native TinyAgents model interface, OpenHuman-owned product policy. Every live route is an `Arc<dyn ChatModel<()>>`: TinyAgents OpenAI-compatible clients cover wire-equivalent managed, local, and BYOK routes, while host `ChatModel` implementations cover Claude SDK/Code and Codex-specific transports. OpenHuman still owns credential resolution, OAuth, access gates, endpoint selection, egress disclosure, billing metadata, and error classification.
* `sqlite` feature enabled with one native sqlite chain. OpenHuman's root and Tauri Cargo worlds pin `rusqlite = "=0.40.2"` (`bundled`); the pinned Rust 1.96 toolchain provides the `cfg_select!` macro its build script needs. Both worlds resolve to a single `libsqlite3-sys` chain. Durable graph checkpoints run through TinyAgents' own `SqliteCheckpointer` (see [`orchestration/delegation.rs`](https://github.com/tinyhumansai/openhuman/blob/main/crates/openhuman-core/src/agent/orchestration/delegation.rs)).
* WhatsApp Web storage bridge. `whatsapp-rust`'s Diesel-backed `sqlite-storage` feature links sqlite separately from rusqlite 0.40, so the optional `whatsapp-web` feature (forwarded to `tinychannels/whatsapp-web`) builds against `wacore::store::InMemoryBackend` and logs that sessions are not durable. Web sessions stay non-persistent until a rusqlite-backed durable WhatsApp store exists.
* No scripting runtime. There is no `repl` or Rhai scripting feature and no `rhai_workflows` tool. Rhai still appears as a dependency of `tinyflows` for its own JSON-query stack, unrelated to the agent harness.
* Ownership map: model construction → `inference::provider::create_chat_model*`; durable graph checkpoints → TinyAgents' `SqliteCheckpointer`; generic detached executor state → `DetachedTaskRegistry`; controller-facing durability → OpenHuman SQL/JSON run ledgers (`running_subagents`, `workflow_runs`, `agent_teams`, `command_center`). The generic harness/graph/middleware/event primitives are used as-is.

The agent harness is the runtime that turns a user message (or a webhook fire, or a cron tick) into a complete, tool-using LLM interaction. It owns the tool-call loop, sub-agent dispatch, the trigger-triage pipeline, and the hook surface around them. It does not own provider HTTP transport, tool implementations, prompt-section assembly or memory storage. Those are separate domains the harness composes.

This page walks through what happens in one turn, then zooms in on each of the moving parts.

## The shape of a turn

Every turn goes through the same lifecycle, whether the user just typed a message, a Telegram webhook fired or a 9am cron ticked:

```
┌─ inbound ─────────────────────────────────────────────────────────┐
│ user message · channel inbound · webhook · cron · composio event │
└──────────────────────────┬────────────────────────────────────────┘
                           │
                           ▼  (external triggers only)
                ┌──────────────────────┐
                │   trigger triage     │  classify → drop / notify /
                │   (small local LLM)  │  spawn reactor / spawn orchestrator
                └──────────┬───────────┘
                           │
                           ▼
            ┌──────────────────────────────┐
            │      Agent::turn()           │
            │  1. resume transcript        │
            │  2. build system prompt*     │
            │  3. recall memory pack      │
            │  4. enter tool-call loop ────┼──► provider call
            │  5. dispatch tool calls  ────┼──► tool exec / sub-agent spawn
            │  6. context guard / compact  │
            │  7. stop-hook check          │
            │  8. final assistant text     │
            └──────────┬───────────────────┘
                       │ async, after the user sees the reply
                       ▼
              ┌─────────────────┐
              │  post-turn      │  cost log · conversation
              │  hooks          │  buffering for memory
              └─────────────────┘

* the system prompt is built only on the first turn. Later
  turns reuse the rendered prompt verbatim so the inference
  backend's KV-cache prefix stays valid.
```

The rest of this page is the same diagram, expanded.

## Sessions and `Agent::turn`

A session is the live conversation an `Agent` instance is running. The `Agent` struct owns:

* The conversation history (system + user + assistant + tool messages).
* The provider client to call (model resolved by the [model router](/openhuman/features/model-routing.md)).
* The tool registry visible to the model.
* The `memory` tool (when a memory engine is on), through which the model recalls, fetches, learns and forgets on demand.
* Per-turn budgets: max tool iterations, max payload size, max USD cost.
* Local action budget: a rolling hourly cap for side-effecting tool actions, read from `config.autonomy.max_actions_per_hour`.

`Agent::turn(user_message)` is the hot path. In one turn it:

1. Resumes the session transcript if this is a fresh process, re-loading the exact provider messages from disk so the inference backend's KV-cache prefix still hits.
2. Builds the system prompt (only on the first turn). This pulls in identity, soul, profile, memory, connected integrations, available tools, safety preamble, assembled by the prompt section builder.
3. Recalls a memory pack. TinyMemory's agent lifecycle logs the user turn and recalls a token-budgeted pack (learnings and beliefs, brain documents, this agent's history, other agents' turns), wrapped in `<memory-context>` and added to that turn's model requests ephemerally; it is never persisted in the transcript, so the prompt cache is unaffected. The model can still call the `memory` tool for a cited answer (see [Memory architecture](/openhuman/developing/architecture/memory.md)).
4. Enters the tool-call loop (next section).
5. Spawns post-turn hooks in the background: the user gets their answer before cost logging and conversation logging finish.

The system prompt is not rebuilt on subsequent turns. Even cosmetic byte changes invalidate the KV-cache prefix and force a full re-prefill, so dynamic per-turn context (tool results, memory recall answers) is appended as user-visible message content rather than spliced into the system prompt.

#### The cacheable prefix is wider than the system prompt

Freezing the prompt is only one third of the contract, and the other two are easier to break because nothing about them looks like caching:

* The tool block counts, and it comes *first*. Every prefix cache in production renders the tool catalogue ahead of the conversation, a chat template has to put it somewhere the model reads before the first user turn. OpenAI's automatic cache, Anthropic's `cache_control` (tools → system → messages), DeepSeek's context cache and any vLLM/SGLang radix cache all work this way. So a `tools` array that changes between turns invalidates the frozen system prompt too, and the JSON key order of the request body (which puts `messages` before `tools`) says nothing about it. The mid-session Composio reconcile is still correct, a tool surface that lies about what the model can call is worse than a cold prefill, but it must fire only on a real capability change and be byte-stable otherwise. `connected_set_hash` sorts before hashing so a reordered backend response never reaches a rebuild, and `collect_orchestrator_tools` sorts the connected-toolkit enum for the same reason.
* History must be append-only. Turn N's serialization has to survive verbatim as the opening of turn N+1. `pair_tool_cycles` drops half-finished tool cycles at serialization time, so its verdict for an entry must depend only on that entry and its immediate neighbour, never on anything appended later, or an earlier message's presence flips retroactively and the prefix moves under the cache. Context compaction is the one deliberate exception; it rewrites the middle and pays for a re-prefill.

A resumed session's replayed prefix is folded into `Agent::history` rather than spliced into one request, so the request, the following turn and the persisted transcript all read the same sequence. Splicing it cost the conversation twice: the next turn went out without it, and the transcript written afterwards (serialized from `history`) held only the new turn, which the *next* resume then read back, truncating the thread a little further on every restart.

#### A stable prefix only pays out when the wire says so

Keeping the bytes stable is necessary, not sufficient: the provider still has to be *told* to cache, and two providers need telling explicitly.

* Anthropic caches nothing without `cache_control` markers, and its OpenAI-compatible endpoint cannot carry them, Anthropic documents prompt caching as unsupported on that path and reports `prompt_tokens_details` as always empty. A `cloud_providers` entry with `auth_style = "anthropic"` is therefore built as the crate's native Messages adapter (`tinyinference_llm::providers::anthropic`), which places markers on the last tool, the last system block, and the final message so a growing tool loop reuses the previous iteration's cache rather than only the system prompt. The one exception is text mode (`native_tools = false`, prompt-guided tools), which only the Chat Completions adapter implements and which keeps the compat client, and so keeps paying full price.
* OpenRouter forwards markers to Anthropic and Gemini but adds none itself; hosted OpenAI rejects unknown content-part fields. The OpenRouter slug (and any endpoint on `openrouter.ai`) enables `OpenAiModel::with_explicit_cache_control`, which marks the last system and last user message; every other Chat Completions endpoint stays unmarked.
* The routing hint is separate from the markers. The harness derives a `prompt_cache_key` from the declared stable prefix and puts it in `provider_options`, so every turn of a thread, and every sub-agent sharing its system prompt and tool set, routes to the same cache shard on providers that shard (OpenAI's `prompt_cache_key`). Adapters that have no such concept drop it. The managed backend receives it as a top-level body field alongside `thread_id`.

All three hang off `RunPolicy::cache.protect_prompt_prefix`, which the host sets unconditionally, and off the cacheable prefix the tinyagents loop declares on every request from the session's frozen system prefix (system prompt tiers, then the tool catalogue); the host has no middleware of its own for this. Until tinyagents began stamping that effective policy onto the outgoing request, both readers consulted `request.cache_policy` alone, always `None` here, so the flag was diagnostic only and no marker or key ever reached the wire. The `[cache]` debug line on every model call now reports whether a key was injected and how many cacheable segments were declared; `cache_read_tokens` on the usage that comes back is the number that proves it worked.

Measure this rather than reasoning about it. `CAPTURE_ALL=1 node scripts/debug/capture-first-inference.mjs` records a whole session's requests, `scripts/debug/run-multi-turn-capture.mjs` drives a multi-turn thread through the production RPC, and `scripts/debug/audit-inference-prefix.mjs` reports the first divergence and attributes it. A single-turn capture cannot see any of these, the question is never what turn 1 costs, it is whether turn 2 can reuse it.

### AGENTS.md project instructions

Alongside identity/soul/profile/memory, the system prompt pulls in AGENTS.md instruction files, OpenHuman's analog of Claude Code's `CLAUDE.md` / Codex's `AGENTS.md`. Two layers are loaded once, at system-prompt build time (never re-read per turn, so the frozen-prefix / KV-cache contract holds):

* Global, `<workspace_dir>/AGENTS.md`, the user's OpenHuman workspace (where `SOUL.md` / `USER.md` live). Applies to every run.
* Project, `<action_dir>/AGENTS.md`, the folder the agent is operating in. For sub-agent runs with a git-worktree override (`SubagentRunOptions.worktree_action_dir`), that override dir is the project layer instead.

The global layer renders first, the project layer second (project instructions layered after, taking precedence on conflict), under a `## Project instructions (AGENTS.md)` heading. When the two dirs resolve to the same path the file is loaded once (deduped). Missing / unreadable / empty files are silently skipped, and each layer is capped at `BOOTSTRAP_MAX_CHARS` (\~20 000 chars) with a `[... truncated]` marker so a large file can't crowd the prompt.

The loader is `agent::prompts::agents_md` (pure functions returning pre-loaded strings, bounded at read time so a pathological multi-MB file can't exhaust memory before the render-time cap); the strings are threaded onto `PromptContext` (`agents_md_global` / `agents_md_local`) and rendered by `AgentsInstructionsSection`, which sits after the user-files section and before the tool catalogue in the default and sub-agent builders. The primary / orchestrator agent and the other built-in dynamic agents (`PromptSource::Dynamic`) assemble their own body via `render_*` helpers, so the same `AgentsInstructionsSection` is injected centrally by `SystemPromptBuilder::from_dynamic`, appended after the agent's own body, before the central grounding contract, rather than by each `agents/<id>/prompt.rs` builder. The feature is gated by `agent.agents_md_enabled` (default on); when off, no AGENTS.md content is loaded or injected.

## The tool-call loop

Inside `Agent::turn`, the tool-call loop is the inner engine. It is the published tinyagents crate's `AgentHarness` loop, assembled per turn by `run_turn_via_tinyagents_shared` ([`crates/openhuman-core/src/agent/tinyagents/mod.rs`](https://github.com/tinyhumansai/openhuman/blob/main/crates/openhuman-core/src/agent/tinyagents/mod.rs)). It runs up to the turn's iteration cap, resolved in `session_host/builder/iteration_cap.rs`: an explicit `[agent] max_tool_iterations_override` (or `OPENHUMAN_AGENT_MAX_TOOL_ITERATIONS`) wins, then the agent definition's cap (the orchestrator's is 200), then `max_tool_iterations` (default 10) for a turn with no definition. The model is told how many calls remain at 50% and 80% of the budget, and `inference_agent_chat` reports a capped turn with `hit_cap: true` and its `checkpoint` text:

```
loop {
    1. context guard      - if history is too big, microcompact / autocompact
    2. stop-hook check    - budget caps, max-iterations, custom kill switches
    3. provider call      - send messages + tool specs, stream the response
    4. parse response     - split assistant text from tool calls
    5. if no tool calls   - return final text
    6. execute tool calls - dispatch each one (next section)
    7. summarize oversize - route huge tool outputs through the summarizer agent
    8. append results     - push tool results into history, loop again
}
```

Every iteration emits a real-time `AgentProgress` event so the UI can render token-by-token streaming, "calling tool X" status, and per-iteration cost updates.

One engine, three entry points. The loop lives in TinyAgents. OpenHuman's chat host, channel/CLI route and sub-agent host each supply product policy at the adapter seam; spawned sub-agents enter through `agent/subagent_host`, which drives the neutral `tinyagents_orchestration::subagent::SubagentDriver` rather than an OpenHuman runner. The host owns provider/model choice, tool and security narrowing, workspace/artifact handling, progress and durable product projection. The neutral driver owns lifecycle ordering, task-key coalescing and the exclusive pause-or-terminal persistence decision. There is no `harness/subagent_runner` path, and it must not return as a facade or compatibility export.

### Tool dispatch and tool-call dialects

`agent.tool_dispatcher` (overridable for one launch with `OPENHUMAN_TOOL_DISPATCHER`) picks how tools are spoken to the model. `auto` (the default) uses native tool calling, structured tool specs through the `ChatModel` adapter and structured calls back, whenever the provider profile supports it, and falls back to JSON-in-tag for prompt-guided providers such as local Ollama. `python` and `typescript` are opt-in: they render the catalogue as code signatures and read code-style calls back, but mis-parse on some models. The session composes its prompt for the chosen dialect and pins the same dialect on the turn harness, so a text dialect keeps its schemas off the wire and the harness recovers calls with the matching grammar.

Canonical `tinytools_agent::dialect::ToolDialect` implementations provide transcript-compatible parsing and rendering directly; OpenHuman converts durable/provider records only at those I/O boundaries:

* Native (`native`), structured tool-call fields.
* XML (`xml`), `<tool_call>{...}</tool_call>` tags in assistant text, with full JSON schemas in the prompt.
* P-Format (`pformat`), compact positional `<tool_call>name[0|a|1|b]</tool_call>` with `name[0|<a>|1|<b>]` signatures in the prompt; opt-in.
* Code (`python` / `typescript`), the catalogue is a list of function signatures (`def read_file(path: str, limit: int = None) -> str` or `function read_file(path: string, limit?: number): string;`) and the model writes a function call inside the tag: `read_file(path="src/main.rs", limit=20)` or `read_file({path: "src/main.rs", limit: 20})`. Compact like P-Format but a syntax small code-trained models already write; `python` is the default.

Every text dialect shares one parser: a `<tool_call>` body is tried as P-Format, then as a code call, then as JSON, so a model that mixes forms is still understood. Persisted session histories can contain suffixes in any of these shapes, so the session shell keeps the dispatcher around to parse and replay them faithfully when a transcript is resumed.

### Context management mid-loop

Long tool-calling chains can blow past the context window. Two layers handle that:

* Tool-result budget: every tool result is checked against a per-call byte budget, enforced as tinyagents tool middleware. Anything over is hard-truncated with an explanatory marker so the model knows it didn't see the full output.
* Microcompact / autocompact: when total history is creeping toward the context window, tinyagents middleware (message trimming + the compression hooks installed in `tinyagents/middleware.rs` and `tinyagents/harness_context_ladder.rs`) compacts older turns into summaries before the next provider call. The compacted history keeps the system prompt and the most recent turns intact (KV-cache stability) and rewrites the middle.

### Oversized tool results: the summarizer detour

Some tool calls return enormous payloads: a Composio action dumping 200 KB of JSON, a web scrape returning 50 KB of markdown, a `file_read` over a multi-thousand-line log. Hard-truncating mid-payload drops whatever happens to land past the cut.

When a tool result exceeds the summarizer's threshold (`summarizer_payload_threshold_tokens`, default 4 000), TinyJuice's summary stage summarizes it before it enters the parent's history. The parent agent sees only the summary. There is one summarizer, and TinyJuice owns it: it decides when a summary is worth writing, writes the extraction prompt (identifiers and key facts first), caches identical summaries, trips a per-thread breaker after three failures, and offloads the original to CCR so the `juice_retrieve` footer can recover it exactly. The model call itself belongs to the host. `ToolOutputMiddleware` binds a unary child of the current turn through `PayloadSummarizer::prepare`, registers the capped model on a one-off harness (`tinyagents/payload_summarizer.rs`), and the module calls back through `MlHost.Generate`. Only the orchestrator carries a summary model, and it always sees `juice_retrieve`, even with the compaction router off, because a summary's footer names it. A result that qualified for a summary and did not get one is labelled as unsummarized, whether the module or the host failed, so the model does not re-run the tool for one. Summary reuse and the breaker are scoped to the thread, or to the run when there is no thread. Hard truncation remains the downstream backstop when summarization fails, or when the payload is so large that an LLM call on it makes no economic sense.

Caller focus. A tool can opt in to an optional `summary_focus` argument by adding `tokenjuice::focus::summary_focus_property()` to its schema; `web_fetch` and `web_search_tool` do. The model uses it to say what it needs from the result, for example "the rate limits". `before_tool` removes the argument from the call before validation, so the tool never sees it. Only a tool whose schema carries that exact property loses it; any other tool with a parameter of the same name keeps it. It reaches TinyJuice with the result, where it steers the summary, keys its cache, and ranks text for the deterministic compressors. A tool that caps its own output, such as `web_fetch`, is normally left to cap-and-spill. When the caller gives a focus, it is summarized as well, because paging the raw page cannot answer a question.

### Filesystem offload: `outputs/` and `workspace/`

Compression alone does not survive a long-horizon run. Summaries still accumulate step after step, and no amount of compressing restores the fidelity that was thrown away. So for minutes-to-hours tasks the harness moves large results out of context and onto disk, and hands the next step a path.

Two directories under the agent's existing `action_dir` at runtime (the implementation lives in `crates/openhuman-core/src/agent/harness/artifact_offload/`):

| Directory               | Holds                                                                      |
| ----------------------- | -------------------------------------------------------------------------- |
| `action_dir/outputs/`   | Deliverables. Meant to outlive the step that made them, handed on by path. |
| `action_dir/workspace/` | Scratch. Intermediate files a worker needs but does not hand back.         |

Note `action_dir/workspace/` is a scratch folder inside the agent's action root. It is not the core's internal `workspace_dir`, which agent writes may never reach.

Two halves enforce the convention:

* Prompt. A sub-agent that actually holds `file_write` gets a Long-horizon Artifact Offload contract in its system prompt: results past roughly 2 000 tokens go to a file under `outputs/`, and the reply is that relative path plus a short abstract. The gate is deliberate, a prompt may only name tools the agent can really call, or the model emits calls that fail. Skill-filtered specialists get no contract text, and a dedicated guard asserts their prompts never mention a filesystem tool. They stay covered by the harness half below, which needs no cooperation from the model. The `planner` prompt spells out what the convention means for its own work; the planner is told to reference artifact paths across DAG nodes rather than pasting payloads forward.
* Harness. `offload_oversized_result` runs on every sub-agent outcome, so an oversized result is offloaded even when the worker inlined it anyway. It fires before the definition's `max_result_chars` cap, so the full body lands on disk instead of being cut.

What the parent receives is a pointer, not a payload:

```
[artifact] kind=output path=outputs/presentation_agent/sub-1234-result.md bytes=52318
read_with: file_read {"path":"outputs/presentation_agent/sub-1234-result.md"}
note: The full result was written to the action workspace instead of being inlined. …

[abstract]
HEADLINE FINDING …
```

`SubagentRunOutcome.artifact_paths` carries the same paths structurally, parsed out of the `[artifact]` pointers in the output, so the handoff carries paths whether the harness offloaded the result or the worker wrote the file itself. Any path the parent takes delivery of is recoverable with an ordinary `file_read` long after the child's context is gone.

The summarizer is the fallback, not the first resort. Offload catches the common case; the summarizer detour and the `tool_result_budget_bytes` truncation above still handle everything it does not, including every failure mode here. A refused or failed offload is deliberately soft: the caller keeps its inline payload and falls through to those backstops.

Path hardening is fail-closed. `resolve_artifact_path` rejects absolute paths, `..` traversal, and anything that escapes its convention root after lexical normalization. When a `SecurityPolicy` is available it also refuses anything under `workspace_dir`, both by blanket containment and via `is_workspace_internal_path`. Offload targets resolve under `action_dir`, never `workspace_dir`, including the case where someone configures `action_dir` inside the workspace root, where every offload is refused rather than quietly writing to internal state.

Writes log `[artifact] wrote worker artifact under action_dir`; each path a handoff carries logs `[artifact] handoff carried an artifact path to the parent`, on both the producing and the consuming side, so a run journal shows both ends of every pointer.

### TokenJuice: content-aware tool-output compaction

Before a fresh tool result enters history (and ahead of the byte-budget backstop), it passes through the TokenJuice content router in the vendored TinyJuice crate (`vendor/tinyjuice`), with OpenHuman adapters in `crates/openhuman-core/src/inference/tokenjuice/`. Inspired by Headroom, the router *detects the content kind* (JSON, code, log, search, diff, HTML, plain text) from the bytes and/or a hint derived from the tool name and arguments, then dispatches to a specialised compressor:

* JSON → SmartCrusher: array-of-objects → table (each key once), preserving rows that carry errors or numeric outliers.
* Code → tree-sitter (Rust/TS/JS/Python) signature keeper that collapses function bodies; brace-depth heuristic fallback.
* Log → the 100-rule engine for *command* output (git/cargo/npm/…), signal-based keep-failures otherwise.
* Search → relevance-ranked top-K matches per file with a `+N more` tally.
* Diff → keep changed hunks, collapse unchanged context, summarise lockfile hunks.
* HTML → strip markup to readable text.
* Plain text → the opt-in Python/ML "Kompress" compressor (ModernBERT), or pass-through.

Every lossy compression offloads the original to the CCR (Compress-Cache-Retrieve) store behind a `⟦tj:<hash>⟧` marker, so compaction is effectively lossless: the agent calls `juice_retrieve` (token + optional byte/line range) to fetch the full original on demand. The same engine is exposed as a universal `compress_content(content, hint, opts)` for any large payload (file reads, web fetches), and as read-only `tokenjuice.*` debug RPCs. Compaction is on by default (`context.compaction_enabled`; `OPENHUMAN_COMPACTION=0` opts out). A large result is stored and replaced by a stats line, a 500-character head and a handle that the read-only `juice_find` / `juice_extract` / `juice_summarize` tools query (`tokenjuice.repl_handle_enabled`), so the LLM summary stage is skipped for results that get a handle. Configured via the `[tokenjuice]` block / `OPENHUMAN_TOKENJUICE_*` env. Agent definitions can override tool-result compression with `tokenjuice_compression = "auto" | "full" | "light" | "off"`; `auto` resolves coding-model agents (`[model] hint = "coding"`) to `light`, which disables CCR-backed lossy compression so coding agents keep raw build/test/diff/search text unless a reduction is truly lossless. Other agents default to `full`. The ML (Kompress) path runs as a `kompress` backend of the shared [`runtime_python_server`](https://github.com/tinyhumansai/openhuman/tree/main/crates/openhuman-core/src/runtime/python_server/README.md) (torch + ModernBERT pip-installed at runtime), gated by the `ml_compression_enabled` flag and degrading gracefully to a native compressor when the Python runtime is unavailable.

### Missing capabilities

There is no self-healing archetype. A missing host command is reported back to the user rather than polyfilled.

## Sub-agents: the orchestrator pattern

OpenHuman is multi-agent. The agent the user is chatting with is the Master Agent (stable internal id: `orchestrator`), a capable default agent that answers and completes ordinary work directly, including the inspect → edit → verify coding loop. It spawns specialist sub-agents when parallelism, deeper reasoning, or a specialised capability materially helps.

### Why multi-agent

A single agent that knows everything also has a system prompt the size of a small book. Splitting work across specialists means:

* Each sub-agent gets a narrow system prompt with only the sections it needs (identity / memory / safety preamble can be stripped).
* Each sub-agent gets a filtered tool registry: the image agent doesn't need filesystem tools, the task manager doesn't need the Composio catalog.
* Sub-agent histories never leak back to the parent: the parent sees one compact tool result, not the inner conversation.
* Cheaper models can do the leaf work. The orchestrator is on a strong reasoning model; a research sub-agent might be on a faster, cheaper one.

### The built-in archetypes

Each archetype lives under `agents/<name>/` with an `agent.toml` (metadata, tool scope, model hint) and a prompt. The orchestrator's chat delegates are the ones in its `[subagents]` allowlist:

| Archetype            | When the orchestrator picks it                                                                |
| -------------------- | --------------------------------------------------------------------------------------------- |
| `orchestrator`       | The Master Agent: top-level, direct-capable default. Never spawned by another orchestrator.   |
| `task_manager_agent` | Adding, previewing, fetching, updating or removing task sources, workflow bundles, artifacts. |
| `vision_agent`       | Anything that depends on the content of an image (describe, OCR, UI elements).                |
| `image_agent`        | Generating or editing an image.                                                               |
| `video_agent`        | Generating a video or animating an image.                                                     |
| `presentation_agent` | Building a slide deck with grounding and citations (feature `documents`).                     |
| `skill_setup`        | Finding, installing and managing skills from registries (`setup_skills`).                     |
| `workflow_builder`   | Authoring or editing a saved workflow (`build_workflow`); returns a proposal, never saves.    |
| `flow_discovery`     | Suggesting workflows worth setting up (`discover_workflows`).                                 |

Other built-ins are never chat delegates:

| Archetype                            | Who runs it                                                                                                                   |
| ------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------- |
| `planner`, `critic`                  | Only the `parallel_research_cross_check` workflow-run template (decompose and research = `planner`, cross\_check = `critic`). |
| `summarizer`                         | The harness, to compress oversized tool results; also the synthesize phase of that workflow-run template.                     |
| `trigger_triage` / `trigger_reactor` | Classifying an incoming external event, and the lightweight reaction to one.                                                  |
| `morning_briefing`                   | Cron: the daily digest, on a named read-only tool belt.                                                                       |

The orchestrator itself handles coding, crypto, settings, scheduling, product docs and MCP servers through inline skills: `use_skill` loads a tool pack's guide and schemas (`coding`, `web3`, `system`, `scheduling`, `docs`, `mcp`), and the same tools are `Deferred` so `tool_search` finds any single one. Running an installed skill is the orchestrator's own `run_workflow`. See [`tools/toolpacks/README.md`](https://github.com/tinyhumansai/openhuman/blob/main/crates/openhuman-core/src/tools/toolpacks/README.md).

Custom archetypes ship as TOML files under `$OPENHUMAN_WORKSPACE/agents/*.toml` (or `~/.openhuman/agents/*.toml` for user-global specialists). Custom definitions override built-ins on id collision.

### Running a reusable sub-agent

When the orchestrator calls `spawn_subagent`, the default contract is durable and asynchronous. The tool builds a deterministic compatibility selector from the parent session/thread, agent id, model override, sandbox mode, action root, and normalized task key/title. It then checks `orchestration::subagent_sessions` before spawning:

* If a compatible worker is already running, the instruction is injected through its `RunQueue` and the parent gets a quick `subagent_session_id` / `task_id` reference.
* If a compatible worker is idle or paused with reusable history, the harness starts a new transient run for the same durable `subagent_session_id` and passes the saved child history through `SubagentRunOptions.initial_history`, with the new instruction appended as a user-visible follow-up.
* If the shape is incompatible, the worker was closed, `fresh: true` was passed, or no session exists, the harness creates a new durable session and worker thread.

The child run is prepared and executed by the OpenHuman host through the neutral TinyAgents lifecycle:

1. Receives parent cancellation and workspace from its typed run context; remaining legacy provider, sandbox, and transcript inputs stay scoped until their explicit migration lands.
2. Resolves the sub-agent's model: inline `model` override first, then config-level pins (`[orchestrator].model`, `[teams.*].lead_model`, `[teams.*].agent_model`), then the archetype hint or inherited parent model.
3. Filters the parent's tool registry per the definition's `tools`, `disallowed_tools`, and `skill_filter`. In `fork` mode, the parent's full registry is inherited verbatim.
4. Builds a narrow system prompt, omitting the sections the definition asks to strip.
5. Passes the prepared child to `tinyagents-orchestration` for lifecycle ordering and exactly one pause-or-terminal persistence action.
6. Persists the child history and worker thread pointer under the durable `subagent_session_id` so later turns can resume or inspect it. Pause continuation also retains its complete TinyAgents task key.

`wait_subagent` and `steer_subagent` accept either the durable `subagent_session_id` or the transient `task_id`; durable ids are preferred across turns. `list_subagents` shows reusable children for the current parent thread, and `close_subagent` marks a worker non-reusable and cancels it if it is still running. Inline blocking is explicit via `blocking: true`.

The synthesized archetype delegations (`delegate_*`, `build_workflow`, and the other `delegate_name` tools) follow the same contract: they route through the durable async path by default, returning an `[async_subagent_ref]` (with `subagent_session_id` + `task_id`) immediately, and the finished result is inserted into the parent chat as a new system turn via `background_completions`/`background_delivery`. The delivery turn persists its own closing message, `sender: "agent"`, id `agent:<run_id>`, `extraMetadata.requestId = run_id`, before it emits `chat_done` as `client_id: "system"`; the frontend reuses that id, so its usual `chat_done` append collapses onto the same row (the conversation store is idempotent for these deterministic `agent:`-prefixed ids; every other id is UUID-fresh and keeps the constant-time append path) instead of persisting the delivered result a second time. Interactive turns follow the same contract: `web_chat::presentation::deliver_response` stores the reply under `agent:<request_id>` before publishing `chat_done`, so an answer the core produced exists on disk whether or not a client is there to receive the announcement, a dropped socket or a reloaded webview costs a repaint, not the reply. The exception is a segmented delivery, where the client owns one row per segment and the core stores none; the frontend keeps generated ids there for exactly that reason. They fall back to inline blocking automatically when there is no parent agent turn or no chat thread to deliver into (cron/CLI), or when `blocking: true` is passed. Cross-turn continuity comes from three pieces: the per-turn `[active_subagents]` roster merges the live in-memory registry with the durable `subagent_sessions` store (so a cold-booted orchestrator still sees earlier workers); `continue_subagent` falls back from pause checkpoints to the durable store, resuming an idle worker with its persisted history; and a `workflow_proposal` payload found in a finished child's history is persisted as a parent-thread message (`extraMetadata.scope = "workflow_proposal"`) that the frontend rehydrates into the proposal card on thread load.

### Spawn hierarchy and tiers

Not every agent is allowed to spawn every other agent. The harness models a three-tier hierarchy that mirrors the cost / latency / depth-of-thought split between models:

```
Primary     (direct-capable, Master Agent on `coding` hint)
  │
  ├─► Worker      ◄─── fast path: one delegation, leaf does the work
  │
  └─► Reasoning   (slow, deep-thinking, e.g. planner on `reasoning` hint)
        │
        └─► Worker  ◄─── deep path: reasoning decomposes, workers execute
```

Each `AgentDefinition` carries an `agent_tier` field (`chat` / `reasoning` / `worker`, default `worker`). The contract:

| Tier        | May spawn             | Must NOT spawn                  | Typical members                               |
| ----------- | --------------------- | ------------------------------- | --------------------------------------------- |
| `chat`      | `reasoning`, `worker` | another `chat`                  | `orchestrator`                                |
| `reasoning` | `worker`              | another `reasoning`, any `chat` | `planner` (today the canonical one)           |
| `worker`    | nothing[^1]           | anything                        | critic, image\_agent, task\_manager\_agent, … |

Why the rules.

* *Chat → chat is meaningless.* The chat tier exists for snappy UX. A chat agent spawning another chat agent just doubles TTFT and burns tokens without buying any new capability.
* *Reasoning → reasoning blows up depth.* The reasoning tier is expensive. Chains of reasoning agents tend to re-decompose the same problem and create runaway hierarchies.
* *Worker → anything mixes execution and orchestration.* Workers are leaves so the parent always sees one compact result, not a transcript of nested delegations.

Enforcement. Two layers:

1. Loader-time (static). [`agents::loader::validate_tier_hierarchy`](https://github.com/tinyhumansai/openhuman/blob/main/crates/openhuman-core/src/agent/registry/agents/loader.rs) runs over the merged registry (built-ins + workspace TOMLs) and refuses to boot a registry that lists a same-tier or worker-with-subagents entry. Built-in archetypes are checked at compile-test time; user-shipped TOMLs are checked at workspace load.
2. Runtime depth gate (dynamic). Independent of tier, the sub-agent runner caps total spawn chain depth at `MAX_SPAWN_DEPTH = 3` via a task-local counter incremented across `run_subagent`, surfaced as a `SpawnDepthExceeded` agent error. This makes a user-shipped TOML that drops the tier annotation still unable to recurse past three hops.

> Status: the loader-time tier check, `agent_tier` field, and runtime depth-counter task-local are live. Depth is bounded by both the static loader contract and the runtime `MAX_SPAWN_DEPTH = 3` guard.

### Toolkit-specific specialists

For Composio toolkits with hundreds of actions (GitHub alone has 500+), loading every action into the sub-agent's tool set balloons prompt size. The harness ranks the toolkit's actions against the parent-refined task prompt with a cheap CPU-only filter (verb detection, token overlap, verb-alignment boost) and only loads the top-ranked subset into the sub-agent. No model call, pure heuristic, fast and explainable.

## Triage: handling external triggers

When a webhook fires, a cron ticks, or a Composio event arrives, the system can't just hand it straight to the orchestrator. Most triggers are noise; some warrant a notification; only a few deserve a full agent turn. The trigger-triage pipeline is the gate.

```
TriggerEnvelope ──► run_triage ──► TriageDecision ──► apply_decision
                       │                                     │
                       │                                     ├─► drop (noise)
                       │                                     ├─► notify only
                       │                                     ├─► spawn trigger_reactor
                       │                                     └─► spawn orchestrator
                       │
                       └── small local LLM (with cloud-LLM retry fallback)
```

The evaluator is intentionally cheap: a small local model where available, falling back to a remote model on retry. The decision is cached so identical triggers don't re-classify. Only triggers that escalate to "spawn orchestrator" go through the full `Agent::turn` machinery.

## Hooks: observability and policy levers

Two hook surfaces wrap the loop, on opposite ends:

### Stop hooks (mid-turn)

Stop hooks fire between iterations of the tool-call loop. They're the policy lever for budget caps, rate limits, and custom kill switches. Built-in hooks:

* Budget stop hook: caps cumulative turn cost in USD using the per-iteration cost accumulator.
* Max-iterations stop hook: caps iteration count from outside the agent's persistent config.
* Action budget policy: `SecurityPolicy` enforces `config.autonomy.max_actions_per_hour` for side-effecting tool operations. Users can tune it in Settings, Advanced, Agent autonomy, or operators can override it with `OPENHUMAN_MAX_ACTIONS_PER_HOUR`.

A hook returning `Stop` aborts the loop with a clear reason the caller can surface to the user. Stop hooks are distinct from interrupts (next section): they're policy-driven, not user-driven.

### Post-turn hooks

Post-turn hooks fire after the turn completes, in the background. They get a `TurnContext` snapshot, user message, assistant response, every tool call with arguments and outcome, total wall-clock, iteration count, session ID. Built-in consumers:

* Cost log: final per-turn cost line.
* Conversation logging: the `memory` lifecycle logs the reply after the turn, with one-line tool results (never arguments); there is no batching or idle flush. See [Memory architecture](/openhuman/developing/architecture/memory.md).

Long-term facts are written explicitly: the agent calls the `memory` tool's `learn` action. There is no automatic archivist.

Hooks run via `tokio::spawn`, so the user gets their answer before any of them finish.

## Interrupts: graceful cancellation

Cancellation is the tinyagents steering channel. When the user hits Ctrl+C or sends `/stop`, the runner forwards the request into the harness's steering/cancellation seam, which stops the loop at safe points: before each tool execution, before each sub-agent spawn and before each provider call:

* Every running sub-agent shares the cancellation scope and bails at its next checkpoint.
* In-flight provider streams are dropped.
* Buffered conversation turns are still stored (idle flush, or the exit flush on quit).

Interrupts are user-driven; stop hooks are policy-driven. Both enter the same harness pause/stop plumbing, but from different sides.

## Cost accounting

Every provider response carries a `UsageInfo` block, input tokens, output tokens, cached input tokens, and an authoritative `charged_amount_usd` populated by the OpenHuman backend. `TurnCost` sums those across every provider call inside one turn so the harness can:

* Emit per-iteration cost telemetry over the progress channel.
* Feed the budget stop hook so a runaway turn cuts itself off mid-loop.
* Log accurate end-of-turn cost lines.

When the backend doesn't surface a charged amount (older builds, providers that don't bill through it), a small per-tier rate table provides a token-rate floor estimate. Direct cost from the backend always wins when available.

## Web channel events and RPCs (assistant-ui elements pass)

The assistant-ui-elements integration added a batch of additive `WebChannelEvent`s and RPCs so the frontend can render tool args/timing, plan/goal/queue state, and turn lifecycle without polling. `EVENTS_VERSION` (`core/bus.rs`) is `1.4.0`; every new field is optional/defaulted so an older subscriber keeps parsing what a newer publisher emits.

New/extended socket events (bridged from `DomainEvent` onto `WebChannelEvent` by `web_chat::event_bus` and `openhuman_rpc::server::socketio`):

* `ts` (epoch ms) is stamped on every event by `publish_web_channel_event` when the producer left it unset, so the frontend can order/measure latency without guessing at receive time.
* `chat_done.timing`: `{ first_token_ms, first_tool_ms, total_ms, tokens_per_second }`, threaded from the progress bridge's per-turn `TurnTiming` through `ProgressBridgeHandle::timing_snapshot()`. `tokens_per_second` is derived from `output_tokens / (total_ms / 1000)` when both are known and `total_ms > 0`.
* `chat_done.usage.context_tokens`: input plus output tokens from the root agent's final model call. It feeds the context-window gauge and excludes earlier calls and sub-agents; `input_tokens` and `output_tokens` remain summed turn spend.
* `chat_cancelled`: `{ thread_id, client_id, request_id, cancel_reason: "user_stop" | "superseded", superseded_by }`. Emitted alongside (not instead of) the existing `chat_error{error_type:"cancelled"}`, on both the unscoped/scoped stop path and the superseded-by-a-newer-request path, and on the parallel-turn cooperative-cancel path (which previously published no terminal event at all).
* `chat_error{error_type:"guardrail"}`: carries a `guardrail: { verdict, score, reasons: [{code,message}] }` payload when `start_chat` rejects a message via `StartChatError::Guardrail` (the prompt-injection/security guardrail). Every other rejection stays `error_type:"inference"`. The RPC surface (`channel.web_chat`'s `Result<_, String>`) encodes the same structured verdict as a `GUARDRAIL:<json>` sentinel string (`web_chat::ops::start_chat::{GUARDRAIL_ERROR_PREFIX, is_guardrail_error_message}`, mirroring the existing `BACKEND_UNAVAILABLE:` pattern).
* `turn_cost`: live per-turn cost readout from `AgentProgress::TurnCostUpdated`, throttled to at most one emission per 750ms per turn (the first update always emits immediately): `{ thread_id, client_id, request_id, round, usage: { input_tokens, output_tokens, cached_input_tokens, cost_usd, context_window, context_tokens, subagents } }`. `context_tokens` is the root agent's final-call input plus output, for its context-window gauge; it is not summed over calls or sub-agents. `subagents` is always empty on this live event (it is the parent's cumulative rollup only); per-sub-agent attribution still only shows up on the terminal `chat_done.usage`.
* `approval_decided` / `plan_review_decided`: bridged from `DomainEvent::ApprovalDecided` / `PlanReviewDecided`, which gained `thread_id`, `client_id`, `tool_call_id`, and `resolution` (`"expired"` on TTL/sweep, `"cancelled"` on a dropped decision channel, `None` for an ordinary user decision, carried on the wire as the existing `cancel_reason` field). Only surfaced when the original park had both `thread_id` and `client_id` (chat-routed).
* `run_mode_changed`: `{ thread_id, client_id: "", message: "plan" | "build" }`, bridged from `DomainEvent::ThreadRunModeChanged`, published by `agent::tinyagents::run_mode::set_mode`.
* `thread_todos_changed`, `queue_item_queued` / `queue_item_delivered` / `queue_item_removed`, `queue_item` carries `{ id, lane, text_preview }`.
* Sub-agent correlation: `subagent_spawned`/`subagent_completed`/`subagent_failed`/`subagent_awaiting_user` carry `subagent.parent_call_id` (the spawning tool call's id: every worker of one `spawn_parallel_agents` call shares it); `subagent_completed` also carries `subagent.output` (capped final text). `tool_call.args`/`tool_result.args`+`elapsed_ms` now carry real arguments and timing instead of `null` (from tinyagents' `AgentEvent::ToolStarted.input`).
* Artifact events (`ArtifactPending`/`ArtifactReady`/`ArtifactFailed`) carry `tool_call_id` and `turn_request_id`, the latter filled from the originating turn's `ApprovalChatContext::request_id` so the frontend can bind an artifact card to the exact turn that produced it (`None` for CLI/cron/sub-agent producers with no chat context).

New RPCs:

* `agent.set_run_mode { thread_id, mode }` / `agent.get_run_mode { thread_id }`, read/flip a thread's Plan/Build `RunModeHandle`. `channel.web_chat` also accepts an optional `run_mode: "plan" | "build"` param (mirroring the socket `chat:start` payload) so a turn can start with the thread already in the requested mode instead of racing a separate RPC call; unrecognized values are logged and ignored.
* `threads.token_usage { thread_id }`: returns persisted thread spend totals plus `last_turn_context_tokens`, the root agent's final-call input plus output used by the context-window gauge. Older turns without per-call usage fall back to their summed input plus output.
* `threads.goal_get` / `threads.todos_get`: read a thread's current goal/todo state directly .
* `threads.edit_message { thread_id, message_id, content, client_id? }` / `threads.regenerate { thread_id, message_id?, client_id? }`, both return `{ request_id }`. They cancel the thread's in-flight turn, fork the session transcript at a cut point via `TranscriptLocator::truncate_into_next_generation` (the sealed generation the fork was cut from is never touched, same "compaction never erases" guarantee), truncate the conversation-store message log and the turn-state snapshots for every dropped turn, then restart the turn with the edited content (`edit_message`) or the original prompt (`regenerate`). See `threads::ops::edit`'s module doc for the full UI-message-id -> transcript-cut-point mapping.
* `channel.web_queue_remove { client_id, thread_id, item_id }`, remove one specific queued item from a thread's run queue.
* `agent.context_breakdown { agent_id?, thread_id? }`, a UI-friendly view over the agent's rendered prompt size (system/tools/history split) for the composer's context-usage indicator.
* `commands.list`: merges built-in commands with `skills.list` and `flows.list` into one slash-command catalog.

`plan_exit` (the tool that flips a thread from Plan back to Build) only flips the mode when there is no plan review still parked on that thread (`agent::plan_review::gate::PlanReviewGate::parked_review_for_thread`), a review resolves before `request_plan_review` returns control to the agent, so this only matters for a mis-timed/concurrent `plan_exit` call racing a still-pending review, which must not unlock every tool before the user has actually approved anything.

## Explicit run context and host capabilities

`OpenHumanRunContext` is the live carrier at the shared chat, channel, and sub-agent turn seam. Roots snapshot their currently scoped origin, progress, stop hooks, dispatch state, thread, route slot, and workspace grant, then own or explicitly receive one cancellation token before passing the context to the runner; a recursive sub-agent forks it with `child()`, preserving shared cancellation/policy handles while isolating route observation and usage accounting. The shared runner creates `RunContext<OpenHumanRunContext>` through `into_tinyagents`, and the OpenHuman assembly and middleware registry consume that typed context directly. Route, and thread scopes still surround the drive only for legacy tool/model APIs outside the typed harness boundary. Recursive fan-out receives cancellation and workspace from its typed parent `RunContext` through `ToolDispatch`; it never reads ambient cancellation state.

`OpenHumanHostBundleFactory` is the B1 host composition point. It constructs the context, definition, security, model, memory, budget, progress, tool-outcome and experience adapters from the same session/runtime inputs. OpenHuman retains all policy decisions; TinyAgents receives only the resulting canonical run context today. Wiring the complete host-capability bundle into every invocation remains a later cutover step.

## Self-healing recap

A few small adaptive systems sit on top of the main loop:

* Tool-output summary circuit-breaker: three consecutive summary failures in a thread make TinyJuice stop asking for summaries in that thread, so results fall back to compaction and truncation.
* Triage local-vs-remote retry: local LLM first; remote fallback on parse failure.
* Unknown-tool and malformed-argument recovery: middleware rewrites an invalid model tool call into a recoverable result instead of aborting the run.

None of these change the loop's shape, they just make the common failure modes recoverable without the user having to intervene.

## Where to look in the code

The harness shell lives under `crates/openhuman-core/src/agent/`, with the tinyagents adapter seam in `crates/openhuman-core/src/agent/tinyagents/` and archetype definitions in `crates/openhuman-core/src/agent/registry/`. The README in `crates/openhuman-core/src/agent/` enumerates the public surface; the most load-bearing files (paths relative to `crates/openhuman-core/src/agent/` unless prefixed) are:

| File / dir                         | What lives there                                                                                                                                                             |
| ---------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `session_host/runtime_session.rs`  | `Agent::turn`: the lifecycle described above; routes into the tinyagents runner via `tinyagents/turn_runner.rs`.                                                             |
| `tinyagents/mod.rs`                | `run_turn_via_tinyagents_shared`: the shared tinyagents harness assembly (the live loop).                                                                                    |
| `tinyagents/middleware.rs`         | The named OpenHuman middleware stack (approval/security, tool policy, recovery, budgets, circuit breaker).                                                                   |
| `harness/graph.rs`                 | The channel/CLI bus turn route into the tinyagents runner.                                                                                                                   |
| `subagent_host/`                   | Direct OpenHuman planner/executor/persistence adapters for `tinyagents-orchestration::subagent`; product prompt/tool/model/security/artifact/progress behavior remains here. |
| `orchestration/subagent_sessions/` | Durable reusable sub-agent identity, compatibility matching, persisted status/history.                                                                                       |
| `harness/definition.rs`            | `AgentDefinition`: what an archetype declares.                                                                                                                               |
| `subagent_host/ops/runner.rs`      | Integration-tool ranking and the host execution leaf; generic lifecycle stays in `tinyagents-orchestration`.                                                                 |
| `tinyagents/payload_summarizer.rs` | The model call behind TinyJuice's oversized-tool-result summary.                                                                                                             |
| `session_host/tool_progress.rs`    | Surviving OpenHuman seam: `TurnProgress`.                                                                                                                                    |
| `message_convert.rs`               | Concrete durable/provider conversion around canonical tool-call dialect APIs.                                                                                                |
| `triage/`                          | External-trigger classification + escalation.                                                                                                                                |
| `registry/agents/`                 | Built-in archetypes: one subdirectory per agent.                                                                                                                             |
| `hooks.rs` / `stop_hooks.rs`       | Post-turn and mid-turn hook surfaces.                                                                                                                                        |
| `cost.rs`                          | Per-turn USD/token accounting.                                                                                                                                               |
| `progress.rs`                      | Real-time progress events to the UI.                                                                                                                                         |

## Agent engine and orchestration on tinyagents

Every agent turn uses TinyAgents. Chat and channel/CLI callers invoke the session/turn host, while a sub-agent call enters `agent/subagent_host` and its direct `tinyagents-orchestration::subagent` lifecycle. There is no `harness/subagent_runner` compatibility path. The surrounding OpenHuman seam:

| File (`crates/openhuman-core/src/agent/tinyagents/`) | Role                                                                                                                                                                                                                                                                                                          |
| ---------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `mod.rs`                                             | The runner (`run_turn_via_tinyagents_shared`): installs the native `ChatModel`, host tool adapters, and middleware on an `AgentHarness`; runs one turn; caps output through `MaxTokensModel` (`agent/tinyagents/model.rs`); mirrors progress; forwards steering; and pauses gracefully at the model-call cap. |
| `mod.rs` / `model.rs` / `tools.rs` / `convert.rs`    | `RunPolicy` / `ChatModel` / `Tool` / message adapters (incl. unknown-tool policy and out-of-band reasoning forwarding).                                                                                                                                                                                       |
| `observability.rs`                                   | Harness `AgentEvent` → `AgentProgress` + cost; `GraphTracingSink` for graph events.                                                                                                                                                                                                                           |
| `orchestration.rs`                                   | Re-exported `graph::orchestration` task-store types; map-reduce fanout uses the TinyAgents SDK surface directly.                                                                                                                                                                                              |
| `delegation.rs`                                      | The durable `plan → execute ⇄ review → finalize` delegation graph (production worker wired in `orchestration::delegation`), checkpointed via TinyAgents' own `SqliteCheckpointer`.                                                                                                                            |

Orchestration on graphs (`crates/openhuman-core/src/agent/orchestration/`):

* Workflow phase DAG (`workflow_runs/engine.rs`) runs on a `dispatch ⇄ run_phase → done` conditional-routing graph; each phase fans its agents out via `graph::parallel::map_reduce`. The durable `workflow_runs` row stays the source of truth (controllers + resume read it).
* Team member runtime (`agent_teams/runtime.rs`, `run_member_graph`) is a conditional-routing graph (`execute → complete|fail → done`).
* Multi-stage delegation (`orchestration::delegation` + the `delegate` tool) runs `delegation.rs`, checkpointed to the session DB.
* Detached sub-agents (`running_subagents.rs`) use TinyAgents `DetachedTaskRegistry` for ownership-aware snapshots, wait/timeout, steering lookup, cooperative cancellation, hard abort, and terminal cleanup. OpenHuman retains durable task-store projection, product/session metadata, RPC and delivery semantics, and the `RunQueue` compatibility fallback.

Deliberately kept off the crate's primitives (these are design decisions, not gaps):

* Sub-agent host policy (`subagent_host/`) stays OpenHuman-owned: definition resolution, archetype tool filtering, provider resolution, narrow prompt building, memory context, worker-thread mirror, handoff configuration and product checkpoint/session projection. It implements the TinyAgents planner/executor/persistence traits directly. The neutral driver, not a host-local runner, owns lifecycle sequencing and task-key coalescing.
* Durable run ledgers (`workflow_runs`, `agent_teams`, `command_center`, `subagent_sessions`) stay on openhuman SQLite/JSON until their controller projections and restart semantics are mapped onto TinyAgents task/status/journal records. The `agent_teams` race-safe SQL compare-and-swap task claim remains OpenHuman-owned.

> TinyAgents ships harness store/cache/session primitives (`harness::store` with JSONL append stores, `harness::cache`, `harness::subagent`, lineage-aware status), graph task stores, the detached runtime registry, and conformance contracts. The session shell and product-specific sub-agent build/delivery pipeline remain OpenHuman-owned.

## Reliability: breakers, handback and classified failures

Three cooperating mechanisms keep runs from wandering or dying silently:

No-progress circuit breaker (`RepeatedToolFailureMiddleware`, `crates/openhuman-core/src/agent/tinyagents/middleware.rs`) is a thin driver over the crate's `NoProgressTracker`. It fingerprints each tool call's arguments and feeds outcomes into an escalation ladder: `Continue` → `Nudge` (a structured "no progress since step X" corrective injected via `SteeringCommand::InjectMessage`, which is safe inside interactive turns) → `Halt` (record a root-cause summary into the `HaltSummarySlot`, pause via the steering handle). Identical arguments retried count toward the trip (threshold 3 consecutive identical failures); *recoverable* failures (timeouts, connection resets, rate limits, 5xx) get an extended headroom ladder instead of the fixed crate thresholds.

Sub-agent handback (`subagent_host/`): a sub-agent run resolves to one of three statuses:

* `Completed`: clean final response.
* `AwaitingUser { question, options }`: the child called `ask_user_clarification`; the host projects its full checkpoint (history, question, options and overrides) under the complete TinyAgents task key `(root_run_id, parent_run_id, thread_id, task_id)`. Resume recovers that original key rather than deriving a new one from the fresh parent turn.
* `Incomplete { reason }`: the child was halted by the breaker or hit its model-call cap. The delegating parent relays the blocker instead of treating a halted child as a finished answer or re-spinning the identical delegation.

A breaker halt at the top level is likewise never a silent finish, and the breaker's root-cause summary is not shown to the user as is either: it is worded for a model ("Report this back instead of retrying"). `hit_cap` / `breaker_halt` are surfaced on the turn result, and the chat turn closes the halted run the same way it closes a tool turn that ended without final text (`session_host/driver/grounded_close.rs`):

1. A tools-disabled wrap-up call whose instruction restates the turn's tool records, each failure's own message included, with the breaker summary passed as a stop note to explain rather than repeat.
2. A separate check call that sees only the request, the records and the candidate reply. It rejects a reply that only narrates intent, contradicts a record, or leaves out the failure that explains an unfinished request.
3. A deterministic fallback for an empty, tool-calling, rejected or unverified reply (a check that failed or gave no verdict). It quotes each tool result and the stop note.

Accepted text is streamed only after the check, so a rejected reply never renders.

Classified tool failures (`crates/openhuman-core/src/tools/status/`): every failed tool call is classified into a transport-agnostic `ClassifiedFailure { class, category, cause_plain, next_action, recoverable }`. Classes cover `MissingPermission`, `MissingApp`, `ServiceUnavailable`, `BadCredentials`, `BlockedByPolicy`, `ModelConnection`, `Timeout`, `Denied`, `ApprovalExpired`; categories map 1:1 to UI states: *recoverable* (safe auto-retry), *blocked by policy* (change settings), *needs user confirmation* (sign in / install / grant), *user declined* (never auto-retried). The classification rides `AgentProgress::ToolCallCompleted.failure` (including for sub-agent calls) into the chat timeline.

## Journals, replay and migration shadows

Every run appends to a durable event journal (`tinyagents/journal.rs`): a `StoreEventJournal` over a JSONL append store at `{workspace}/tinyagents_store/journal`, composed as `FanOutSink` (live bridge + journal) → `RedactingSink` (credential masking before persistence), with restart-stable event ids (`{run_id}-evt-{offset}`). Even an unobserved background turn is reconstructable after the fact. Three read-only RPCs expose it: `agent_run_events` (paged, late-attach replay by `run_id`/`offset`/`limit`), `agent_run_status` (latest harness status), and `agent_runs_active` (active runs, filterable by thread or root run).

The remaining store cutover runs on shadow scaffolding. Product behavior is unchanged and divergences are logged:

* Session dual-write / shadow read (`session_host/runtime_session.rs`): session messages dual-write into the TinyAgents store (default-ON flag `config.session_dual_write`); loads shadow-read for parity while the legacy file store stays authoritative.

Goals and todos are crate-backed outright, with no shadow: thread goals live in the crate `graph.goals` KV store (`agent/goals/store.rs`), and the session todo list lives in the in-process crate `graph.todos` store (`agent/todos/ops.rs`); see [Goals and todos](/openhuman/features/goals-and-todos.md).

## Workload routes and the burst tier

`tinyagents/routes.rs` is the declarative TinyAgents `ModelRouter` for the OpenHuman workload roles `chat`, `reasoning`, `agentic`, `coding`, `burst`, `summarization`, and `vision`, keyed by their `hint:*` aliases. It owns fallback chains and capability gates; `inference::provider::factory` resolves each selected role to its configured native `ChatModel`, on the managed backend, the pinned default model (`openrouter/deepseek/deepseek-v4-flash` unless changed under Settings → Routing → Default model); there are no per-role tier endpoints. The `burst` role serves low-context, high-fanout workers.

## See also

* [Architecture overview](/openhuman/developing/architecture.md): where the harness sits in the bigger picture.
* [Memory](/openhuman/features/memory.md) and [Memory architecture](/openhuman/developing/architecture/memory.md): the `memory` tool, the per-turn memory lifecycle and conversation logging.
* [Automatic model routing](/openhuman/features/model-routing.md): how `model: "hint:reasoning"` resolves to a concrete provider+model.
* [Native tools: agent coordination](/openhuman/features/native-tools/agent-coordination.md): the user-facing surface for `spawn_subagent`, `delegate_*`, `todo`.

[^1]: Skill-wildcard entries (`{ skills = "*" }`) are exempt because they name no agent: they expand to the connected Composio actions as `Deferred` tools the agent reaches through `tool_search`, not to a spawn.
