> ## Documentation Index
> Fetch the complete documentation index at: https://docs.oao.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# Adding models and providers

> Configure organization-shared OpenRouter, OpenAI, Anthropic, or xAI credentials and link approved models to agents.

OAO stores hosted model configuration in PostgreSQL. Provider connections are
organization-shared — every project uses the same pool — while model presets
stay per project. The MVP supports four provider types: **OpenRouter**,
**OpenAI**, **Anthropic**, and **xAI (Grok)**.
An agent version never contains a raw model identifier or credential. It names
an immutable model preset, and that preset references a provider connection in
the same organization.

## Security model

Provider API keys are write-only:

* the API encrypts each key with AES-256-GCM before inserting it;
* the authenticated encryption context includes organization ID, provider ID,
  provider type, and credential version;
* PostgreSQL stores ciphertext, a unique nonce, authentication tag, key
  version, and SHA-256 fingerprint;
* list and create responses expose only the first 12 fingerprint characters;
* the runtime decrypts a key only while activating the selected project preset;
* API keys never appear in model presets, agent snapshots, SSE, audit detail,
  logs, or traces.

The encryption key is a platform secret, not a provider credential. Generate a
32-byte key and supply the same value to the API and runtime worker:

```bash theme={null}
openssl rand -base64 32
```

```dotenv theme={null}
OAO_CREDENTIAL_ENCRYPTION_KEY=replace-with-the-generated-base64-value
```

Do not put `OPENROUTER_API_KEY`, `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, or
`XAI_API_KEY` in `.env`. Those keys are entered through the project UI or API
and stored as tenant-scoped ciphertext.
Without `OAO_CREDENTIAL_ENCRYPTION_KEY`, provider creation, preset publication,
and runtime execution are unavailable. OAO does not fall back to a local fake
model.

## Add a provider in the console

1. Open **Models**.
2. Select **Add provider**.
3. Choose **OpenRouter**, **OpenAI**, **Anthropic**, or **xAI (Grok)**.
4. Enter a stable project-local key, display name, and API key.
5. Select **Add provider**.

The console clears the submitted secret when the dialog closes. It can display
the provider type, credential fingerprint, and version, but it cannot retrieve
the key.

To replace a key, select **Rotate key** next to the provider. Rotation writes a
new nonce, authentication tag, ciphertext, fingerprint, and incremented
credential version. Existing presets continue to reference the same provider.
The runtime replaces its in-memory provider on the next activation.

To remove a connection, select **Remove** next to the provider and confirm.
Removal is refused while any model preset still routes through the connection;
archive those presets first. Removing wipes the encrypted key, releases the
provider key for a new connection, and makes runs of agents pinned to its
archived presets fail with a provider-removed error.

## Add a model preset

1. Select **Add model preset**.
2. Choose a provider connection.
3. In **Model**, type to search the live catalog available to that connection
   and pick a model or OpenRouter saved preset. OAO uses each connection's
   authenticated live catalog. xAI connections use `GET /v1/language-models`
   so image-generation, video, voice, and embedding-only IDs are not offered.
   The picker shows the first
   matches with their model identifier, context window, and reasoning support,
   and reports how many further matches the search is hiding.
4. Check the suggested preset key and display name. Both are derived from the
   chosen model — for example `claude-sonnet-4-6-v1` — with the version suffix
   moved forward past keys this project already uses. Editing either field
   stops the suggestion from overwriting it.
5. For OpenRouter, open **Routing and data policy** to set routing and
   data-handling constraints. It stays folded away while it is unset, and the
   summary next to it always shows the policy that will be stored.
6. For OpenAI, review **Model settings**. New presets default to plain text,
   standard reasoning mode, medium reasoning effort, medium verbosity, and an
   automatic reasoning summary. GPT-5.6 presets can select standard or pro
   mode and reasoning effort from none through max. GPT-6 Astra supports low,
   medium, high, xhigh, and max; reasoning cannot be disabled.
7. For Anthropic, review **Claude settings**. Thinking can be disabled or set
   to adaptive when the model advertises that capability. `maxTokens` is the
   total output ceiling shared by thinking and the visible answer. Effort is
   limited to the levels advertised by the selected model; new adaptive
   presets use 20,000 tokens and high effort when supported.
8. For a supported xAI reasoning model, review **Grok settings**. The response
   format is text and reasoning effort defaults to high. Grok 4.6 offers low,
   medium, high, and xhigh; the picker only offers levels OAO has verified for
   the selected model. Tools and the system prompt stay on the agent version,
   where their schemas and instructions are immutable together.
9. Select **Add model preset**.

The dialog lists what is still missing under **Validation** and keeps the
submit button disabled until nothing is.

Preset rows are append-only. A published agent version keeps referring to the
same reviewed model, routing policy, and generation settings. To change any of
them, create a new preset key and publish a new agent version.

OpenRouter presets may use live OpenRouter models or saved OpenRouter presets
(`@preset/<slug>`), plus routing controls including fallbacks, zero-data
retention, data collection, upstream allow/deny lists, preference order, sort,
and price caps. OpenAI, Anthropic, and xAI are direct provider connections, so
OpenRouter routing controls must be empty.

For direct OpenAI presets, OAO maps the stored controls to the Responses API as
`text.format`, `text.verbosity`, `reasoning.mode`, `reasoning.effort`, and
`reasoning.summary`. Reasoning effort is also the Flue default for the agent;
an explicitly configured Harness Operation may still override its own effort.

For direct Anthropic presets, OAO maps the stored controls to the Messages API
as `thinking.type`, `max_tokens`, and `output_config.effort`. Anthropic counts
thinking tokens against `max_tokens`, so this setting is the hard ceiling for
the combined thinking and visible response. Adaptive thinking and the available
effort values are model capabilities rather than assumptions made by OAO.
Manual `thinking.type: "enabled"` and `budget_tokens` are not part of this
preset contract; use adaptive thinking for supported current models or disable
thinking.

For direct xAI presets, OAO stores the exact `xai/<catalog-id>` returned by the
live language-model catalog and calls `https://api.x.ai/v1`. Newly available
text-output Grok IDs can be selected without an OAO catalog release. Known Pi
models retain their pinned context/output metadata; a newly discovered ID uses
conservative runtime limits until the pinned Pi catalog adds richer metadata.
For documented xAI reasoning models, OAO stores `textFormat` and `effort`, then
maps them to `text.format` and `reasoning.effort` on the Responses API. xAI
documents high as the default, with low, medium, and high on Grok 4.5 and xhigh
from Grok 4.6 onward. Reasoning cannot be disabled on these models. Structured
JSON schemas are operation-specific rather than a fixed model default. xAI
server-side tools are not part of this preset contract; OAO's versioned agent
tools remain available.

### Inspect, duplicate, or archive a preset

Select a preset in the **Model presets** table to open it. The dialog shows the
key, model, provider connection, routing policy or generation settings as JSON,
and the agents that currently pin the key.

* **Duplicate as new preset** opens the add dialog prefilled with the same
  provider, model, policy, and settings under the next free key. This is how
  you "change" a preset: the original stays exactly as published agent versions
  reviewed it.
* **Archive preset** removes it from the table and from the agent editor's
  preset picker. Agents that pin the key keep running on it, but must move to
  another preset before they can publish a new version. The key stays reserved.

## Model timeouts and retries

Each model attempt has a **five-minute deadline**, including connection setup
and the entire response stream. Partial output does not reset the clock. OAO
aborts an expired request and discards late results, including late tool calls.
This deadline applies to all hosted providers and is fixed by the runtime;
there is no per-preset timeout setting.

The deadline bounds receiving the provider response. Database admission before
that request, processing buffered output, and recording run progress are separate
work, so elapsed run time can exceed a single attempt's deadline. A model-start
event records the requested turn; it does not prove that the provider has received
the request.

OAO checks the durable model-turn budget once per model operation, including new
retry and recovery operations. Reading additional response chunks does not repeat
that database transaction or consume extra turns. The public progress recorder
filters unused text, reasoning, and tool-argument deltas before database access;
model outcomes, retries, tool results, and recovery events remain recorded.

For agent model turns, timeouts and other transient provider failures use Flue's
durable retry budget: **up to three retries after the initial attempt**, or four
attempts total. Retries back off for approximately 2, 4, and 8 seconds with
jitter. Provider SDK retries are disabled so they cannot multiply this budget.
Completed tools remain committed and are not executed again by a model retry.
Cancellation, authentication failures, and the platform model-turn guard do not
trigger these retries. Retry attempts still consume the agent's model-turn
budget, and a run or Harness Operation deadline can stop execution sooner.

Four consecutive five-minute timeouts can therefore take about **20 minutes
plus retry delays**. Reduce reasoning effort or requested output size when a
model repeatedly exceeds the deadline; retrying the same large generation can
time out again. Internal context-compaction requests share the five-minute
deadline but do not use the agent-turn retry loop.

The session event inspector shows **Model call started** and **Model retry
scheduled** entries while an agent turn is pending. These events contain safe
timing and retry metadata, never provider credentials or raw error bodies.
Timed-out attempts may still incur provider charges; without final provider
usage, OAO cannot report their complete token usage or cost.

## Prompt caching

OAO passes each Flue conversation's stable session identity to OpenRouter as
`x-session-id`, allowing OpenRouter to keep later turns on the same cache-aware
provider endpoint. For direct `openrouter/anthropic/*` models, OAO also adds
Anthropic-compatible ephemeral cache markers to the stable system prompt, the
last tool definition, and the latest conversation content. A qualifying first
request can therefore report `cacheWriteTokens`; later requests with the same
prefix can report `cacheReadTokens` in the session usage panel and API.

Some providers use implicit prompt caching and do not report writes because
creating the cache is free. In that case `cacheWriteTokens` remains zero even
when later calls report cache reads. OpenRouter saved presets
(`openrouter/@preset/<slug>`) can resolve to different model families, so their
explicit cache behavior must be configured in the saved preset; OAO still
sends session affinity. Provider minimum prompt sizes and cache TTLs continue
to apply.

## Link a model to an agent

Open **Agents**, create or edit an agent, and choose an available model preset.
Publishing validates the preset inside the same tenant-scoped transaction. A
project preset is available only when it references a provider connection in
that project and platform credential encryption is configured.

## HTTP API

All routes are scoped to `/v1/projects/{projectId}`. Reads require
`agent:read`; writes require `project:admin`.

### Create a provider connection

```bash theme={null}
curl -sS -X POST "$API/v1/projects/$PROJECT_ID/model-providers" \
  -H 'content-type: application/json' \
  -H 'idempotency-key: provider-openrouter-1' \
  -b "$COOKIE_JAR" \
  --data '{
    "key": "openrouter-primary",
    "displayName": "OpenRouter primary",
    "providerType": "openrouter",
    "apiKey": "replace-with-project-api-key"
  }'
```

The response contains `credentialConfigured: true`, a 12-character
`credentialFingerprint`, and `credentialVersion: 1`. It never contains
`apiKey`, ciphertext, nonce, or authentication tag.

### List providers

```bash theme={null}
curl -sS -b "$COOKIE_JAR" \
  "$API/v1/projects/$PROJECT_ID/model-providers?limit=200"
```

### Rotate a provider key

```bash theme={null}
curl -sS -X PUT \
  "$API/v1/projects/$PROJECT_ID/model-providers/$PROVIDER_ID/credential" \
  -H 'content-type: application/json' \
  -H 'idempotency-key: provider-rotation-2' \
  -b "$COOKIE_JAR" \
  --data '{"apiKey":"replace-with-rotated-key"}'
```

### Search the matching catalog

```bash theme={null}
curl -sS -b "$COOKIE_JAR" \
  "$API/v1/projects/$PROJECT_ID/model-catalog?providerId=$PROVIDER_ID&search=gpt&limit=200"
```

The provider ID is required. OAO reads its type from PostgreSQL and returns only
the matching provider catalog. For OpenRouter connections, OAO decrypts the
project provider key server-side and fetches OpenRouter's live `/models` catalog
plus saved `/presets`, exposing saved presets as `openrouter/@preset/<slug>`
options. For direct OpenAI and Anthropic connections, OAO decrypts the selected
project key server-side and fetches the provider's live `/v1/models` catalog.
For xAI, OAO fetches `/v1/language-models` and exposes text-output entries as
`xai/<catalog-id>` without relying on a hard-coded Grok list. OpenAI uses the
live list for account membership and discovers new GPT-4.1, GPT-5-and-later, and
o3-and-later family IDs without a pinned catalog update. Known
non-Responses models and specialized audio, realtime, transcription, search,
image, embedding, moderation, and deep-research variants are excluded.
Models without verified runtime capabilities and pricing appear with
`runtimeSupported: false`. The console shows them disabled with an explanation;
the CLI setup only offers supported models. Preset creation rejects unsupported
entries, so naming patterns cannot authorize execution or produce zero-cost usage.
Anthropic responses are intersected with OAO's pinned runtime models.
Anthropic entries also expose
the model's adaptive-thinking, thinking-disable, and effort capabilities.
Claude models whose live capabilities do not support effort are not offered,
because every OAO Anthropic preset pins that setting.
Catalog responses contain public metadata only; no provider key is returned.

### Create a preset

```bash theme={null}
curl -sS -X POST "$API/v1/projects/$PROJECT_ID/model-presets" \
  -H 'content-type: application/json' \
  -H 'idempotency-key: model-preset-1' \
  -b "$COOKIE_JAR" \
  --data "{
    \"key\": \"support-model-v1\",
    \"displayName\": \"Support model\",
    \"providerId\": \"$PROVIDER_ID\",
    \"model\": \"openrouter/anthropic/claude-sonnet-4.6\",
    \"routing\": {
      \"dataCollection\": \"deny\",
      \"zeroDataRetention\": true,
      \"allowFallbacks\": false
    }
  }"
```

For a direct OpenAI connection, use an `openai/<catalog-id>` model returned by
the catalog, send an empty `routing` object, and include the immutable model
settings:

```json theme={null}
{
  "key": "gpt-5-6-terra-v1",
  "displayName": "GPT-5.6 Terra",
  "providerId": "55555555-5555-4555-8555-555555555555",
  "model": "openai/gpt-5.6-terra",
  "routing": {},
  "settings": {
    "textFormat": "text",
    "mode": "standard",
    "effort": "medium",
    "verbosity": "medium",
    "summary": "auto"
  }
}
```

For a direct Anthropic connection, use an `anthropic/<catalog-id>` model
returned by the catalog, send an empty `routing` object, and include the
immutable Claude settings:

```json theme={null}
{
  "key": "claude-sonnet-5-v1",
  "displayName": "Claude Sonnet 5",
  "providerId": "55555555-5555-4555-8555-555555555555",
  "model": "anthropic/claude-sonnet-5",
  "routing": {},
  "settings": {
    "thinking": "adaptive",
    "maxTokens": 20000,
    "effort": "high"
  }
}
```

For a direct xAI connection, use an `xai/<catalog-id>` returned by the live
language-model catalog. Routing must be empty. Supported reasoning models
accept immutable text-format and effort settings:

```json theme={null}
{
  "key": "grok-4-6-v1",
  "displayName": "Grok 4.6",
  "providerId": "55555555-5555-4555-8555-555555555555",
  "model": "xai/grok-4.6",
  "routing": {},
  "settings": {
    "textFormat": "text",
    "effort": "low"
  }
}
```

## Tenant and lifecycle guarantees

* Provider and preset primary keys include organization and project identity.
* Composite foreign keys prevent cross-project provider references.
* Row-level security is enabled and forced on both relations.
* Provider identity and type are immutable; only encrypted credential material
  may rotate, and its version must increase.
* Model presets are immutable and cannot be updated or deleted.
* Existing presets created before provider connections were introduced remain
  visible as unavailable legacy history. Recreate them with a new versioned key
  linked to a provider connection.

## Troubleshooting

### Provider creation says encryption is not configured

Set the same canonical base64 32-byte `OAO_CREDENTIAL_ENCRYPTION_KEY` for the
API and runtime worker, then restart both processes. Do not replace it while
ciphertext encrypted with the old key still exists; key-ring rotation is not
part of this MVP.

### A provider key cannot be decrypted

Confirm the API and runtime use the same platform encryption key and that the
database row was not copied between tenants or providers. Tenant/provider
identity is authenticated, so copied ciphertext intentionally fails.

### A model is not in the catalog

Search `/model-catalog` with the intended provider connection. OpenRouter,
OpenAI, Anthropic, and xAI use separate catalogs and prefixes. If an OpenRouter model
or saved preset is missing, confirm the provider API key can read the OpenRouter
catalog. If a direct OpenAI model is missing, confirm the selected key can list
it through OpenAI's `/v1/models` endpoint and belongs to a supported text/reasoning
family. Each catalog request fetches the current account list; OAO does not
cache it or require a library release to discover new family members.

Known OpenAI models retain their runtime metadata. GPT-6 Astra has documented
metadata for its 1,050,000-token context window, 128,000-token output limit,
image input, pricing, and reasoning levels, so it can be selected immediately
when the connected key lists it. Models without verified metadata show unknown
context/output limits and remain disabled until runtime support is added.
Discovery is automatic; enabling a new model still requires verified endpoint,
tool, reasoning, limit, and pricing metadata. This prevents failed runs and
incorrect cost reporting.

If a direct Anthropic model is missing, confirm the selected key can list it
through Anthropic's `/v1/models` endpoint. OAO must also have matching native
Anthropic runtime metadata. The model catalog controls whether adaptive
thinking can be enabled or disabled and which effort levels can be saved.

If a Grok model is missing, confirm the selected xAI key can list it through
xAI's `/v1/language-models` endpoint and that the entry has text output. Image,
video, voice, and embedding-only models are intentionally excluded from agent
presets.

### A preset is unavailable

Confirm it has a non-null provider connection, the connection belongs to the
same organization, and the runtime has the platform encryption key. Legacy
unbound presets are history only.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.