Skip to main content
OAO stores hosted model configuration in PostgreSQL. Provider connections are organization-shared — every project uses the same pool — while model presets stay per project. The MVP supports four provider types: OpenRouter, OpenAI, Anthropic, and xAI (Grok). An agent version never contains a raw model identifier or credential. It names an immutable model preset, and that preset references a provider connection in the same organization.

Security model

Provider API keys are write-only:
  • the API encrypts each key with AES-256-GCM before inserting it;
  • the authenticated encryption context includes organization ID, provider ID, provider type, and credential version;
  • PostgreSQL stores ciphertext, a unique nonce, authentication tag, key version, and SHA-256 fingerprint;
  • list and create responses expose only the first 12 fingerprint characters;
  • the runtime decrypts a key only while activating the selected project preset;
  • API keys never appear in model presets, agent snapshots, SSE, audit detail, logs, or traces.
The encryption key is a platform secret, not a provider credential. Generate a 32-byte key and supply the same value to the API and runtime worker:
Do not put OPENROUTER_API_KEY, OPENAI_API_KEY, ANTHROPIC_API_KEY, or XAI_API_KEY in .env. Those keys are entered through the project UI or API and stored as tenant-scoped ciphertext. Without OAO_CREDENTIAL_ENCRYPTION_KEY, provider creation, preset publication, and runtime execution are unavailable. OAO does not fall back to a local fake model.

Add a provider in the console

  1. Open Models.
  2. Select Add provider.
  3. Choose OpenRouter, OpenAI, Anthropic, or xAI (Grok).
  4. Enter a stable project-local key, display name, and API key.
  5. Select Add provider.
The console clears the submitted secret when the dialog closes. It can display the provider type, credential fingerprint, and version, but it cannot retrieve the key. To replace a key, select Rotate key next to the provider. Rotation writes a new nonce, authentication tag, ciphertext, fingerprint, and incremented credential version. Existing presets continue to reference the same provider. The runtime replaces its in-memory provider on the next activation. To remove a connection, select Remove next to the provider and confirm. Removal is refused while any model preset still routes through the connection; archive those presets first. Removing wipes the encrypted key, releases the provider key for a new connection, and makes runs of agents pinned to its archived presets fail with a provider-removed error.

Add a model preset

  1. Select Add model preset.
  2. Choose a provider connection.
  3. In Model, type to search the live catalog available to that connection and pick a model or OpenRouter saved preset. OAO uses each connection’s authenticated live catalog. xAI connections use GET /v1/language-models so image-generation, video, voice, and embedding-only IDs are not offered. The picker shows the first matches with their model identifier, context window, and reasoning support, and reports how many further matches the search is hiding.
  4. Check the suggested preset key and display name. Both are derived from the chosen model — for example claude-sonnet-4-6-v1 — with the version suffix moved forward past keys this project already uses. Editing either field stops the suggestion from overwriting it.
  5. For OpenRouter, open Routing and data policy to set routing and data-handling constraints. It stays folded away while it is unset, and the summary next to it always shows the policy that will be stored.
  6. For OpenAI, review Model settings. New presets default to plain text, standard reasoning mode, medium reasoning effort, medium verbosity, and an automatic reasoning summary. GPT-5.6 presets can select standard or pro mode and reasoning effort from none through max. GPT-6 Astra supports low, medium, high, xhigh, and max; reasoning cannot be disabled.
  7. For Anthropic, review Claude settings. Thinking can be disabled or set to adaptive when the model advertises that capability. maxTokens is the total output ceiling shared by thinking and the visible answer. Effort is limited to the levels advertised by the selected model; new adaptive presets use 20,000 tokens and high effort when supported.
  8. For a supported xAI reasoning model, review Grok settings. The response format is text and reasoning effort defaults to high. Grok 4.6 offers low, medium, high, and xhigh; the picker only offers levels OAO has verified for the selected model. Tools and the system prompt stay on the agent version, where their schemas and instructions are immutable together.
  9. Select Add model preset.
The dialog lists what is still missing under Validation and keeps the submit button disabled until nothing is. Preset rows are append-only. A published agent version keeps referring to the same reviewed model, routing policy, and generation settings. To change any of them, create a new preset key and publish a new agent version. OpenRouter presets may use live OpenRouter models or saved OpenRouter presets (@preset/<slug>), plus routing controls including fallbacks, zero-data retention, data collection, upstream allow/deny lists, preference order, sort, and price caps. OpenAI, Anthropic, and xAI are direct provider connections, so OpenRouter routing controls must be empty. For direct OpenAI presets, OAO maps the stored controls to the Responses API as text.format, text.verbosity, reasoning.mode, reasoning.effort, and reasoning.summary. Reasoning effort is also the Flue default for the agent; an explicitly configured Harness Operation may still override its own effort. For direct Anthropic presets, OAO maps the stored controls to the Messages API as thinking.type, max_tokens, and output_config.effort. Anthropic counts thinking tokens against max_tokens, so this setting is the hard ceiling for the combined thinking and visible response. Adaptive thinking and the available effort values are model capabilities rather than assumptions made by OAO. Manual thinking.type: "enabled" and budget_tokens are not part of this preset contract; use adaptive thinking for supported current models or disable thinking. For direct xAI presets, OAO stores the exact xai/<catalog-id> returned by the live language-model catalog and calls https://api.x.ai/v1. Newly available text-output Grok IDs can be selected without an OAO catalog release. Known Pi models retain their pinned context/output metadata; a newly discovered ID uses conservative runtime limits until the pinned Pi catalog adds richer metadata. For documented xAI reasoning models, OAO stores textFormat and effort, then maps them to text.format and reasoning.effort on the Responses API. xAI documents high as the default, with low, medium, and high on Grok 4.5 and xhigh from Grok 4.6 onward. Reasoning cannot be disabled on these models. Structured JSON schemas are operation-specific rather than a fixed model default. xAI server-side tools are not part of this preset contract; OAO’s versioned agent tools remain available.

Inspect, duplicate, or archive a preset

Select a preset in the Model presets table to open it. The dialog shows the key, model, provider connection, routing policy or generation settings as JSON, and the agents that currently pin the key.
  • Duplicate as new preset opens the add dialog prefilled with the same provider, model, policy, and settings under the next free key. This is how you “change” a preset: the original stays exactly as published agent versions reviewed it.
  • Archive preset removes it from the table and from the agent editor’s preset picker. Agents that pin the key keep running on it, but must move to another preset before they can publish a new version. The key stays reserved.

Model timeouts and retries

Each model attempt has a five-minute deadline, including connection setup and the entire response stream. Partial output does not reset the clock. OAO aborts an expired request and discards late results, including late tool calls. This deadline applies to all hosted providers and is fixed by the runtime; there is no per-preset timeout setting. The deadline bounds receiving the provider response. Database admission before that request, processing buffered output, and recording run progress are separate work, so elapsed run time can exceed a single attempt’s deadline. A model-start event records the requested turn; it does not prove that the provider has received the request. OAO checks the durable model-turn budget once per model operation, including new retry and recovery operations. Reading additional response chunks does not repeat that database transaction or consume extra turns. The public progress recorder filters unused text, reasoning, and tool-argument deltas before database access; model outcomes, retries, tool results, and recovery events remain recorded. For agent model turns, timeouts and other transient provider failures use Flue’s durable retry budget: up to three retries after the initial attempt, or four attempts total. Retries back off for approximately 2, 4, and 8 seconds with jitter. Provider SDK retries are disabled so they cannot multiply this budget. Completed tools remain committed and are not executed again by a model retry. Cancellation, authentication failures, and the platform model-turn guard do not trigger these retries. Retry attempts still consume the agent’s model-turn budget, and a run or Harness Operation deadline can stop execution sooner. Four consecutive five-minute timeouts can therefore take about 20 minutes plus retry delays. Reduce reasoning effort or requested output size when a model repeatedly exceeds the deadline; retrying the same large generation can time out again. Internal context-compaction requests share the five-minute deadline but do not use the agent-turn retry loop. The session event inspector shows Model call started and Model retry scheduled entries while an agent turn is pending. These events contain safe timing and retry metadata, never provider credentials or raw error bodies. Timed-out attempts may still incur provider charges; without final provider usage, OAO cannot report their complete token usage or cost.

Prompt caching

OAO passes each Flue conversation’s stable session identity to OpenRouter as x-session-id, allowing OpenRouter to keep later turns on the same cache-aware provider endpoint. For direct openrouter/anthropic/* models, OAO also adds Anthropic-compatible ephemeral cache markers to the stable system prompt, the last tool definition, and the latest conversation content. A qualifying first request can therefore report cacheWriteTokens; later requests with the same prefix can report cacheReadTokens in the session usage panel and API. Some providers use implicit prompt caching and do not report writes because creating the cache is free. In that case cacheWriteTokens remains zero even when later calls report cache reads. OpenRouter saved presets (openrouter/@preset/<slug>) can resolve to different model families, so their explicit cache behavior must be configured in the saved preset; OAO still sends session affinity. Provider minimum prompt sizes and cache TTLs continue to apply. Open Agents, create or edit an agent, and choose an available model preset. Publishing validates the preset inside the same tenant-scoped transaction. A project preset is available only when it references a provider connection in that project and platform credential encryption is configured.

HTTP API

All routes are scoped to /v1/projects/{projectId}. Reads require agent:read; writes require project:admin.

Create a provider connection

The response contains credentialConfigured: true, a 12-character credentialFingerprint, and credentialVersion: 1. It never contains apiKey, ciphertext, nonce, or authentication tag.

List providers

Rotate a provider key

Search the matching catalog

The provider ID is required. OAO reads its type from PostgreSQL and returns only the matching provider catalog. For OpenRouter connections, OAO decrypts the project provider key server-side and fetches OpenRouter’s live /models catalog plus saved /presets, exposing saved presets as openrouter/@preset/<slug> options. For direct OpenAI and Anthropic connections, OAO decrypts the selected project key server-side and fetches the provider’s live /v1/models catalog. For xAI, OAO fetches /v1/language-models and exposes text-output entries as xai/<catalog-id> without relying on a hard-coded Grok list. OpenAI uses the live list for account membership and discovers new GPT-4.1, GPT-5-and-later, and o3-and-later family IDs without a pinned catalog update. Known non-Responses models and specialized audio, realtime, transcription, search, image, embedding, moderation, and deep-research variants are excluded. Models without verified runtime capabilities and pricing appear with runtimeSupported: false. The console shows them disabled with an explanation; the CLI setup only offers supported models. Preset creation rejects unsupported entries, so naming patterns cannot authorize execution or produce zero-cost usage. Anthropic responses are intersected with OAO’s pinned runtime models. Anthropic entries also expose the model’s adaptive-thinking, thinking-disable, and effort capabilities. Claude models whose live capabilities do not support effort are not offered, because every OAO Anthropic preset pins that setting. Catalog responses contain public metadata only; no provider key is returned.

Create a preset

For a direct OpenAI connection, use an openai/<catalog-id> model returned by the catalog, send an empty routing object, and include the immutable model settings:
For a direct Anthropic connection, use an anthropic/<catalog-id> model returned by the catalog, send an empty routing object, and include the immutable Claude settings:
For a direct xAI connection, use an xai/<catalog-id> returned by the live language-model catalog. Routing must be empty. Supported reasoning models accept immutable text-format and effort settings:

Tenant and lifecycle guarantees

  • Provider and preset primary keys include organization and project identity.
  • Composite foreign keys prevent cross-project provider references.
  • Row-level security is enabled and forced on both relations.
  • Provider identity and type are immutable; only encrypted credential material may rotate, and its version must increase.
  • Model presets are immutable and cannot be updated or deleted.
  • Existing presets created before provider connections were introduced remain visible as unavailable legacy history. Recreate them with a new versioned key linked to a provider connection.

Troubleshooting

Provider creation says encryption is not configured

Set the same canonical base64 32-byte OAO_CREDENTIAL_ENCRYPTION_KEY for the API and runtime worker, then restart both processes. Do not replace it while ciphertext encrypted with the old key still exists; key-ring rotation is not part of this MVP.

A provider key cannot be decrypted

Confirm the API and runtime use the same platform encryption key and that the database row was not copied between tenants or providers. Tenant/provider identity is authenticated, so copied ciphertext intentionally fails.

A model is not in the catalog

Search /model-catalog with the intended provider connection. OpenRouter, OpenAI, Anthropic, and xAI use separate catalogs and prefixes. If an OpenRouter model or saved preset is missing, confirm the provider API key can read the OpenRouter catalog. If a direct OpenAI model is missing, confirm the selected key can list it through OpenAI’s /v1/models endpoint and belongs to a supported text/reasoning family. Each catalog request fetches the current account list; OAO does not cache it or require a library release to discover new family members. Known OpenAI models retain their runtime metadata. GPT-6 Astra has documented metadata for its 1,050,000-token context window, 128,000-token output limit, image input, pricing, and reasoning levels, so it can be selected immediately when the connected key lists it. Models without verified metadata show unknown context/output limits and remain disabled until runtime support is added. Discovery is automatic; enabling a new model still requires verified endpoint, tool, reasoning, limit, and pricing metadata. This prevents failed runs and incorrect cost reporting. If a direct Anthropic model is missing, confirm the selected key can list it through Anthropic’s /v1/models endpoint. OAO must also have matching native Anthropic runtime metadata. The model catalog controls whether adaptive thinking can be enabled or disabled and which effort levels can be saved. If a Grok model is missing, confirm the selected xAI key can list it through xAI’s /v1/language-models endpoint and that the entry has text output. Image, video, voice, and embedding-only models are intentionally excluded from agent presets.

A preset is unavailable

Confirm it has a non-null provider connection, the connection belongs to the same organization, and the runtime has the platform encryption key. Legacy unbound presets are history only.