Security model
Provider API keys are write-only:- the API encrypts each key with AES-256-GCM before inserting it;
- the authenticated encryption context includes organization ID, provider ID, provider type, and credential version;
- PostgreSQL stores ciphertext, a unique nonce, authentication tag, key version, and SHA-256 fingerprint;
- list and create responses expose only the first 12 fingerprint characters;
- the runtime decrypts a key only while activating the selected project preset;
- API keys never appear in model presets, agent snapshots, SSE, audit detail, logs, or traces.
OPENROUTER_API_KEY, OPENAI_API_KEY, ANTHROPIC_API_KEY, or
XAI_API_KEY in .env. Those keys are entered through the project UI or API
and stored as tenant-scoped ciphertext.
Without OAO_CREDENTIAL_ENCRYPTION_KEY, provider creation, preset publication,
and runtime execution are unavailable. OAO does not fall back to a local fake
model.
Add a provider in the console
- Open Models.
- Select Add provider.
- Choose OpenRouter, OpenAI, Anthropic, or xAI (Grok).
- Enter a stable project-local key, display name, and API key.
- Select Add provider.
Add a model preset
- Select Add model preset.
- Choose a provider connection.
- In Model, type to search the live catalog available to that connection
and pick a model or OpenRouter saved preset. OAO uses each connection’s
authenticated live catalog. xAI connections use
GET /v1/language-modelsso image-generation, video, voice, and embedding-only IDs are not offered. The picker shows the first matches with their model identifier, context window, and reasoning support, and reports how many further matches the search is hiding. - Check the suggested preset key and display name. Both are derived from the
chosen model — for example
claude-sonnet-4-6-v1— with the version suffix moved forward past keys this project already uses. Editing either field stops the suggestion from overwriting it. - For OpenRouter, open Routing and data policy to set routing and data-handling constraints. It stays folded away while it is unset, and the summary next to it always shows the policy that will be stored.
- For OpenAI, review Model settings. New presets default to plain text, standard reasoning mode, medium reasoning effort, medium verbosity, and an automatic reasoning summary. GPT-5.6 presets can select standard or pro mode and reasoning effort from none through max. GPT-6 Astra supports low, medium, high, xhigh, and max; reasoning cannot be disabled.
- For Anthropic, review Claude settings. Thinking can be disabled or set
to adaptive when the model advertises that capability.
maxTokensis the total output ceiling shared by thinking and the visible answer. Effort is limited to the levels advertised by the selected model; new adaptive presets use 20,000 tokens and high effort when supported. - For a supported xAI reasoning model, review Grok settings. The response format is text and reasoning effort defaults to high. Grok 4.6 offers low, medium, high, and xhigh; the picker only offers levels OAO has verified for the selected model. Tools and the system prompt stay on the agent version, where their schemas and instructions are immutable together.
- Select Add model preset.
@preset/<slug>), plus routing controls including fallbacks, zero-data
retention, data collection, upstream allow/deny lists, preference order, sort,
and price caps. OpenAI, Anthropic, and xAI are direct provider connections, so
OpenRouter routing controls must be empty.
For direct OpenAI presets, OAO maps the stored controls to the Responses API as
text.format, text.verbosity, reasoning.mode, reasoning.effort, and
reasoning.summary. Reasoning effort is also the Flue default for the agent;
an explicitly configured Harness Operation may still override its own effort.
For direct Anthropic presets, OAO maps the stored controls to the Messages API
as thinking.type, max_tokens, and output_config.effort. Anthropic counts
thinking tokens against max_tokens, so this setting is the hard ceiling for
the combined thinking and visible response. Adaptive thinking and the available
effort values are model capabilities rather than assumptions made by OAO.
Manual thinking.type: "enabled" and budget_tokens are not part of this
preset contract; use adaptive thinking for supported current models or disable
thinking.
For direct xAI presets, OAO stores the exact xai/<catalog-id> returned by the
live language-model catalog and calls https://api.x.ai/v1. Newly available
text-output Grok IDs can be selected without an OAO catalog release. Known Pi
models retain their pinned context/output metadata; a newly discovered ID uses
conservative runtime limits until the pinned Pi catalog adds richer metadata.
For documented xAI reasoning models, OAO stores textFormat and effort, then
maps them to text.format and reasoning.effort on the Responses API. xAI
documents high as the default, with low, medium, and high on Grok 4.5 and xhigh
from Grok 4.6 onward. Reasoning cannot be disabled on these models. Structured
JSON schemas are operation-specific rather than a fixed model default. xAI
server-side tools are not part of this preset contract; OAO’s versioned agent
tools remain available.
Inspect, duplicate, or archive a preset
Select a preset in the Model presets table to open it. The dialog shows the key, model, provider connection, routing policy or generation settings as JSON, and the agents that currently pin the key.- Duplicate as new preset opens the add dialog prefilled with the same provider, model, policy, and settings under the next free key. This is how you “change” a preset: the original stays exactly as published agent versions reviewed it.
- Archive preset removes it from the table and from the agent editor’s preset picker. Agents that pin the key keep running on it, but must move to another preset before they can publish a new version. The key stays reserved.
Model timeouts and retries
Each model attempt has a five-minute deadline, including connection setup and the entire response stream. Partial output does not reset the clock. OAO aborts an expired request and discards late results, including late tool calls. This deadline applies to all hosted providers and is fixed by the runtime; there is no per-preset timeout setting. The deadline bounds receiving the provider response. Database admission before that request, processing buffered output, and recording run progress are separate work, so elapsed run time can exceed a single attempt’s deadline. A model-start event records the requested turn; it does not prove that the provider has received the request. OAO checks the durable model-turn budget once per model operation, including new retry and recovery operations. Reading additional response chunks does not repeat that database transaction or consume extra turns. The public progress recorder filters unused text, reasoning, and tool-argument deltas before database access; model outcomes, retries, tool results, and recovery events remain recorded. For agent model turns, timeouts and other transient provider failures use Flue’s durable retry budget: up to three retries after the initial attempt, or four attempts total. Retries back off for approximately 2, 4, and 8 seconds with jitter. Provider SDK retries are disabled so they cannot multiply this budget. Completed tools remain committed and are not executed again by a model retry. Cancellation, authentication failures, and the platform model-turn guard do not trigger these retries. Retry attempts still consume the agent’s model-turn budget, and a run or Harness Operation deadline can stop execution sooner. Four consecutive five-minute timeouts can therefore take about 20 minutes plus retry delays. Reduce reasoning effort or requested output size when a model repeatedly exceeds the deadline; retrying the same large generation can time out again. Internal context-compaction requests share the five-minute deadline but do not use the agent-turn retry loop. The session event inspector shows Model call started and Model retry scheduled entries while an agent turn is pending. These events contain safe timing and retry metadata, never provider credentials or raw error bodies. Timed-out attempts may still incur provider charges; without final provider usage, OAO cannot report their complete token usage or cost.Prompt caching
OAO passes each Flue conversation’s stable session identity to OpenRouter asx-session-id, allowing OpenRouter to keep later turns on the same cache-aware
provider endpoint. For direct openrouter/anthropic/* models, OAO also adds
Anthropic-compatible ephemeral cache markers to the stable system prompt, the
last tool definition, and the latest conversation content. A qualifying first
request can therefore report cacheWriteTokens; later requests with the same
prefix can report cacheReadTokens in the session usage panel and API.
Some providers use implicit prompt caching and do not report writes because
creating the cache is free. In that case cacheWriteTokens remains zero even
when later calls report cache reads. OpenRouter saved presets
(openrouter/@preset/<slug>) can resolve to different model families, so their
explicit cache behavior must be configured in the saved preset; OAO still
sends session affinity. Provider minimum prompt sizes and cache TTLs continue
to apply.
Link a model to an agent
Open Agents, create or edit an agent, and choose an available model preset. Publishing validates the preset inside the same tenant-scoped transaction. A project preset is available only when it references a provider connection in that project and platform credential encryption is configured.HTTP API
All routes are scoped to/v1/projects/{projectId}. Reads require
agent:read; writes require project:admin.
Create a provider connection
credentialConfigured: true, a 12-character
credentialFingerprint, and credentialVersion: 1. It never contains
apiKey, ciphertext, nonce, or authentication tag.
List providers
Rotate a provider key
Search the matching catalog
/models catalog
plus saved /presets, exposing saved presets as openrouter/@preset/<slug>
options. For direct OpenAI and Anthropic connections, OAO decrypts the selected
project key server-side and fetches the provider’s live /v1/models catalog.
For xAI, OAO fetches /v1/language-models and exposes text-output entries as
xai/<catalog-id> without relying on a hard-coded Grok list. OpenAI uses the
live list for account membership and discovers new GPT-4.1, GPT-5-and-later, and
o3-and-later family IDs without a pinned catalog update. Known
non-Responses models and specialized audio, realtime, transcription, search,
image, embedding, moderation, and deep-research variants are excluded.
Models without verified runtime capabilities and pricing appear with
runtimeSupported: false. The console shows them disabled with an explanation;
the CLI setup only offers supported models. Preset creation rejects unsupported
entries, so naming patterns cannot authorize execution or produce zero-cost usage.
Anthropic responses are intersected with OAO’s pinned runtime models.
Anthropic entries also expose
the model’s adaptive-thinking, thinking-disable, and effort capabilities.
Claude models whose live capabilities do not support effort are not offered,
because every OAO Anthropic preset pins that setting.
Catalog responses contain public metadata only; no provider key is returned.
Create a preset
openai/<catalog-id> model returned by
the catalog, send an empty routing object, and include the immutable model
settings:
anthropic/<catalog-id> model
returned by the catalog, send an empty routing object, and include the
immutable Claude settings:
xai/<catalog-id> returned by the live
language-model catalog. Routing must be empty. Supported reasoning models
accept immutable text-format and effort settings:
Tenant and lifecycle guarantees
- Provider and preset primary keys include organization and project identity.
- Composite foreign keys prevent cross-project provider references.
- Row-level security is enabled and forced on both relations.
- Provider identity and type are immutable; only encrypted credential material may rotate, and its version must increase.
- Model presets are immutable and cannot be updated or deleted.
- Existing presets created before provider connections were introduced remain visible as unavailable legacy history. Recreate them with a new versioned key linked to a provider connection.
Troubleshooting
Provider creation says encryption is not configured
Set the same canonical base64 32-byteOAO_CREDENTIAL_ENCRYPTION_KEY for the
API and runtime worker, then restart both processes. Do not replace it while
ciphertext encrypted with the old key still exists; key-ring rotation is not
part of this MVP.
A provider key cannot be decrypted
Confirm the API and runtime use the same platform encryption key and that the database row was not copied between tenants or providers. Tenant/provider identity is authenticated, so copied ciphertext intentionally fails.A model is not in the catalog
Search/model-catalog with the intended provider connection. OpenRouter,
OpenAI, Anthropic, and xAI use separate catalogs and prefixes. If an OpenRouter model
or saved preset is missing, confirm the provider API key can read the OpenRouter
catalog. If a direct OpenAI model is missing, confirm the selected key can list
it through OpenAI’s /v1/models endpoint and belongs to a supported text/reasoning
family. Each catalog request fetches the current account list; OAO does not
cache it or require a library release to discover new family members.
Known OpenAI models retain their runtime metadata. GPT-6 Astra has documented
metadata for its 1,050,000-token context window, 128,000-token output limit,
image input, pricing, and reasoning levels, so it can be selected immediately
when the connected key lists it. Models without verified metadata show unknown
context/output limits and remain disabled until runtime support is added.
Discovery is automatic; enabling a new model still requires verified endpoint,
tool, reasoning, limit, and pricing metadata. This prevents failed runs and
incorrect cost reporting.
If a direct Anthropic model is missing, confirm the selected key can list it
through Anthropic’s /v1/models endpoint. OAO must also have matching native
Anthropic runtime metadata. The model catalog controls whether adaptive
thinking can be enabled or disabled and which effort levels can be saved.
If a Grok model is missing, confirm the selected xAI key can list it through
xAI’s /v1/language-models endpoint and that the entry has text output. Image,
video, voice, and embedding-only models are intentionally excluded from agent
presets.

