Docs
Gateway reference
Base URLs, API keys, identity headers, refusals and limits: the contract your servers call.
What a team's servers call. The gateway speaks each provider's own API: a team changes the base
URL, the API key and a few Alectura-* headers, and everything this page doesn't mention is the
provider's API, passed through unchanged. The management API is specified in
openapi.yaml.
Base URLs
An environment lives in one region, and its requests go to that region's host.
| Region | OpenAI SDKs | Anthropic SDKs | Gemini SDKs |
|---|---|---|---|
| US | https://us.gateway.alecturalabs.com/openai/v1 |
https://us.gateway.alecturalabs.com/anthropic |
https://us.gateway.alecturalabs.com/gemini |
| EU | https://eu.gateway.alecturalabs.com/openai/v1 |
https://eu.gateway.alecturalabs.com/anthropic |
https://eu.gateway.alecturalabs.com/gemini |
| Australia | https://au.gateway.alecturalabs.com/openai/v1 |
https://au.gateway.alecturalabs.com/anthropic |
https://au.gateway.alecturalabs.com/gemini |
import os
from openai import OpenAI
client = OpenAI(
base_url="https://eu.gateway.alecturalabs.com/openai/v1",
api_key=os.environ["ALECTURA_API_KEY"],
)
response = client.responses.create(
model="gpt-5-mini",
input="Say hello in French.",
extra_headers={
"Alectura-Organization": "globex",
"Alectura-User": "user_42",
"Alectura-Plan": "enterprise",
},
)
print(response.output_text)
import os
from anthropic import Anthropic
client = Anthropic(
base_url="https://eu.gateway.alecturalabs.com/anthropic",
api_key=os.environ["ALECTURA_API_KEY"],
)
message = client.messages.create(
model="claude-haiku-4-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Say hello in French."}],
extra_headers={"Alectura-Organization": "globex", "Alectura-User": "user_42"},
)
print(message.content[0].text)
import os
from google import genai
client = genai.Client(
api_key=os.environ["ALECTURA_API_KEY"],
http_options={
"base_url": "https://eu.gateway.alecturalabs.com/gemini",
"headers": {"Alectura-Organization": "globex", "Alectura-User": "user_42"},
},
)
response = client.models.generate_content(
model="gemini-3.5-flash",
contents="Say hello in French.",
)
print(response.text)
API keys
The gateway takes a sending key. Management keys are for the management API, and the gateway refuses them.
| Key | Format |
|---|---|
| Sending | sk-alectura-{region}-{mode}-{secret} |
| Management | sk-alectura-mgmt-{region}-{mode}-{secret} |
{region}isus,euorau.{mode}islivefor an environment marked as production andtestfor any other.{secret}is 40 random characters and then a 6-character checksum, all from[0-9A-Za-z]. The checksum is the CRC32 of everything before it, in base 62 and padded with zeros, so a scanner can tell a real key from a typo without calling us. Keys matchsk-alectura-(?:mgmt-)?(?:us|eu|au)-(?:live|test)-[0-9A-Za-z]{46}, which is registered with GitHub secret scanning.- The key goes in whichever header the SDK already sends:
Authorization: Bearer …,x-api-key: …orx-goog-api-key: …, on any path. A request carrying two different keys is refused. Gemini's?key=query parameter isn't read, so a key never sits in a URL that a proxy or log might keep; Google's SDKs send the header. - A key for another region is refused with
wrong_region, and the message names the right host.
Identity
| Header | Carries |
|---|---|
Alectura-Organization |
The team's own identifier for the organization the request is for |
Alectura-User |
The team's own identifier for the user |
Alectura-Plan |
The plan the organization is on |
Alectura-Conversation |
Optional. Requests that share it form a conversation; without it, conversations are linked by their blocks. |
- Each value is 1 to 256 characters of printable ASCII, and case matters. A missing header is null, and policies can match null.
- Only these headers set identity. Any other
Alectura-*header is refused asinvalid_request, so a typo can't quietly make a request anonymous. - The provider fields that carry identity never set it: OpenAI's
user,safety_identifierandprompt_cache_key, and Anthropic'smetadata.user_id. A value a client sends in one of them goes upstream as its keyed hash, so equal values stay equal and prompt caching still works. - The gateway always sends the provider a safety identifier (
safety_identifierfor OpenAI,metadata.user_idfor Anthropic and Bedrock): a keyed hash of the user, or of the organization when there's no user. A provider that acts on abuse then acts on one of the team's users, not the team's whole account. The hash key is the environment's, so the same user has different identifiers in different environments. Gemini's API has no such field, so its requests carry none.
Endpoints
| Method and path | Format | Served by |
|---|---|---|
POST /openai/v1/chat/completions |
OpenAI Chat Completions | OpenAI connections |
POST /openai/v1/responses |
OpenAI Responses | OpenAI connections |
POST /anthropic/v1/messages |
Anthropic Messages | Anthropic connections, or Bedrock running Claude |
POST /gemini/{version}/models/{model}:generateContent |
Gemini | Gemini connections |
POST /gemini/{version}/models/{model}:streamGenerateContent?alt=sse |
Gemini, streamed | Gemini connections |
GET /openai/v1/models |
OpenAI's model list | |
GET /anthropic/v1/models |
Anthropic's model list | |
GET /gemini/{version}/models |
Gemini's model list |
Anything else under /openai/, /anthropic/ or /gemini/ is refused as unsupported_endpoint:
stored responses (GET /openai/v1/responses/{id}), files, batches, realtime, embeddings, cached
contents and token counting.
Gemini
Gemini's format is the Gemini API's own (generativelanguage.googleapis.com), which Google's SDKs
speak. The paths are Google's under /gemini, as LiteLLM's Gemini pass-through has them;
OpenRouter serves Gemini models only through OpenAI's format, which would translate one format
into another.
{version}isv1beta, the SDKs' default, orv1, and goes upstream unchanged.- The model is in the path, not the body, and a request is streamed when its path says
streamGenerateContent. A stream is served only as server-sent events, withalt=sse, which Google's SDKs always send; without it, Google streams one JSON array, which is refused asinvalid_request. - Its parts are read as the other formats' blocks are:
textis text (a part withthought: trueis reasoning),inlineDataan attachment,fileDataa pointer to one,functionCalla tool call andfunctionResponsea tool result.systemInstructionis the system prompt. - A
thoughtSignaturealways passes through byte for byte, and so does a function call or thought that carries one, since Google checks those. Google also puts one on answer text, and there it doesn't check the text, so that text can still be edited. - The most output is
generationConfig.maxOutputTokens.
Models
The model lists hold the models this request's policies allow that an active connection of the
path's format serves, so they take the identity headers too. Ids are the providers' own; the
gateway never renames a model. For OpenAI's list, created is when the first connection serving the
model was created. For Anthropic's, display_name is the id and the list is one page. The
simulator serves every format, so with it connected each list holds every provider's models.
A policy names models by id. An id without a date matches every dated version of it, and Bedrock's
ids match by the model they name, so claude-sonnet-4-5 matches claude-sonnet-4-5-20250929 and
eu.anthropic.claude-sonnet-4-5-20250929-v1:0. A dated id matches only itself.
Routing
A request goes to a connection of its own format family: OpenAI's formats to OpenAI, at its EU or
global endpoint; Anthropic's to Anthropic, or to Bedrock running Claude, whose envelope route
translates in both directions, errors included; Gemini's to the Gemini API, whose one endpoint
promises no region, so its connections are global. No format is ever translated into another.
Which connections are candidates, and in what order, is set by the policy's ordered lists
routing.connections and routing.regions. A failure before the provider's first byte moves the
request to the next candidate, for up to three attempts. After the first byte, there's no fallback.
When a provider was reached, the response carries Alectura-Connection (the connection's id) and
Alectura-Region (the region that served the request).
Streaming
- A streamed response is the provider's own server-sent events, held until the provider's first event arrives. A failure before then is an ordinary HTTP error with its real status, and can fall back.
secretsandpersonal_dataread the answer when theirpartsselectoutput. To redact a value a stream splits across deltas, they hold back the word still being written, and before it the words that could begin a value of several words (a number or a capital), so text reaches the client a few words behind the provider. Monitor mode holds back the same text.- While a guardrail holds back text, the gateway sends keep-alives:
: keep-alivecomment lines on the OpenAI formats andpingevents on Anthropic's, as the providers do. Gemini's stream has no keep-alive of its own, and Google's SDKs parse every line as JSON, so there it's a chunk whose one part has empty text:data: {"candidates":[{"content":{"role":"model","parts":[{"text":""}]},"index":0}]}. - So a
200can still end in an error. A cut, or a provider failure after the first byte, ends the stream with the format's own terminal error:
# Chat Completions: a chunk carrying the error and an empty choice, then [DONE]
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","created":1790680024,"model":"gpt-5",
"choices":[{"index":0,"delta":{"content":""},"finish_reason":"error"}],
"error":{"message":"Stopped by the prompt_injection guardrail.","type":"alectura_policy",
"param":null,"code":"cut_by_guardrail","request_id":"req_01M3PQJ64T…"}}
data: [DONE]
# Responses: response.failed, numbered one past the last event sent
event: response.failed
data: {"type":"response.failed","sequence_number":42,
"response":{"id":"resp_…","object":"response","status":"failed",
"error":{"code":"cut_by_guardrail",
"message":"Stopped by the prompt_injection guardrail."}, …}}
# Anthropic Messages
event: error
data: {"type":"error","error":{"type":"permission_error","code":"cut_by_guardrail",
"message":"Stopped by the prompt_injection guardrail."},"request_id":"req_01M3PQJ64T…"}
# Gemini: the error object Google sends when a stream fails, which its SDKs raise
data: {"error":{"code":403,"message":"Stopped by the prompt_injection guardrail.",
"status":"PERMISSION_DENIED","details":[…]}}
The JSON is wrapped here to fit; on the wire each data: is one line. A non-streamed response
that a guardrail stops is a 403 refusal with the code cut_by_guardrail, since nothing has been
sent yet.
Refusals
A refusal comes back in the format's own error shape, so the team's error handling already works,
and the provider never sees the request. Every error body carries the request id, as does the
Alectura-Request-Id header on every response.
{
"error": {
"message": "Requests for this user are suspended.",
"type": "alectura_policy",
"param": null,
"code": "user_suspended",
"request_id": "req_01M3PQJ64T…"
}
}
{
"type": "error",
"error": {
"type": "permission_error",
"code": "user_suspended",
"message": "Requests for this user are suspended."
},
"request_id": "req_01M3PQJ64T…"
}
{
"error": {
"code": 403,
"message": "Requests for this user are suspended.",
"status": "PERMISSION_DENIED",
"details": [
{
"@type": "type.googleapis.com/google.rpc.ErrorInfo",
"reason": "user_suspended",
"domain": "alecturalabs.com",
"metadata": {}
},
{
"@type": "type.googleapis.com/google.rpc.RequestInfo",
"requestId": "req_01M3PQJ64T…"
}
]
}
}
On the OpenAI formats type is always alectura_policy, a bad key's too: OpenAI's SDKs raise
their authentication error from the 401 alone, and code says which refusal it is. On Anthropic's, type is Anthropic's own for the status, and code is ours.
On Gemini's, code is the status and status Google's own name for it, as Google's errors have
them; our code is the reason of an ErrorInfo in the domain alecturalabs.com, and the request
id is a RequestInfo, the places Google's own APIs put them.
Each code has one status, and only a code that waiting fixes gets a status that SDKs retry on their
own (429 and 5xx).
| Status | Anthropic type |
Gemini status |
Code | When |
|---|---|---|---|---|
| 400 | invalid_request_error |
INVALID_ARGUMENT |
invalid_request |
The body can't be parsed, so the guardrails can't run; or an Alectura-* header is unknown or invalid. |
| 400 | invalid_request_error |
INVALID_ARGUMENT |
conflicting_api_keys |
The request carries two different keys. |
| 401 | authentication_error |
UNAUTHENTICATED |
invalid_api_key |
No key, or an unknown, expired or revoked one. |
| 402 | billing_error |
FAILED_PRECONDITION |
budget_exhausted |
A limit over a day, week or month is spent. The error names the limit and when it resets. |
| 403 | permission_error |
PERMISSION_DENIED |
permission_denied |
The key can't send, such as a management key. |
| 403 | permission_error |
PERMISSION_DENIED |
wrong_region |
The key belongs to another region. The message names its host. |
| 403 | permission_error |
PERMISSION_DENIED |
user_suspended, organization_suspended |
A suspension. |
| 403 | permission_error |
PERMISSION_DENIED |
model_not_allowed, region_not_allowed, no_allowed_connection |
The routing settings leave nowhere to send the request; or a spend limit applies and the model has no list price, so its cost can't be counted. |
| 403 | permission_error |
PERMISSION_DENIED |
blocked_by_guardrail, attachment_uninspectable, cut_by_guardrail |
A guardrail blocks the request, or the policy blocks content nothing can inspect, or a guardrail stopped a non-streamed response. |
| 404 | not_found_error |
NOT_FOUND |
unsupported_endpoint |
A provider path the gateway doesn't serve. |
| 413 | request_too_large |
INVALID_ARGUMENT |
request_too_large |
The body is over the provider's cap: 32 MB for Anthropic, 72 MiB for OpenAI, 100 MB for Gemini. |
| 429 | rate_limit_error |
RESOURCE_EXHAUSTED |
rate_limited |
A per-minute limit, or a trial's on the simulator, with Retry-After in seconds. |
| 502 | api_error |
UNAVAILABLE |
provider_unreachable |
No candidate could be reached. |
| 503 | api_error |
UNAVAILABLE |
limiter_unavailable |
Spend so far isn't known, so a spend limit can't be checked. |
| 504 | timeout_error |
DEADLINE_EXCEEDED |
provider_timeout |
No candidate answered in time. |
A spent budget is FAILED_PRECONDITION rather than Google's RESOURCE_EXHAUSTED, which Google
uses for its own rate limits and its SDKs retry.
A refusal over a limit adds a limit object beside code, or on Gemini's format the same fields
as limit_kind, limit_per and so on in the ErrorInfo's metadata, which holds only strings:
"limit": {
"kind": "spend",
"per": "user",
"period": "day",
"amount": "2.00",
"used": "1.9962",
"requested": "0.0079",
"resets_at": "2026-09-30T00:00:00Z"
}
used is what the period has used so far, and requested what the refused request would have
used at most: its estimated cost, its input and maximum output tokens, or one request. A request
is refused when the two together would pass amount, so it can be refused before used reaches
it.
Provider errors
An error the provider returns passes through unchanged: its status, its body and its headers. Two exceptions:
- Anything in it that exposes a connection is masked with
***: an endpoint URL, an AWS account or role, the provider's organization or project id, or a credential the request carried, such as the session token a Bedrock signature error quotes. - Bedrock's errors are rewritten into Anthropic's shape, keeping the status, since the client speaks Anthropic's format.
When every attempt fails, the client gets the last attempt's failure: the provider's error, or one of ours if the provider never answered.
Headers
| Header | Direction | What happens |
|---|---|---|
Authorization, x-api-key, x-goog-api-key |
in | The Alectura key; never sent upstream |
Alectura-* |
in | As above; never sent upstream |
anthropic-version, anthropic-beta |
in | Passed to the provider |
OpenAI-Organization, OpenAI-Project |
in | Dropped: the connection names the provider account |
| Other request headers | in | Passed to the provider, except hop-by-hop headers and those a proxy adds: Forwarded, X-Forwarded-*, Via, X-Real-IP and X-Amzn-Trace-Id |
Alectura-Request-Id |
out | On every response, refusals included |
Alectura-Connection, Alectura-Region |
out | When a provider was reached |
Retry-After |
out | On rate_limited |
| The provider's response headers | out | Passed through, rate-limit headers included, except those naming the provider account |
Limits
limits.max_output_tokenslowers a request's maximum output to the limit, and sets it when the request sets none.- Each kind of limit takes its own periods:
requestsandtokensaminuteor aday, andspendaday, aweekor amonth. So there's no weekly or monthly token limit, and no spend limit per minute. - A limit over a minute refuses with
429andRetry-After. A limit over a day, week or month refuses with402, which SDKs don't retry, until it resets at the start of the next UTC period; weeks start on Monday. - A request reserves its estimated tokens and cost when it's admitted, at the connection's list prices, and settles them from the provider's usage when it ends. Events fire when a limit passes 80% and 95% of its amount, and when it starts refusing.
Compatibility
- A new field or event in a provider's API passes through without a gateway release.
- Once its key is known, a request is recorded whatever happens to it, refusals included; see
GET /v1/requestsin the management API.