Docs

Gateway reference

Base URLs, API keys, identity headers, refusals and limits: the contract your servers call.

What a team's servers call. The gateway speaks each provider's own API: a team changes the base URL, the API key and a few Alectura-* headers, and everything this page doesn't mention is the provider's API, passed through unchanged. The management API is specified in openapi.yaml.

Base URLs

An environment lives in one region, and its requests go to that region's host.

Region OpenAI SDKs Anthropic SDKs Gemini SDKs
US https://us.gateway.alecturalabs.com/openai/v1 https://us.gateway.alecturalabs.com/anthropic https://us.gateway.alecturalabs.com/gemini
EU https://eu.gateway.alecturalabs.com/openai/v1 https://eu.gateway.alecturalabs.com/anthropic https://eu.gateway.alecturalabs.com/gemini
Australia https://au.gateway.alecturalabs.com/openai/v1 https://au.gateway.alecturalabs.com/anthropic https://au.gateway.alecturalabs.com/gemini
import os

from openai import OpenAI

client = OpenAI(
    base_url="https://eu.gateway.alecturalabs.com/openai/v1",
    api_key=os.environ["ALECTURA_API_KEY"],
)
response = client.responses.create(
    model="gpt-5-mini",
    input="Say hello in French.",
    extra_headers={
        "Alectura-Organization": "globex",
        "Alectura-User": "user_42",
        "Alectura-Plan": "enterprise",
    },
)
print(response.output_text)
import os

from anthropic import Anthropic

client = Anthropic(
    base_url="https://eu.gateway.alecturalabs.com/anthropic",
    api_key=os.environ["ALECTURA_API_KEY"],
)
message = client.messages.create(
    model="claude-haiku-4-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Say hello in French."}],
    extra_headers={"Alectura-Organization": "globex", "Alectura-User": "user_42"},
)
print(message.content[0].text)
import os

from google import genai

client = genai.Client(
    api_key=os.environ["ALECTURA_API_KEY"],
    http_options={
        "base_url": "https://eu.gateway.alecturalabs.com/gemini",
        "headers": {"Alectura-Organization": "globex", "Alectura-User": "user_42"},
    },
)
response = client.models.generate_content(
    model="gemini-3.5-flash",
    contents="Say hello in French.",
)
print(response.text)

API keys

The gateway takes a sending key. Management keys are for the management API, and the gateway refuses them.

Key Format
Sending sk-alectura-{region}-{mode}-{secret}
Management sk-alectura-mgmt-{region}-{mode}-{secret}
  • {region} is us, eu or au. {mode} is live for an environment marked as production and test for any other.
  • {secret} is 40 random characters and then a 6-character checksum, all from [0-9A-Za-z]. The checksum is the CRC32 of everything before it, in base 62 and padded with zeros, so a scanner can tell a real key from a typo without calling us. Keys match sk-alectura-(?:mgmt-)?(?:us|eu|au)-(?:live|test)-[0-9A-Za-z]{46}, which is registered with GitHub secret scanning.
  • The key goes in whichever header the SDK already sends: Authorization: Bearer …, x-api-key: … or x-goog-api-key: …, on any path. A request carrying two different keys is refused. Gemini's ?key= query parameter isn't read, so a key never sits in a URL that a proxy or log might keep; Google's SDKs send the header.
  • A key for another region is refused with wrong_region, and the message names the right host.

Identity

Header Carries
Alectura-Organization The team's own identifier for the organization the request is for
Alectura-User The team's own identifier for the user
Alectura-Plan The plan the organization is on
Alectura-Conversation Optional. Requests that share it form a conversation; without it, conversations are linked by their blocks.
  • Each value is 1 to 256 characters of printable ASCII, and case matters. A missing header is null, and policies can match null.
  • Only these headers set identity. Any other Alectura-* header is refused as invalid_request, so a typo can't quietly make a request anonymous.
  • The provider fields that carry identity never set it: OpenAI's user, safety_identifier and prompt_cache_key, and Anthropic's metadata.user_id. A value a client sends in one of them goes upstream as its keyed hash, so equal values stay equal and prompt caching still works.
  • The gateway always sends the provider a safety identifier (safety_identifier for OpenAI, metadata.user_id for Anthropic and Bedrock): a keyed hash of the user, or of the organization when there's no user. A provider that acts on abuse then acts on one of the team's users, not the team's whole account. The hash key is the environment's, so the same user has different identifiers in different environments. Gemini's API has no such field, so its requests carry none.

Endpoints

Method and path Format Served by
POST /openai/v1/chat/completions OpenAI Chat Completions OpenAI connections
POST /openai/v1/responses OpenAI Responses OpenAI connections
POST /anthropic/v1/messages Anthropic Messages Anthropic connections, or Bedrock running Claude
POST /gemini/{version}/models/{model}:generateContent Gemini Gemini connections
POST /gemini/{version}/models/{model}:streamGenerateContent?alt=sse Gemini, streamed Gemini connections
GET /openai/v1/models OpenAI's model list
GET /anthropic/v1/models Anthropic's model list
GET /gemini/{version}/models Gemini's model list

Anything else under /openai/, /anthropic/ or /gemini/ is refused as unsupported_endpoint: stored responses (GET /openai/v1/responses/{id}), files, batches, realtime, embeddings, cached contents and token counting.

Gemini

Gemini's format is the Gemini API's own (generativelanguage.googleapis.com), which Google's SDKs speak. The paths are Google's under /gemini, as LiteLLM's Gemini pass-through has them; OpenRouter serves Gemini models only through OpenAI's format, which would translate one format into another.

  • {version} is v1beta, the SDKs' default, or v1, and goes upstream unchanged.
  • The model is in the path, not the body, and a request is streamed when its path says streamGenerateContent. A stream is served only as server-sent events, with alt=sse, which Google's SDKs always send; without it, Google streams one JSON array, which is refused as invalid_request.
  • Its parts are read as the other formats' blocks are: text is text (a part with thought: true is reasoning), inlineData an attachment, fileData a pointer to one, functionCall a tool call and functionResponse a tool result. systemInstruction is the system prompt.
  • A thoughtSignature always passes through byte for byte, and so does a function call or thought that carries one, since Google checks those. Google also puts one on answer text, and there it doesn't check the text, so that text can still be edited.
  • The most output is generationConfig.maxOutputTokens.

Models

The model lists hold the models this request's policies allow that an active connection of the path's format serves, so they take the identity headers too. Ids are the providers' own; the gateway never renames a model. For OpenAI's list, created is when the first connection serving the model was created. For Anthropic's, display_name is the id and the list is one page. The simulator serves every format, so with it connected each list holds every provider's models.

A policy names models by id. An id without a date matches every dated version of it, and Bedrock's ids match by the model they name, so claude-sonnet-4-5 matches claude-sonnet-4-5-20250929 and eu.anthropic.claude-sonnet-4-5-20250929-v1:0. A dated id matches only itself.

Routing

A request goes to a connection of its own format family: OpenAI's formats to OpenAI, at its EU or global endpoint; Anthropic's to Anthropic, or to Bedrock running Claude, whose envelope route translates in both directions, errors included; Gemini's to the Gemini API, whose one endpoint promises no region, so its connections are global. No format is ever translated into another.

Which connections are candidates, and in what order, is set by the policy's ordered lists routing.connections and routing.regions. A failure before the provider's first byte moves the request to the next candidate, for up to three attempts. After the first byte, there's no fallback.

When a provider was reached, the response carries Alectura-Connection (the connection's id) and Alectura-Region (the region that served the request).

Streaming

  • A streamed response is the provider's own server-sent events, held until the provider's first event arrives. A failure before then is an ordinary HTTP error with its real status, and can fall back.
  • secrets and personal_data read the answer when their parts select output. To redact a value a stream splits across deltas, they hold back the word still being written, and before it the words that could begin a value of several words (a number or a capital), so text reaches the client a few words behind the provider. Monitor mode holds back the same text.
  • While a guardrail holds back text, the gateway sends keep-alives: : keep-alive comment lines on the OpenAI formats and ping events on Anthropic's, as the providers do. Gemini's stream has no keep-alive of its own, and Google's SDKs parse every line as JSON, so there it's a chunk whose one part has empty text: data: {"candidates":[{"content":{"role":"model","parts":[{"text":""}]},"index":0}]}.
  • So a 200 can still end in an error. A cut, or a provider failure after the first byte, ends the stream with the format's own terminal error:
# Chat Completions: a chunk carrying the error and an empty choice, then [DONE]
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","created":1790680024,"model":"gpt-5",
       "choices":[{"index":0,"delta":{"content":""},"finish_reason":"error"}],
       "error":{"message":"Stopped by the prompt_injection guardrail.","type":"alectura_policy",
                "param":null,"code":"cut_by_guardrail","request_id":"req_01M3PQJ64T…"}}

data: [DONE]

# Responses: response.failed, numbered one past the last event sent
event: response.failed
data: {"type":"response.failed","sequence_number":42,
       "response":{"id":"resp_…","object":"response","status":"failed",
                   "error":{"code":"cut_by_guardrail",
                            "message":"Stopped by the prompt_injection guardrail."}, …}}

# Anthropic Messages
event: error
data: {"type":"error","error":{"type":"permission_error","code":"cut_by_guardrail",
       "message":"Stopped by the prompt_injection guardrail."},"request_id":"req_01M3PQJ64T…"}

# Gemini: the error object Google sends when a stream fails, which its SDKs raise
data: {"error":{"code":403,"message":"Stopped by the prompt_injection guardrail.",
       "status":"PERMISSION_DENIED","details":[…]}}

The JSON is wrapped here to fit; on the wire each data: is one line. A non-streamed response that a guardrail stops is a 403 refusal with the code cut_by_guardrail, since nothing has been sent yet.

Refusals

A refusal comes back in the format's own error shape, so the team's error handling already works, and the provider never sees the request. Every error body carries the request id, as does the Alectura-Request-Id header on every response.

{
  "error": {
    "message": "Requests for this user are suspended.",
    "type": "alectura_policy",
    "param": null,
    "code": "user_suspended",
    "request_id": "req_01M3PQJ64T…"
  }
}
{
  "type": "error",
  "error": {
    "type": "permission_error",
    "code": "user_suspended",
    "message": "Requests for this user are suspended."
  },
  "request_id": "req_01M3PQJ64T…"
}
{
  "error": {
    "code": 403,
    "message": "Requests for this user are suspended.",
    "status": "PERMISSION_DENIED",
    "details": [
      {
        "@type": "type.googleapis.com/google.rpc.ErrorInfo",
        "reason": "user_suspended",
        "domain": "alecturalabs.com",
        "metadata": {}
      },
      {
        "@type": "type.googleapis.com/google.rpc.RequestInfo",
        "requestId": "req_01M3PQJ64T…"
      }
    ]
  }
}

On the OpenAI formats type is always alectura_policy, a bad key's too: OpenAI's SDKs raise their authentication error from the 401 alone, and code says which refusal it is. On Anthropic's, type is Anthropic's own for the status, and code is ours. On Gemini's, code is the status and status Google's own name for it, as Google's errors have them; our code is the reason of an ErrorInfo in the domain alecturalabs.com, and the request id is a RequestInfo, the places Google's own APIs put them. Each code has one status, and only a code that waiting fixes gets a status that SDKs retry on their own (429 and 5xx).

Status Anthropic type Gemini status Code When
400 invalid_request_error INVALID_ARGUMENT invalid_request The body can't be parsed, so the guardrails can't run; or an Alectura-* header is unknown or invalid.
400 invalid_request_error INVALID_ARGUMENT conflicting_api_keys The request carries two different keys.
401 authentication_error UNAUTHENTICATED invalid_api_key No key, or an unknown, expired or revoked one.
402 billing_error FAILED_PRECONDITION budget_exhausted A limit over a day, week or month is spent. The error names the limit and when it resets.
403 permission_error PERMISSION_DENIED permission_denied The key can't send, such as a management key.
403 permission_error PERMISSION_DENIED wrong_region The key belongs to another region. The message names its host.
403 permission_error PERMISSION_DENIED user_suspended, organization_suspended A suspension.
403 permission_error PERMISSION_DENIED model_not_allowed, region_not_allowed, no_allowed_connection The routing settings leave nowhere to send the request; or a spend limit applies and the model has no list price, so its cost can't be counted.
403 permission_error PERMISSION_DENIED blocked_by_guardrail, attachment_uninspectable, cut_by_guardrail A guardrail blocks the request, or the policy blocks content nothing can inspect, or a guardrail stopped a non-streamed response.
404 not_found_error NOT_FOUND unsupported_endpoint A provider path the gateway doesn't serve.
413 request_too_large INVALID_ARGUMENT request_too_large The body is over the provider's cap: 32 MB for Anthropic, 72 MiB for OpenAI, 100 MB for Gemini.
429 rate_limit_error RESOURCE_EXHAUSTED rate_limited A per-minute limit, or a trial's on the simulator, with Retry-After in seconds.
502 api_error UNAVAILABLE provider_unreachable No candidate could be reached.
503 api_error UNAVAILABLE limiter_unavailable Spend so far isn't known, so a spend limit can't be checked.
504 timeout_error DEADLINE_EXCEEDED provider_timeout No candidate answered in time.

A spent budget is FAILED_PRECONDITION rather than Google's RESOURCE_EXHAUSTED, which Google uses for its own rate limits and its SDKs retry.

A refusal over a limit adds a limit object beside code, or on Gemini's format the same fields as limit_kind, limit_per and so on in the ErrorInfo's metadata, which holds only strings:

"limit": {
  "kind": "spend",
  "per": "user",
  "period": "day",
  "amount": "2.00",
  "used": "1.9962",
  "requested": "0.0079",
  "resets_at": "2026-09-30T00:00:00Z"
}

used is what the period has used so far, and requested what the refused request would have used at most: its estimated cost, its input and maximum output tokens, or one request. A request is refused when the two together would pass amount, so it can be refused before used reaches it.

Provider errors

An error the provider returns passes through unchanged: its status, its body and its headers. Two exceptions:

  • Anything in it that exposes a connection is masked with ***: an endpoint URL, an AWS account or role, the provider's organization or project id, or a credential the request carried, such as the session token a Bedrock signature error quotes.
  • Bedrock's errors are rewritten into Anthropic's shape, keeping the status, since the client speaks Anthropic's format.

When every attempt fails, the client gets the last attempt's failure: the provider's error, or one of ours if the provider never answered.

Headers

Header Direction What happens
Authorization, x-api-key, x-goog-api-key in The Alectura key; never sent upstream
Alectura-* in As above; never sent upstream
anthropic-version, anthropic-beta in Passed to the provider
OpenAI-Organization, OpenAI-Project in Dropped: the connection names the provider account
Other request headers in Passed to the provider, except hop-by-hop headers and those a proxy adds: Forwarded, X-Forwarded-*, Via, X-Real-IP and X-Amzn-Trace-Id
Alectura-Request-Id out On every response, refusals included
Alectura-Connection, Alectura-Region out When a provider was reached
Retry-After out On rate_limited
The provider's response headers out Passed through, rate-limit headers included, except those naming the provider account

Limits

  • limits.max_output_tokens lowers a request's maximum output to the limit, and sets it when the request sets none.
  • Each kind of limit takes its own periods: requests and tokens a minute or a day, and spend a day, a week or a month. So there's no weekly or monthly token limit, and no spend limit per minute.
  • A limit over a minute refuses with 429 and Retry-After. A limit over a day, week or month refuses with 402, which SDKs don't retry, until it resets at the start of the next UTC period; weeks start on Monday.
  • A request reserves its estimated tokens and cost when it's admitted, at the connection's list prices, and settles them from the provider's usage when it ends. Events fire when a limit passes 80% and 95% of its amount, and when it starts refusing.

Compatibility

  • A new field or event in a provider's API passes through without a gateway release.
  • Once its key is known, a request is recorded whatever happens to it, refusals included; see GET /v1/requests in the management API.