Supported surface
The Gateway implementsPOST /v1/agent/chat/completions (OpenAI
Chat Completions wire format), streaming and non-streaming. Configure
your client with base_url="https://app.egisai.co/v1/agent" — the
OpenAI client appends /chat/completions itself. The older
POST /v1/chat/completions path stays live as a backward-compat alias,
so existing integrations keep working without changes. Other OpenAI
endpoints (Responses API, embeddings, images, audio) are not proxied
yet — use the SDK for those, or call the provider directly.
The upstream provider is chosen per call from the model-name prefix,
via each provider’s OpenAI-compatible endpoint:
Send the matching provider’s key in the
Authorization header — the
Gateway forwards it verbatim to whichever upstream the model selected.
There is no cross-provider fallback (a gpt-… request is never
silently retried on Anthropic); routing is deterministic from the
model name.
Error envelope
All Gateway-originated errors use the OpenAI error shape, so existing client error handling keeps working. Thecode field tells you whether
Egis or the provider produced the error:
Every Gateway-originated refusal — including auth, entitlement, and
rate-limit failures — uses this envelope, so an OpenAI-compatible client
raises its normal error type and your existing handling keeps working.
Provider-side errors (bad provider key, rate limits, model not found)
pass through verbatim with the provider’s original status code —
those are between you and your provider, though the call is still
audited.
Response-phase blocks are not errors: they return HTTP 200 with a
well-formed completion stub (
finish_reason: "content_filter"), because
the model call itself succeeded. The stub’s egis field carries the
phase, verdict, and reason code.
Key handling
- Your provider key (
Authorization) is forwarded upstream verbatim. It is never logged, never persisted, and never included in an error string. The Gateway holds it in memory for the duration of the request only. - Your Egis key (
X-Egis-Api-Key) is subject to the same revocation, expiry, and per-key rate limits as SDK keys — one key space, one management page.
Failure modes
This mirrors the SDK’s posture: fail closed on PII, fail open on
availability.
When Egis can’t govern
A policy block and an Egis failure are different things, and the Gateway reports them differently. A block is a decision:400 with
egis_policy_blocked. A failure on our side means we learned nothing
about your call, and it answers 503 with egis_unavailable — a
retryable status, never a 500.
Before it gets there, the Gateway degrades in order:
- Last-known-good policies. Each successful policy load is cached per agent. If a later load fails, the Gateway keeps enforcing what it enforced a moment ago. Normal operation is unaffected — the cache is only ever read on failure, so policy edits still take effect immediately.
-
Your degraded-mode setting, on the dashboard’s Gateway page
(owners and admins, or anyone with
manage_gateway):Every degraded response carries theX-Egis-Degradedheader (rules_stalewhen served from cache,rules_unavailablewhen forwarded with no policy) so you can alert on it.
init(gateway=True)), you
have a second, client-side control:
gateway_on_outage. Its
default re-runs the call against your own provider client under
in-process governance when the Gateway can’t be reached at all —
including on a 503 from the rule above. A 4xx is never retried
locally, because that’s a decision, not an outage.
Latency
Every Gateway call pays one extra network hop plus policy evaluation. Deterministic rules (PII patterns, regex, size caps, model allow-lists) add single-digit milliseconds. LLM-backed rules (semantic_guard) add
a judge round-trip on the calls they match — the same cost they carry
in the SDK. Streaming with response-phase rules trades time-to-first-
token for enforcement; see Streaming.