Skip to main content
This page documents the Gateway’s contract in the corners: what errors look like, what happens when things fail, and where the current boundaries are.

Supported surface

The Gateway implements POST /v1/agent/chat/completions (OpenAI Chat Completions wire format), streaming and non-streaming. Configure your client with base_url="https://app.egisai.co/v1/agent" — the OpenAI client appends /chat/completions itself. The older POST /v1/chat/completions path stays live as a backward-compat alias, so existing integrations keep working without changes. Other OpenAI endpoints (Responses API, embeddings, images, audio) are not proxied yet — use the SDK for those, or call the provider directly. The upstream provider is chosen per call from the model-name prefix, via each provider’s OpenAI-compatible endpoint: Send the matching provider’s key in the Authorization header — the Gateway forwards it verbatim to whichever upstream the model selected. There is no cross-provider fallback (a gpt-… request is never silently retried on Anthropic); routing is deterministic from the model name.

Error envelope

All Gateway-originated errors use the OpenAI error shape, so existing client error handling keeps working. The code field tells you whether Egis or the provider produced the error: Every Gateway-originated refusal — including auth, entitlement, and rate-limit failures — uses this envelope, so an OpenAI-compatible client raises its normal error type and your existing handling keeps working. Provider-side errors (bad provider key, rate limits, model not found) pass through verbatim with the provider’s original status code — those are between you and your provider, though the call is still audited. Response-phase blocks are not errors: they return HTTP 200 with a well-formed completion stub (finish_reason: "content_filter"), because the model call itself succeeded. The stub’s egis field carries the phase, verdict, and reason code.

Key handling

  • Your provider key (Authorization) is forwarded upstream verbatim. It is never logged, never persisted, and never included in an error string. The Gateway holds it in memory for the duration of the request only.
  • Your Egis key (X-Egis-Api-Key) is subject to the same revocation, expiry, and per-key rate limits as SDK keys — one key space, one management page.

Failure modes

This mirrors the SDK’s posture: fail closed on PII, fail open on availability.

When Egis can’t govern

A policy block and an Egis failure are different things, and the Gateway reports them differently. A block is a decision: 400 with egis_policy_blocked. A failure on our side means we learned nothing about your call, and it answers 503 with egis_unavailable — a retryable status, never a 500. Before it gets there, the Gateway degrades in order:
  1. Last-known-good policies. Each successful policy load is cached per agent. If a later load fails, the Gateway keeps enforcing what it enforced a moment ago. Normal operation is unaffected — the cache is only ever read on failure, so policy edits still take effect immediately.
  2. Your degraded-mode setting, on the dashboard’s Gateway page (owners and admins, or anyone with manage_gateway): Every degraded response carries the X-Egis-Degraded header (rules_stale when served from cache, rules_unavailable when forwarded with no policy) so you can alert on it.
If you reach the Gateway through the SDK (init(gateway=True)), you have a second, client-side control: gateway_on_outage. Its default re-runs the call against your own provider client under in-process governance when the Gateway can’t be reached at all — including on a 503 from the rule above. A 4xx is never retried locally, because that’s a decision, not an outage.

Latency

Every Gateway call pays one extra network hop plus policy evaluation. Deterministic rules (PII patterns, regex, size caps, model allow-lists) add single-digit milliseconds. LLM-backed rules (semantic_guard) add a judge round-trip on the calls they match — the same cost they carry in the SDK. Streaming with response-phase rules trades time-to-first- token for enforcement; see Streaming.

What the Gateway does not govern

The Gateway sees network traffic, so it governs what crosses the wire: prompts, completions, and the tool-call requests inside completions. It cannot intercept tool code that executes locally in your process after the response arrives. If you need enforcement at the tool execution boundary (blocking a shell command before it runs, gating an MCP call in-process), that is the SDK’s job — run both where it matters.