Skip to main content
The Gateway supports stream: true on every call, but how the stream is delivered depends on whether your organization has any policies that run on the response phase of model calls. The Gateway picks the mode automatically per call — there is nothing to configure.

The two modes

Passthrough

No response-phase rules configured. The provider’s SSE bytes are relayed to your client unchanged. Response evaluation still runs afterwards for the audit row, but with nothing to enforce there is nothing to hold the stream for.

Evaluate-then-release

At least one response-phase rule is active. The Gateway requests the full completion from the provider, evaluates it, and only then emits an SSE sequence to your client. Enforcement is correct-by-construction — refused content never reaches you, not even partially.
The trade-off is deliberate. A blocked response that has already streamed half its tokens to the client is not blocked in any meaningful sense — the leak already happened. So when response enforcement is in play, the Gateway holds the stream until the verdict is in.

What evaluate-then-release looks like to your client

Your client still receives a valid text/event-stream — role chunk, content chunk, finish chunk, [DONE] — so OpenAI SDKs parse it without modification. The observable difference is timing: the first byte arrives after the full completion is ready instead of ~immediately. If the response is blocked, the stream carries a well-formed stub completion with finish_reason: "content_filter" and a block notice as the message content. Agent loops keep parsing; the refused text is never transmitted.
Rules whose phase is both count as response-phase rules for this decision, as do response rules scoped (via applies_to) to tool or mcp surfaces — the model’s tool-call requests ride on the model call’s response side.

Choosing your posture

You control the mode through your policy set, per the table: If time-to-first-token is critical for a product surface and you can accept audit-only response coverage there, keep that agent’s rules on the request phase (agent-scoped policies make this per-agent). If response enforcement matters more than latency — the common posture for regulated content — accept the buffered delivery.

Non-streaming calls

Calls without stream: true are unaffected by any of this: the Gateway always has the complete response in hand before returning, so response policies are always fully enforced at no additional cost.