Build

Your gateway (InvokeModel)

invokeModelGateway() — Anthropic models behind your organisation's gateway that forwards the Bedrock InvokeModel operation over plain HTTPS with an API-key header. Streaming, tool calls, a key per model, typed errors.

Many organisations reach Claude through their own gateway: plain HTTPS, an API-key header, and the Bedrock InvokeModel operation. bedrock() signs with AWS credentials against the Converse operation, and providerFromEnv() has no base URL for this wire — so neither reaches it. invokeModelGateway() does, with no SDK.

Use

const provider = withRetry(  invokeModelGateway({    baseUrl: 'https://llm-gateway.example.com/bedrock',    apiKeyHeader: 'api-key',    // One key per model: this gateway scopes keys, and the wrong one is a 403.    apiKey: { [HAIKU]: 'haiku-key', [SONNET]: 'sonnet-key' },    model: HAIKU,    fetch: scriptedGateway(log), // ← drop this line to talk to the real gateway  }),); // a 429, a 5xx or a network failure is retried by withRetry, never inside the adapter

Then Agent.create({ provider, model: 'invoke-model-gateway' }). Prefer new? new InvokeModelGatewayProvider(options) is the same provider as a class. The shorthand 'invoke-model-gateway' sends the model you configured; any other model name is sent as given, so one provider can serve several models.

The options object is typed InvokeModelGatewayOptions:

OptionWhat it is
baseUrlThe gateway root. Requests go to {baseUrl}/model/{modelId}/invoke and …/invoke-with-response-stream.
apiKeyHeaderThe header your gateway reads the key from ('api-key', 'x-api-key', …). Required — a guessed default would fail as a 401 that looks like a bad key. For a bearer scheme, use 'authorization' and put 'Bearer …' in the key.
apiKeyAn InvokeModelGatewayKey: a key; a map from model id to key; or a function of the model id, called before every request (a rotated key is picked up without rebuilding the provider).
modelThe model id your gateway knows, e.g. us.anthropic.claude-haiku-4-5-20251001-v1:0.
defaultMaxTokensmax_tokens when the request sets none. Default 4096.
parallelToolCallsfalse limits the model to one tool per reply.
fetchAn InvokeModelGatewayFetch to use instead of the global fetch: an mTLS agent, a proxy, or a test double.

A key per model

Gateways often scope a key to some models, and asking with the wrong one gets a 403. Give each model id its own key with a map (as above) or a function. A model the map does not name is refused before any request is sent, and the error lists the ids the map does name. The key itself is never in an error message.

What goes on the wire

The body is Anthropic's Messages API, with anthropic_version: "bedrock-2023-05-31" and no model field (the model is in the path). The system prompt rides the top-level system field; tool calls go out as tool_use blocks and results come back as tool_result blocks on a user turn. A forced tool (tool_choice) is supported.

stream() reads the SSE the gateway sends after decoding AWS's binary framing. Bare data: lines and event: + data: pairs both work. Text arrives as chunks; tool arguments are reassembled from their input_json_delta fragments, and the final chunk carries the full response (tool calls, usage, stop reason).

Errors

Every failure is an InvokeModelGatewayError with a reason (typed InvokeModelGatewayErrorReason) and the modelId:

reasonMeaningAlso set
http-statusThe gateway answered non-2xx. A 403 message says the key is not allowed that model.status, bodyExcerpt, retryAfterSeconds (from Retry-After)
networkNo answer at all (DNS, TLS, connection reset).cause
no-keyNo key for this model id.—
no-modelThe shorthand was asked for and no model is configured.—
unreadable-responseA 2xx body that is not a Messages response (for example an HTML page from a proxy).bodyExcerpt
streamThe stream sent an error event, a forwarded AWS exception, or an event that is not JSON.—
malformed-tool-argsStreamed tool arguments that are not JSON. The call is refused, never run with {}.toolName
invalid-optionsThrown when you create the provider: the options cannot make a request.—

"Prompt is too long" becomes a ContextWindowExceededError instead, and a cancelled request is passed through as the AbortError it is.

Retries: compose withRetry

The provider makes one attempt per call and has no retry loop of its own. Every InvokeModelGatewayError carries retryable: true for a 429, a 5xx and a network failure, false for everything else. withRetry's default policy honours the false, so it retries exactly those three:

  • A 403 or 404 is an answer, not a passing fault, and fails straight away.
  • A refusal raised before any request (no-key, no-model) is never repeated, and nothing waits out a backoff for it.
  • A 2xx the provider could not read (unreadable-response) is never re-sent: the model may already have run, and billed.
const log: string[] = [];const provider = withRetry(  invokeModelGateway({    baseUrl: 'https://llm-gateway.example.com/bedrock',    apiKeyHeader: 'api-key',    apiKey: 'sonnet-key',    model: SONNET,    fetch: scriptedGateway(log, true),  }),  { initialDelayMs: 10 },);await provider.complete({  model: 'invoke-model-gateway',  messages: [{ role: 'user', content: 'hi' }],});

Today withRetry retries complete() only; it passes stream() through unchanged. An agent that streams gets no retry until withRetry learns to retry a stream that fails before its first chunk. Every failure this provider raises on stream() (an HTTP status, a network error) happens before the first chunk, so that change will cover it without any change to this provider.

Next steps

  • AWS Bedrock — bedrock(), for AWS credentials and the Converse API
  • Resilience — withRetry / withFallback
  • Custom provider — createProvider({ kind: 'invoke-model-gateway', ... })

On this page