Your gateway (InvokeModel)
invokeModelGateway() — Anthropic models behind your organisation's gateway that forwards the Bedrock InvokeModel operation over plain HTTPS with an API-key header. Streaming, tool calls, a key per model, typed errors.
Many organisations reach Claude through their own gateway: plain HTTPS, an API-key header, and the Bedrock InvokeModel operation.
bedrock()signs with AWS credentials against the Converse operation, andproviderFromEnv()has no base URL for this wire — so neither reaches it.invokeModelGateway()does, with no SDK.
Use
const provider = withRetry( invokeModelGateway({ baseUrl: 'https://llm-gateway.example.com/bedrock', apiKeyHeader: 'api-key', // One key per model: this gateway scopes keys, and the wrong one is a 403. apiKey: { [HAIKU]: 'haiku-key', [SONNET]: 'sonnet-key' }, model: HAIKU, fetch: scriptedGateway(log), // ← drop this line to talk to the real gateway }),); // a 429, a 5xx or a network failure is retried by withRetry, never inside the adapterThen Agent.create({ provider, model: 'invoke-model-gateway' }). Prefer new? new InvokeModelGatewayProvider(options) is the same provider as a class. The shorthand 'invoke-model-gateway' sends the model you configured; any other model name is sent as given, so one provider can serve several models.
The options object is typed InvokeModelGatewayOptions:
| Option | What it is |
|---|---|
baseUrl | The gateway root. Requests go to {baseUrl}/model/{modelId}/invoke and …/invoke-with-response-stream. |
apiKeyHeader | The header your gateway reads the key from ('api-key', 'x-api-key', …). Required — a guessed default would fail as a 401 that looks like a bad key. For a bearer scheme, use 'authorization' and put 'Bearer …' in the key. |
apiKey | An InvokeModelGatewayKey: a key; a map from model id to key; or a function of the model id, called before every request (a rotated key is picked up without rebuilding the provider). |
model | The model id your gateway knows, e.g. us.anthropic.claude-haiku-4-5-20251001-v1:0. |
defaultMaxTokens | max_tokens when the request sets none. Default 4096. |
parallelToolCalls | false limits the model to one tool per reply. |
fetch | An InvokeModelGatewayFetch to use instead of the global fetch: an mTLS agent, a proxy, or a test double. |
A key per model
Gateways often scope a key to some models, and asking with the wrong one gets a 403. Give each model id its own key with a map (as above) or a function. A model the map does not name is refused before any request is sent, and the error lists the ids the map does name. The key itself is never in an error message.
What goes on the wire
The body is Anthropic's Messages API, with anthropic_version: "bedrock-2023-05-31" and no model field (the model is in the path). The system prompt rides the top-level system field; tool calls go out as tool_use blocks and results come back as tool_result blocks on a user turn. A forced tool (tool_choice) is supported.
stream() reads the SSE the gateway sends after decoding AWS's binary framing. Bare data: lines and event: + data: pairs both work. Text arrives as chunks; tool arguments are reassembled from their input_json_delta fragments, and the final chunk carries the full response (tool calls, usage, stop reason).
Errors
Every failure is an InvokeModelGatewayError with a reason (typed InvokeModelGatewayErrorReason) and the modelId:
reason | Meaning | Also set |
|---|---|---|
http-status | The gateway answered non-2xx. A 403 message says the key is not allowed that model. | status, bodyExcerpt, retryAfterSeconds (from Retry-After) |
network | No answer at all (DNS, TLS, connection reset). | cause |
no-key | No key for this model id. | — |
no-model | The shorthand was asked for and no model is configured. | — |
unreadable-response | A 2xx body that is not a Messages response (for example an HTML page from a proxy). | bodyExcerpt |
stream | The stream sent an error event, a forwarded AWS exception, or an event that is not JSON. | — |
malformed-tool-args | Streamed tool arguments that are not JSON. The call is refused, never run with {}. | toolName |
invalid-options | Thrown when you create the provider: the options cannot make a request. | — |
"Prompt is too long" becomes a ContextWindowExceededError instead, and a cancelled request is passed through as the AbortError it is.
Retries: compose withRetry
The provider makes one attempt per call and has no retry loop of its own. Every InvokeModelGatewayError carries retryable: true for a 429, a 5xx and a network failure, false for everything else. withRetry's default policy honours the false, so it retries exactly those three:
- A 403 or 404 is an answer, not a passing fault, and fails straight away.
- A refusal raised before any request (
no-key,no-model) is never repeated, and nothing waits out a backoff for it. - A 2xx the provider could not read (
unreadable-response) is never re-sent: the model may already have run, and billed.
const log: string[] = [];const provider = withRetry( invokeModelGateway({ baseUrl: 'https://llm-gateway.example.com/bedrock', apiKeyHeader: 'api-key', apiKey: 'sonnet-key', model: SONNET, fetch: scriptedGateway(log, true), }), { initialDelayMs: 10 },);await provider.complete({ model: 'invoke-model-gateway', messages: [{ role: 'user', content: 'hi' }],});Today withRetry retries complete() only; it passes stream() through unchanged. An agent that streams gets no retry until withRetry learns to retry a stream that fails before its first chunk. Every failure this provider raises on stream() (an HTTP status, a network error) happens before the first chunk, so that change will cover it without any change to this provider.
Next steps
- AWS Bedrock —
bedrock(), for AWS credentials and the Converse API - Resilience —
withRetry/withFallback - Custom provider —
createProvider({ kind: 'invoke-model-gateway', ... })
AWS Bedrock
bedrock() — model-agnostic provider via AWS Bedrock's Converse API. One adapter covers Claude, Llama, Mistral, Titan, Mixtral on Bedrock.
Gemini
gemini() — Google's models through the native @google/genai SDK, on Vertex or the Gemini API. Honest cached- and thinking-token counts, JSON Schema tools with no OpenAPI translation, and a forced single tool that is actually forced.
