Union hosts OpenAI- and Anthropic-compatible endpoints on your own GPUs and puts every model your teams call behind one gateway. Easily manage external LLM providers, virtual keys, budgets, rate limits and guardrails.
Same Python API as the workflows that trained the model, same cloud account, same audit trail.
Six parts of the serving stack, and where each one lives on this page.
The gap between a working checkpoint and a served one is almost entirely infrastructure: the idle GPU bill, the cold start, the key nobody can revoke, and the fact that every prompt your company writes is leaving the building. None of that is a modelling problem.
The GPU is idle most of the day.
An endpoint declares replicas=(0, 2) and a scaledown_after. Between bursts it holds zero replicas and costs nothing; the first request after that brings one back. The replica-count chart in the console is a sawtooth, and that sawtooth is the bill you did not pay.
A 27B checkpoint takes minutes to come up.
Most of that is copying weights to a disk you are about to read once. stream_model=True streams them from blob storage straight into GPU memory instead, and the image itself is cached on the node, so the cold path is the model load rather than the model download.
Nobody knows what the LLM spend is, or whose it is.
Provider keys get pasted into notebooks and CI and never rotated. The gateway inverts that: one endpoint, one key per team or agent, each with its own provider allowlist, monthly budget and rate limit. Spend is attributed to the key that caused it — including on models you host yourself, priced with numbers you supply.
Every prompt leaves your network.
Endpoints run on nodes in your own cloud account and are reachable only from inside it. The gateway runs there too, so the request, the completion and the audit record all stay on your side of the boundary — and when a call does go to a third-party provider, it is the one call you decided to allow.
The same Python API that runs your workflows serves long-running apps next to them. Native integrations wrap vLLM, SGLang, Ollama, FastAPI and Streamlit — declare the app, the image, the resources and the scaling, and Union handles the serving plumbing.
An OpenAI-compatible server in fourteen lines. stream_model=True pulls the weights straight from blob storage into GPU memory, so a cold replica does not wait on a full disk download first.
The LLM Gateway is a Union app that sits in front of everything your teams call — the models you host, and the twenty upstream providers you buy from. Point any OpenAI-compatible SDK at it, authenticate with a virtual key, and address models as provider/model-id. GET /v1/models lists exactly what the calling key can reach.
No client change beyond a base URL and a key. Streaming, tool calls and the rest of the OpenAI surface pass straight through.
Throughput is a number, not a cluster exercise: one replica serves roughly 1,000 req/s, and min/max request rates set the floor that always runs and the burst ceiling above it.
Add an upstream with a Union secret holding its API key — the secret is mounted into the gateway pod, and the control plane never reads it. Restrict a provider to an allowlist of models, and give it a load-balancing weight against the others.
A virtual key is a scoped credential for calling the gateway. Each one carries the providers it may reach, a monthly budget, and a rate limit — so an agent that goes into a loop at 3 a.m. hits its own ceiling instead of your company’s invoice. Revoking a key is one row, not a key rotation across four repos.
An endpoint you host on your own GPUs registers with the gateway like any other provider. Callers address it as union/qwen3-27b and never learn whether it is yours or a vendor’s — which is what makes swapping one for the other a config change.
Upstream providers price themselves. For the models you host, you give the gateway the per-token cost — your amortised GPU rate, your internal chargeback number, whatever your finance team actually uses — and self-hosted traffic lands in the same spend column as everything else.
Requests, spend and keys-with-traffic are reported per key over the last 24 hours, and the logs carry the key on every request. “Who spent that” is a lookup rather than an investigation.
A guardrail you cannot read is a guardrail you cannot trust. Rather than a policy DSL, the gateway calls an ordinary HTTP app you deploy on Union: once before it forwards a request, once after the response comes back. Whatever you can write in Python, you can put in the path.
Same log stream, same metric panes, same scaling events as any endpoint you deploy — because it is the same object. A routing change is a policy version you can point at, not a mystery.
Stand up the gateway, issue a virtual key, and move a single service onto it. Nothing in your application changes except a base URL — and for the first time the answer to “what are we spending on inference” is a number.