Build Data Factories
on your Own Cloud

Union is the durable AI runtime your team owns. Turn your data into models, agents and apps with plain Python, on infrastructure you control. If a step fails, it picks up where it stopped.

Get Started Try for free
Read the docs
Trusted in production by leading AI teams
01
SOVEREIGN AI

Your models, your data,
your perimeter.

Union’s control plane orchestrates the work and holds references, never payloads. Verifiable by inspection — topological, not behavioral.

  • Secrets and key custody
  • Encryption
  • Network isolation
  • Governance
Explore Sovereign AI →
train_grpo ✓ Succeeded Triggered
Run: ufe486be308c71f1d · Task: trainer.train_grpo · Cluster: oc-production
Summary Logs Reports Code
Setup 1s Succeeded 3m 25s
Input
profile_name (string)*
experiment
dataset (file)*
⛁ rl-tasks-dataset@uwbwvdrsf2gzj27gmvgp-5j…
s3://union-oc-production-demo-raw/…/rl_tasks_merged.parquet
Output
o0 (directory)*
⛁ policy-checkpoint@ufe486be308c71f1d-a0-1
s3://union-oc-production-demo-raw/…/policy-checkpoint
Run Logs Kubernetes Events Cloudwatch Logs ↗
Filter logs Timestamps
■Sep 02 23:12:03.256[flyte] WARNING Flyte runtime started for action a0 with run name uqgmjb48ndh88qhp8zcl
■Sep 02 23:12:05.561[flyte] WARNING It is recommended to use a minimum of 2 replicas, to avoid starvation. Options: increase concurrency, increase replicas, or turn off reuse for the parent task.
■Sep 02 23:13:15.095[flyte] WARNING Flyte runtime completed for action a0 with run name uqgmjb48ndh88qhp8zcl
main ⟳ Refresh | Off ▾
GRPO training — grpo-smoke-ufe486be308c71f1d
10
step
0.200
mean reward
0.00%
pass rate
nan
loss
reward (max 1.2)
▪ mean reward▪ pass rate
Code
▾ Files
trainer.py
1import flyte
2from flyte.clustered import ClusteredTaskEnvironment, TorchRun
3 
4trainer = ClusteredTaskEnvironment(
5 name="trainer", resources=flyte.Resources(gpu="A100:4"),
6 replicas=4, nproc_per_node=4, runtime=TorchRun(),
7)
8 
9@trainer.task(trigger=flyte.OnArtifact("rl-tasks-dataset"))
10async def train_grpo(dataset: flyte.io.File) -> flyte.io.Dir:
11 …
02
AI FACTORY

Own how your models get made

A checkpoint you can’t reproduce is a checkpoint you don’t own. The factory wires teams together through named artifacts.

  • Typed, versioned artifacts
  • Multi-node training
  • OnArtifact triggers
  • Evals per checkpoint
  • Batch & real-time inference
Explore the AI Factory →
Artifact lineageThe inspectable assembly line: which run produced each artifact version and which stations consume it downstream.
03
PRODUCTION AGENTS

An agent is a workload,
so run it like one.

Every agent is a loop over a model that calls tools. Union makes each agent step durable and replayable, so a crash on step seven resumes where it left off.

  • Supports OpenAI, Claude, PydanticAI, CrewAI and more
  • Run code in ephemeral container or Monty sandboxes
  • Native memory, MCP servers, human approvals and chat UI
Explore the Agents Platform →
release_agent.py flyte.ai.agents
1import flyte
2from flyte.ai.agents import Agent
3
4env = flyte.TaskEnvironment("agent")
5
6@env.task(cache="auto", retries=3) # a tool is a task
7async def search(q: str) -> str:
8 """Search the corpus for a query."""
9 return await index.query(q)
10
11agent = Agent(
12 name="release-shepherd",
13 model="claude-haiku-4-5",
14 tools=[search, post_digest],
15 max_turns=12,
16)
17
18@env.task(report=True) # every turn traced
19async def main(question: str) -> str:
20 r = await agent.run.aio(question)
21 return r.summary
04
DURABLE AI RUNTIME

Recover from any failure

Owning your AI starts with a runtime that survives your infrastructure. Pick a workload, pick a failure, and press Run — the task fails, and Union recovers it with no manual intervention.

  • Recover, fork, and replay any run
  • Change the hardware inside an except block
Explore the Durable AI Runtime →
Use case:
Break it with:
main ▸Readyidle
· main 7 16.9m
A provider returns 429 mid-shard. Flyte catches it, backs off exponentially per the RetryStrategy, and retries the one failed task — completed shards are never rerun.
05
INFERENCE

Serve the model.
Govern the traffic.

Union puts LLM providers and self-hosted models behind one gateway with virtual keys, budgets, and guardrails you can customize.

  • vLLM, SGLang and Ollama endpoints — plus FastAPI services and Streamlit dashboards
  • Scale to zero when idle; replicas and scaledown are one line of config
  • Logs, request rates, latency and per-card GPU metrics on the app’s own page
Explore Inference →
serve_llm.py flyteplugins-vllm
1from flyteplugins.vllm import VLLMAppEnvironment
2import flyte
3
4vllm_app = VLLMAppEnvironment(
5 name="my-llm-app",
6 model_hf_path="Qwen/Qwen3-0.6B",
7 resources=flyte.Resources(
8 cpu="4", memory="16Gi", gpu="L40s:1"),
9 scaling=flyte.app.Scaling(
10 replicas=(0, 2), scaledown_after=300),
11 stream_model=True, # blob store straight to GPU
12)
13
14app = flyte.serve(vllm_app)
15print(f"OpenAI-compatible API: {app.url}/v1")
06
ADVANCED SCHEDULING

One fleet, many teams,
nobody starved.

Queues carry priority, concurrency, quotas, and a cluster selector, so a backfill never eats production capacity and a GPU workload lands on the cluster you chose.

  • Priority and per-team quotas
  • Gang scheduling
  • Multi-cluster routing
Explore Advanced Scheduling →
8 × H100  ·  gpu-h200-1
Already running Gang · 6 GPU training job Small work · 1–2 GPU Idle capacity
Queue, head first
ganggrpo-train6 GPU · ~5m
taskeval-shard-a2 GPU · ~2m
taskeval-shard-b2 GPU · ~2m
taskfeaturize1 GPU · ~3m

Results, proven in production.

“We can scale to 200,000–300,000 pods with the escalation logic baked right in, and the out-of-memory and scheduling headaches I used to fight are simply gone.”
Jay GanbatPrincipal Bioinformatics Engineer · Prima Mente

See all case studies →

Enterprise grade Flyte

Open source at the core.

Union is built on Flyte, the open-source AI runtime we create and maintain under the Linux Foundation AI & Data.

A Linux Foundation AI & Data Project
4000+ companies using Flyte today
18M+ Flyte SDK downloads

Write workflows in pure

No DSL, no YAML hell. A TaskEnvironment declares images, resources, secrets and retry policy; tasks are plain async functions. Branch, loop and fan out with ordinary control flow — the runtime records every await so a crash resumes instead of restarting.

  • Type-checked I/O between tasks — files, DataFrames, dataclasses
  • Native agents: wrap tools as tasks, every LLM call traced
  • Same code runs locally, on the devbox, or on your cluster
flyteorg/flyte-sdk ↗ pip install flyte
agent.py flyte-sdk
1import flyte
2from flyte.ai.agents import Agent
3
4env = flyte.TaskEnvironment(
5 name="weather-agent",
6 resources=flyte.Resources(cpu=1, memory="250Mi"),
7 secrets=flyte.Secret("OPENAI_API_KEY"),
8)
9
10@env.task
11async def get_weather(city: str) -> dict:
12 ...
13
14agent = Agent(name="Weather agent", tools=[get_weather])
15
16@env.task
17async def main(request: str) -> str:
18 result = await agent.run.aio(agent, input=request)
19 return result.summary
20
21# flyte run agent.py main --request "Weather in Tokyo?"

Make your existing stack durable

All integrations →

Draw the perimeter.
Keep everything inside it.

Bring one workload — training, data, or serving. We'll run it inside your cloud this week, and you can inspect exactly what the control plane sees: references, and nothing else.

Union.ai achieves 9.8x ROI according to analysts. Independent economic validation across compute savings, engineering velocity, and platform consolidation.
View the report