Live
Pay as you go
$0+ usage
No platform fee. Tokens bill at the rates on this page.
- OpenAI-compatible endpoint
- Input, output, and cached input
- Usage dashboard
Public rates are live. Monthly packs and reserved GPUs come next.
Usage is billed against listed token rates. Monthly packs and reserved capacity are next.
Live
$0+ usage
No platform fee. Tokens bill at the rates on this page.
Coming soon
$20/ month
Prepaid months with twice the usage value of what you pay.
Checkout is not open yet.
Coming soon
On request
Committed usage, reserved GPUs, and a private Slack.
USD per million tokens. Input, output, and cached input.
| Model | Input | Output | Cached |
|---|---|---|---|
| qwen | |||
| $0.10 | $1.60 | $0.050 | |
| z-ai | |||
| $0.10$0.075 | $0.40$0.25 | $0.025$0.015 | |
Volume discounts run 10-40% on committed usage.
Coming soon
Autoscale from 1 to 100s of nodes and back to 1. You don't pay for idle compute. Best for bursty workloads. For consistent load, we recommend bare Metal.
NVIDIA RTX PRO 6000
NVIDIA B200
NVIDIA B300
| GPU | Serverless | Metal |
|---|---|---|
| NVIDIA RTX PRO 6000 | $2.49 / hr | $1.69 / hr |
| NVIDIA B200 | $5.89 / hr | $4.29 / hr |
| NVIDIA B300 | $6.29 / hr | $5.49 / hr |
Dedicated
Reserved fleet. No listed rate.
Same day of traffic. Two bills.
Demand spikes in the afternoon. A fleet of 75 metal nodes left running sits full all day. Serverless follows the curve and lets go of idle cards.
Degraded at peak. Idle off-peak.
Always on
$0
75 GPUs × 24 hrs × $4.29
Peak demand without idle capacity
Serverless
$0
Avg 49 GPUs × 24 hrs × $5.89
Reserved capacity for your org. We optimize for your workload to give maximum savings.