Skip to content

Pay for tokens you use

Public rates are live. Monthly packs and reserved GPUs come next.

Plans

Usage is billed against listed token rates. Monthly packs and reserved capacity are next.

Live

Pay as you go

$0+ usage

No platform fee. Tokens bill at the rates on this page.

  • OpenAI-compatible endpoint
  • Input, output, and cached input
  • Usage dashboard

Coming soon

Monthly packs

$20/ month

Prepaid months with twice the usage value of what you pay.

  • $20
  • $60
  • $100
  • $200
  • 2× usage value
  • Priority when the pool is busy

Checkout is not open yet.

Coming soon

Custom

On request

Committed usage, reserved GPUs, and a private Slack.

  • 10-40% off listed token rates
  • Reserved Blackwell capacity
  • SAML and audit logs

Models

USD per million tokens. Input, output, and cached input.

qwen

  • Input
    $0.10
    Output
    $1.60
    Cached
    $0.050

z-ai

  • Input
    $0.10$0.075
    Output
    $0.40$0.25
    Cached
    $0.025$0.015

Volume discounts run 10-40% on committed usage.

Coming soon

GPU

Autoscale from 1 to 100s of nodes and back to 1. You don't pay for idle compute. Best for bursty workloads. For consistent load, we recommend bare Metal.

Join waitlist
  • NVIDIA RTX PRO 6000

    Serverless
    $2.49 / hr
    Metal
    $1.69 / hr
  • NVIDIA B200

    Serverless
    $5.89 / hr
    Metal
    $4.29 / hr
  • NVIDIA B300

    Serverless
    $6.29 / hr
    Metal
    $5.49 / hr

Dedicated

Reserved fleet. No listed rate.

Get a quote

Why serverless?

Same day of traffic. Two bills.

Demand spikes in the afternoon. A fleet of 75 metal nodes left running sits full all day. Serverless follows the curve and lets go of idle cards.

Degraded at peak. Idle off-peak.

Pays for 75 GPUs every hour.

Always on

$0

75 GPUs × 24 hrs × $4.29

Peak demand without idle capacity

Pays for 49 GPUs on average.

Serverless

$0

Avg 49 GPUs × 24 hrs × $5.89

FAQ

Yes. We do not store prompts or completions, and we do not train on them. The only records we keep are usage logs needed for billing: tokens, model, and request counts—not LLM content. KV cache lives only while a request is served. Read the privacy policy.

Dedicated inference

Reserved capacity for your org. We optimize for your workload to give maximum savings.