Usage rates

How usage is charged

Charges apply only to usage past your plan's monthly allowance, and only for two things: the model your agents run on, and the compute your MCP connectors use. Here are the exact rates you are charged, and what makes each one go up or down.
Model usage
Model usage is metered by the token, split between input, the text and context your agents send in, and output, the text the model writes back. Output is billed at the higher rate, so tighter, more focused responses cost less. Input that repeats across a session is largely served from cache at a deep discount, so a stable, reusable context keeps input costs low.
Token figures are illustrative. Each allowance is denominated as a cost budget, which input and output debit at their respective per-token rates. Published figures represent an even allocation of that budget across input and output at prevailing rates; realized token volumes vary with the input to output ratio and with cache utilization. Cache reads are weighted between 0.1 and 0.25 depending on the model, and cache writes at 1.25, so context that stays stable across a session stretches the input side well past its published figure.
Nexus Scale
19.9M input / 4M output tokens
$249 / month
Nexus Pass
3.4M input / 680K output tokens
$39 / month
Pay as you go
No monthly allowance, you're billed from your credit balance for what you use
$0 / month
Overage
Past your plan's monthly allowance, model usage bills to your balance at these rates: Opus $7.00 per 1M input tokens and $32.00 per 1M output tokens, Sonnet $3.00 per 1M input tokens and $15.00 per 1M output tokens, Haiku $1.00 per 1M input tokens and $5.00 per 1M output tokens, Kimi K3 $5.00 per 1M input tokens and $20.00 per 1M output tokens, Grok 4.6 $3.00 per 1M input tokens and $8.00 per 1M output tokens.
On Pay as you go there is no monthly allowance, so model usage is billed from your credit balance as you use it.
MCP compute
Compute is metered per connector session, from container start to session termination, across two dimensions: processor time consumed, and peak resident memory applied over the session's elapsed duration. Usage is reported at fixed intervals and again on termination; sessions may remain resident after a request completes. A minimum memory allocation applies to every session, and a worst case allocation applies where either dimension is unmeasurable. Realized consumption varies with runtime duration, concurrency, and memory resident during execution. Both dimensions are billed past your plan's monthly allowance.
Processor time
CPU time your connectors use while running
$0.179 per vCPU-hour
Memory
Memory held live while your connectors run
$0.0189 per GB-hour
Rates shown are current and subject to change. Usage is metered continuously as your agents and connectors run, and totals may be rounded when they are tallied for billing. Figures shown in the app are estimates, not a guarantee of the final amount billed.