Pricing

Pay per model-minute.

Standard uses shared scale-to-zero capacity and bills model time by the second. Instant reserves one dedicated warm model replica for 60 minutes to 30 days. Customer-connected Policies cost $0.01 per active model-minute plus provider usage.

One walletPrepaid credits. Policy runtime and training jobs draw down the same balance.
Pinned servingEach model revision stays on one qualified RSI runtime. The serving recipe remains internal.
Two speedsUse Standard for shared scale-to-zero calls. Use Instant when one dedicated warm replica must be yours.
Rate card

Every model shows its rate.

Live rates
ModelAccessStandardInstantStatus

Standard uses shared capacity and includes startup, execution, and five minutes of paid warm reuse. It is measured by the second and shown per model-minute. Instant reserves one dedicated warm model replica for 60 minutes to 30 days and replaces Standard runtime charges while active. Requests rejected before hosted compute starts are free. Once compute starts, elapsed runtime remains metered if model execution fails. Customer-connected Policies use your own provider account. RSI bills $0.01 per active model-minute for the Policy control plane, and the provider bills model usage directly. Fathom LoRA tuning costs $0.80 per million training tokens with a $5 minimum. RSI quotes and prepays the job before it starts.

Which speed should I use?

Use Standard for trials and occasional calls. Use Instant for a live robot shift or production period that needs one private warm replica.

Is there a monthly fee?

No. There is no subscription. Top up credits once and usage draws them down. An idle organization spends nothing.

How does Instant work?

Choose 60 minutes through 30 days. RSI reserves one warm replica exclusively for your organization. It starts billing after the replica is ready, and calls during the reservation have no added Standard runtime charge.

What about a tuned Policy?

Fathom LoRA tuning is $0.80 per million training tokens with a $5 minimum. A tuned revision keeps Fathom's normal inference rate. Artifact storage is included.

Can one Policy serve a fleet?

Yes. Every robot can call the same stable Policy name. One Instant reservation runs one model call at a time. Additional calls queue behind your own requests.