Business-model visualizer
cheaperinference.com is a marketplace for idle AI compute. On one side, GPU owners with spare capacity sell it cheap. On the other side, you buy it cheap through one OpenAI-compatible API. The marketplace keeps the spread in the middle. There is no separate fee, no subscription. The gap between wholesale and your price is the business.
Your requests flow down to whoever has spare GPUs right now. Money flows up, with the marketplace skimming the middle.
Sellers apply with their spare capacity; the marketplace reviews the fit and settles supply directly with them. You never touch that side.
Live catalog rates, per million input tokens. The bar is the provider's list price; the red sliver is the part you never pay. Zoomed in so you can see it.
Catalog rates from cheaperinference.com, 9 Oct 2026. Input tokens per million. Output tokens are discounted too (e.g. Sonnet 4.5 output $9.75 vs $15.00).
One token's journey, Claude Sonnet 4.5 input pricing. Two numbers are public. The third is the whole business model.
Claude Sonnet 4.5 at a typical 75/25 input/output mix. Drag to your volume.
Illustrative. Assumes a 75% input / 25% output mix and today's catalog rates. Real rates move with live capacity, and your mix will differ.
You pre-fund a balance with a young startup, and your prompts route through a third party. They say they never store prompt or response bodies (usage metadata only) and offer zero-retention routes, but check those routes before sending anything sensitive. Capacity comes from third-party sellers, so discounts and availability move around. "Always the cheapest" is marketing, not a contract. "Never above list price" is a cap, not a promise.