Skip to content

OpenAI API pricing, explained — per-token linear usage on prepaid credit

OpenAI's API is pure metered pricing — a per-million-token rate that differs by model and by input, cached input and output — drawn against a prepaid credit balance. What the model is, what a statement looks like, how Twilio's per-message pricing adds a pass-through fee to the same shape, and how to bill it.

Sonic AI team

Finance operations

17 Sept 20266 min read

Pricing ModelsHow well-known products price — seats, credits, tiers, blocks — and how that shape is billed.
ModelInput / 1MCached input / 1MOutput / 1M
gpt-5.6-terra$2.00$0.20$12.00
gpt-5.6-sol$4.00$0.40$20.00
gpt-4o-mini$0.15$0.075$0.60

Editorial note. OpenAI rates are from the API pricing page; Twilio rates are US SMS list prices from twilio.com/en-us/sms/pricing/us. Both as of September 2026, both change often (OpenAI with every model launch; Twilio's carrier fees on the carriers' schedule). Neither company is affiliated with Sonic AI. Use the figures to understand the structure, and check the live pages before quoting.

If you want the purest example of usage-based pricing a finance team will ever reconcile, it is an OpenAI API bill. There is no subscription, no seat, no included allowance — just a meter, several rates, and a balance that the meter draws down. Every AI product built on top of it inherits some version of this model, and so does every company that sells "pay per call" of anything — including Twilio, whose per-message pricing adds the one wrinkle OpenAI does not have, covered further down.

What OpenAI charges

Per million tokens, three rates per model:

Model Input Cached input Output
gpt-6-astra $10.00 $1.00 $50.00
gpt-5.6-sol $4.00 $0.40 $20.00
gpt-5.6-terra $2.00 $0.20 $12.00
gpt-4o-mini $0.15 $0.075 $0.60

Batch API requests bill at 50% of standard; cached input at roughly 90% off; usage is paid from a prepaid credit balance that expires after a year.

The model underneath

Both are linear usage pricing — quantity × unit rate, summed per period, billed in arrears. The differences are all in the shape of the meters:

OpenAI Twilio
Meters One per (model × token type × Batch/standard) — a dozen or more per customer One for messages, one per carrier fee — from the same events
Rate Flat per unit; no volume banding on the public list Flat per unit, with volume tiers and committed-volume floors for big senders
Payment Prepaid balance drawn down; statement, not invoice Pay-as-you-go from a balance; invoiced with commitment at enterprise
Volume Hundreds of millions of events a month Millions; each event needs its carrier attribute

Twilio's committed volume is a minimum commitment — the contract says "at least $X a year", the meter says what was sent, and the invoice bills the higher of the two with a shortfall line if usage fell short.

What a month looks like

A customer mostly on gpt-5.6-terra, with some batch work:

Meter Tokens Rate / 1M Amount
terra — input 180,000,000 $2.00 $360.00
terra — cached input 420,000,000 $0.20 $84.00
terra — output 45,000,000 $12.00 $540.00
terra — input, Batch 300,000,000 $1.00 $300.00
terra — output, Batch 60,000,000 $6.00 $360.00
September usage $1,644.00 — drawn from prepaid balance

Who else prices per unit — Twilio, and the pass-through version

OpenAI is the purest metered model. Twilio is the same model with a twist every payments, logistics and telecom product shares: one unit of usage carries more than one price. An SMS is $0.0083 to Twilio plus a carrier fee that depends on which network delivered it, both metered per message on the same invoice — and large senders sign committed annual volumes for a discount, which turns the linear rate into a minimum commitment with a shortfall true-up. Anthropic, Deepgram and SendGrid price the OpenAI way; Stripe (a percentage plus a fixed fee per transaction) and cloud egress price the Twilio way.

Twilio — per message, plus pass-through:

Component Rate
SMS, sent or received $0.0083 per message (long code, toll-free, short code)
Carrier fee, outbound SMS $0.0035 AT&T · $0.0045 T-Mobile, Verizon
MMS, sent $0.022
Volume discounts Automatic at higher monthly volume; tiers not fully published
Committed use "Commit to annual message volumes and receive a discount beyond standard volume tiers"

Twilio — 250,000 outbound US SMS, split 40% AT&T, 35% T-Mobile, 25% Verizon:

Line Qty Rate Amount
SMS sent 250,000 $0.0083 $2,075.00
Carrier fee — AT&T 100,000 $0.0035 $350.00
Carrier fee — T-Mobile 87,500 $0.0045 $393.75
Carrier fee — Verizon 62,500 $0.0045 $281.25
September $3,100.00

The carrier lines are a third of the Twilio bill. A system that only knows "messages × $0.0083" leaves finance rebuilding $1,025 of pass-through by hand every month. On OpenAI the equivalent trap is treating "tokens × rate" as one line when it is five.

Where metered billing goes wrong

  • One event, several lines. A Twilio message counts once for the base rate and once for its carrier's fee; an OpenAI request counts on input and output meters. The template must fan one row out to multiple meters.
  • Idempotency. Usage arrives from logs, and logs get re-shipped. Every event needs a stable key so a replayed hour never double-bills.
  • Rates on someone else's schedule. Carrier fees change when the carriers say so; model prices change at launch. The catalog needs dated rates so an issued statement keeps the rate it was issued at.
  • Volume versus graduated. When a discount kicks in at 300,000 messages, does it apply to all messages that month or only those above? Different totals; the contract must say the word. See graduated vs volume.
  • Prepaid balance is state. Running balance, expiry, auto-recharge threshold — per customer, on the statement, or the first support ticket is "why did my balance drop?".

Billing this shape on Sonic AI

Both are linear list prices per meter on Sonic, on a monthly in-arrears schedule. OpenAI is one meter per (model, token type, tier) with a sum aggregation; Twilio is a message meter plus a meter per carrier fee, all fed from the same events by one usage template with a per-event idempotency key, into the ClickHouse-backed usage store built for this volume. Volume discounts are a volume or graduated list price depending on the contract's wording; a committed annual volume is a minimum commitment group over the covered meters, with the shortfall computed at period end. If you resell prepaid credit, the top-up is an invoice and the monthly usage draws against it, so the customer sees both documents.

See usage & seats, usage-based billing, explained, and minimum commitment billing.

If you price like OpenAI

  • How many meters do you really have once split by SKU, tier and pass-through? Enumerate them.
  • Are events idempotent from the source?
  • Are rates dated in the catalog?
  • Is your discount volume or graduated, and is a commitment's shortfall a computed line?

Bring a month of usage logs and your rate card to a demo — we will run them through a template on the call.

Similar pricing models

Per-unit metered pricing with many meters is the model behind Anthropic, Deepgram and SendGrid; Twilio (above) adds pass-through fees and committed volume; Stripe adds a percentage rate. When a plan includes an allowance before the per-unit rate applies, you are looking at Cloudflare Workers or ElevenLabs; when the rate steps with volume, at AWS S3; when a year of it is bought up front, at Snowflake.

Related terms: usage-based billing, meter, idempotency, minimum commitment.

See it on the platform

Everything above describes how Sonic actually runs it — the product pages show the screens.

Related questions

Still have questions?

Are seats events or snapshots?

Snapshots. Each row is how many they had that day. Sonic compares it to the balance on record and writes an event only when the count changes. Vendor increase/delta columns are not imported — they break at customer boundaries.

What are confirmation days?

An optional hold on seat adds. If you set N days, an increase on date D bills only if the new count still holds through D+N. Removals bill immediately. Incomplete files show as awaiting, not as a guessed invoice line.

How do usage and seats get in?

Templates map vendor files or a Postgres connection. Usage is events. Seats are daily snapshots compared to the balance on record. Large files are a first-class path, not an afterthought.

See the full FAQ →

Next

See it on your contracts

A walkthrough on the agreements and files you actually bill from — not a slide deck.