Usage-based billing (also called metered or consumption billing) charges a customer for what they actually consumed — API calls, compute minutes, gigabytes, transactions — instead of a fixed recurring fee. It's the pricing model behind most infrastructure and AI-native products, because cost and value both scale with consumption rather than headcount.
The pitch is simple: pay for what you use. The mechanics behind that sentence — measuring the usage correctly, pricing it correctly, and never double-counting an event — are where most usage billing systems actually live or die.
The pricing shapes usage billing has to support
"Usage-based" isn't one model. A single contract commonly needs several of these on the same invoice:
| Model | How it prices | Typical use |
|---|---|---|
| Linear | A flat rate per unit | Simple metered APIs |
| Volume (flat-band) | One rate applies to the whole volume once a threshold is crossed | Retroactive volume discounts |
| Graduated (tiered) | Different rates for different slices of volume | Progressive discounting |
| Packaged | Priced in blocks (e.g., per 1,000 units), rounded up | Predictable-feeling pricing |
| Prepaid / drawdown | A pre-purchased balance draws down as usage happens | Commitment-based contracts |
| Minimum commitment | A floor charge regardless of actual usage | Guaranteed revenue with usage upside |
A usage-heavy contract frequently mixes two or three of these — a prepaid balance that overages into a graduated rate once it's exhausted, with a minimum commitment underneath the whole thing. Billing software that only supports one model forces every deal into that shape whether or not it fits.
The real problem isn't pricing — it's ingestion
Pricing formulas are the easy part. The hard part is getting raw usage events from wherever they're generated — a warehouse table, a vendor export, an application log — into something that can be trusted enough to put a dollar figure on it.
Three things go wrong constantly:
- Duplicate events. The same usage record gets counted twice because a file was re-uploaded, or an event stream retried a delivery. Without a deliberate idempotency key per event, a re-run of the same import silently doubles the invoice.
- Unclear scope. A usage file has ten columns; only three actually identify a unique, billable event. Importing "whatever's in the row" instead of a deliberately configured uniqueness rule is how the same usage gets billed under two different keys.
- Volume that breaks the tooling. A usage-heavy customer can generate millions of events a month. A billing system that chokes on a multi-gigabyte file, or takes hours to process one, turns invoice generation into an operational bottleneck at exactly the customers who matter most.
Why "wait and see" beats "guess and bill"
The instinct when a usage file is incomplete — missing the last few days of a billing period — is to bill on what's there and true it up later. That's backwards. An incomplete usage picture should produce a draft that waits, not an invoice that ships with a guess baked into it. The correction for a wrong invoice is always more expensive than the correction for a late one.
What to check before trusting a usage billing setup
- Does it support idempotency keys you configure per data source, or does it assume every row is unique by default?
- Can one contract mix linear, volume, graduated, and prepaid models without a custom engineering project?
- What's the actual processing time for a large file — minutes, or does it silently fall over past a certain size?
- Does an incomplete usage file produce a draft that waits, or does the system guess and finalize?
- Can usage sit on the same invoice as a fixed fee or seat charges, or does it always need a separate line item system?
Sonic AI's usage templates support linear, volume, graduated, packaged, prepaid, and minimum-commitment models on the same schedule as fixed fees and seats, with operator-configured idempotency and multi-gigabyte files processed through ClickHouse in minutes rather than hours. See how usage and seats work, or the deeper mechanics on idempotency as a product feature, usage templates, and Postgres usage sync.
Related terms: usage-based billing, meter, usage template, idempotency, minimum commitment.