Skip to content

What AI transcription actually costs at scale

By Anwar Benhamada · August 6, 2026

Transcription is priced per minute of audio, which makes it feel free. A one-hour podcast costs pennies.

Then you point it at three years of back catalogue and it’s a real number. Worth understanding the cost model before you commit to a provider.

The arithmetic

hours of audio × 60 × price per minute = cost

Two things distort this in practice:

Rounding. Some providers bill per minute started. Two hundred forty-second clips round up to 200 minutes of billing for 133 minutes of audio — a 50% premium invisible on the pricing page.

Reprocessing. You will re-run some files after changing settings, fixing language detection, or switching providers. Budget 20–30% overhead.

Where accuracy actually differs

Most providers are close on clean audio — a single speaker, good microphone, no accent far from the training distribution. On that material, the accuracy difference between the cheapest and most expensive option is often not worth paying for.

The differences appear on hard audio:

So benchmark on your worst audio, not your best. If your material is clean podcasts, buy on price. If it’s four people on a bad conference line, accuracy is worth a large premium and you need to test rather than trust a marketing number.

The features that change the cost calculation

Speaker diarisation (who said what) is often a separate, more expensive tier. If you need it, that’s the price you’re actually comparing.

Word-level timestamps matter if you’re building anything that jumps to a moment. Not all providers give them at the base tier.

Custom vocabulary — the ability to supply product names and jargon upfront — frequently improves accuracy on domain audio more than switching to a more expensive provider does. Check for it before paying for a premium tier.

The cost nobody counts

Correction time. If a transcript needs human review before use, that time dominates everything else.

At 95% accuracy, a 10,000-word transcript has 500 errors to find and fix. Getting to 98% cuts that to 200. If review runs at, say, 20 minutes per hour of audio, and better accuracy halves it, the saving on a large catalogue dwarfs the per-minute price difference.

Which means the cheap provider is frequently the expensive one, and the only way to know is to measure your own correction time on your own audio — not to compare published accuracy figures, which are measured on datasets that aren’t yours.

How to evaluate in an hour

  1. Take five real files: two clean, two hard, one worst-case
  2. Run them through three providers at their base tier
  3. Count actual errors on a fixed passage from each
  4. Time yourself correcting one transcript per provider
  5. Multiply that correction time by your catalogue size

Step 4 is the one everyone skips and the one that decides it.

The pattern that works at volume

Two tiers. Cheap provider for clean audio, premium for hard audio, routed by a quick quality signal — file source, number of speakers, or recording duration.

Most catalogues are mostly clean, so this captures nearly all of the savings while protecting the material that would otherwise cost you hours of correction.

One thing to check before committing

Can you export in a format you own? Plain text, SRT, VTT, or JSON with timestamps. If transcripts live only inside a proprietary editor, migrating a back catalogue later means re-processing all of it — and paying for it twice.