Speech API Cost Calculator (TTS & STT)
Speech APIs bill in two different currencies — characters for text-to-speech, minutes for transcription — and mixing them up wrecks budgets. Enter your monthly audio volume to price TTS, TTS HD, gpt-4o-mini-tts, and Whisper side by side.
Text-to-speech
Speech-to-text (Whisper)
How speech API pricing works
Text-to-speech is billed per character on the classic models: TTS at $15 per million characters and TTS HD at $30 per million. A minute of normal-paced speech is roughly 1,000 characters, so TTS works out near $0.015 per minute of audio and HD doubles that. The newer gpt-4o-mini-tts bills tokens instead — $0.60 per million input text tokens plus $12 per million audio output tokens — which lands close to $0.015 per generated minute with more control over voice and style.
Speech-to-text with Whisper is billed per audio minute at $0.006, rounded to the nearest second, so short files don't carry a per-request penalty the way minute-rounded providers do. At scale, the cost lever is volume: 1,000 hours of audio is $360 with Whisper before you compare against self-hosted or alternative providers.
Reference rates
| Model | Billing unit | List price | ≈ per audio minute |
|---|---|---|---|
| TTS (tts-1) | character | $15 / 1M | ~$0.015 |
| TTS HD (tts-1-hd) | character | $30 / 1M | ~$0.030 |
| gpt-4o-mini-tts | tokens (in + audio out) | $0.60 / $12 per 1M | ~$0.015 |
| Whisper (STT) | audio minute | $0.006 | $0.006 |
List rates as of late 2026, checked against OpenAI's pricing page — confirm current rates before committing a budget.
Methodology: TTS minutes estimated at 1,000 characters/minute (~150 wpm); Whisper billed per audio minute to the nearest second. Byline: KickLLM research. Last reviewed 2026-09-29.
Cost levers
- Match model to purpose. TTS HD only pays off for customer-facing audio; internal narration at standard TTS halves the bill.
- Batch STT for throughput, not price. OpenAI bills Whisper to the nearest second, so short clips no longer carry a per-request penalty — batching is about API efficiency, not cost.
- Cache repeated phrases. IVR prompts and repeated UI lines can be synthesized once and reused forever.