Whisper (OpenAI)
Open-source under MIT. The five model sizes (tiny → large) plus the faster turbo variant give you a tradeoff dial: turbo for speed, large-v3 for max accuracy. 99 languages, strong on accents and code-switching, and it's the de facto baseline most other tools are measured against.
- Cost: $0 self-hosted (your compute), or ~$0.006/min via OpenAI API.
- Latency: 1× realtime on a single GPU for large-v3, faster with whisper.cpp / faster-whisper.
- Quality: very strong on English, strong on European languages, competitive on Asian.
- Weakness: no native diarization, no native word-level timestamps (need WhisperX).
AssemblyAI
Hosted-only API with strong English accuracy and built-in features Whisper lacks: speaker diarization, sentiment analysis, summarization, content moderation. Their universal-1 model is competitive with large-v3 on English benchmarks.
- Cost: ~$0.37/hour for transcription, more for add-ons.
- Latency: under 1× realtime.
- Quality: matches or beats Whisper on English; weaker than Whisper on long-tail languages.
- Strength: production features (diarization, PII redaction, moderation) bundled in.
Deepgram
Hosted API focused on real-time streaming. Their Nova-3 model is one of the fastest in the market and accurate on English with strong domain customization (medical, legal, conversational).
- Cost: ~$0.26/hour batch, more for streaming.
- Latency: streaming with sub-300ms first-word latency.
- Quality: top-tier English, including phone-quality and noisy audio.
- Strength: domain models you can fine-tune.
How to pick
If you're shipping a creator product on a budget, Whisper (self-hosted or via OpenAI API) is the default. If you're building a B2B product that needs diarization and PII redaction, AssemblyAI's bundled features save engineering time. If you're shipping real-time captioning (live events, calls, agents), Deepgram's streaming wins on latency.
