Deepgram pricing starts at $0.0048 per minute for Nova-3 Monolingual streaming, with separate rates for multilingual transcription, batch processing, text-to-speech, and voice agents. But the per-minute rate is only part of the cost. Add-ons, LLM usage, TTS, telephony, and engineering can all increase your total spend.
For enterprises, the bigger question is whether you're pricing speech technology or the full workflow around it. Deepgram is one component of a voice stack, while platforms like HappyRobot combine AI workers, orchestration, integrations, and other infrastructure to run operational workflows end to end.
This guide breaks down Deepgram's current pricing and real-world costs.
NB: Pricing verified August 2026 and subject to change.
How Deepgram Pricing Works
Deepgram uses usage-based pricing, with Pay As You Go requiring no minimum commitment and Growth starting at $4,000+ per year with prepaid credits and discounts of up to 20%. Enterprise pricing is available through sales.

Deepgram bills speech-to-text by the actual audio duration rather than rounding every request to a full minute. Its pricing page states that a 14-second file costs for 14 seconds of usage.
The $200 introductory credit applies to new accounts and does not expire until used. Growth customers receive prepaid credits, while Enterprise customers can negotiate terms around volume, deployment, support, and other requirements.
Deepgram Pricing Per Minute by Model
The table below shows the current Deepgram rates as of August 2026. Streaming rates marked as promotional are subject to change, 2026; pre-recorded rates are separate.
| Model | Streaming Rate | Batch / Pre-Recorded Rate | Best Suited For |
|---|---|---|---|
| Nova-3 Monolingual | $0.0048/min PAYG; $0.0042/min Growth (promotional) | $0.0043/min PAYG; $0.0036/min Growth | General-purpose transcription |
| Nova-3 Multilingual | $0.0058/min PAYG; $0.0050/min Growth (promotional) | $0.0052/min PAYG; $0.0043/min Growth | Multilingual transcription and automatic language detection |
| Flux English | $0.0065/min PAYG; $0.0057/min Growth (promotional) | N/A | Real-time English voice agents with built-in turn detection |
| Flux Multilingual | $0.0078/min PAYG; $0.0068/min Growth | N/A | Real-time multilingual voice agents |
| Aura-2 TTS | N/A | $0.030 per 1K characters PAYG; $0.027 per 1K characters Growth | Higher-quality text-to-speech |
| Aura-1 TTS | N/A | $0.015 per 1K characters PAYG; $0.0135 per 1K characters Growth | Lower-cost text-to-speech |
| Voice Agent API — Standard | $0.056/min through Sept. 12; $0.075/min thereafter PAYG | N/A | Managed real-time voice agents |
| Voice Agent API — Standard, BYO TTS | $0.065/min PAYG; $0.051/min Growth | N/A | Voice agents using your own TTS |
| Voice Agent API — Custom, BYO LLM | $0.050/min through Sept. 12; $0.065/min thereafter PAYG; Growth $0.059/min | N/A | Voice agents using your own LLM |
| Voice Agent API — Custom, BYO LLM + TTS | $0.050/min PAYG; $0.041/min Growth | N/A | Maximum control over model components |
| Voice Agent API — Advanced | $0.122/min through Sept. 12; $0.163/min thereafter PAYG; Growth $0.146/min | N/A | Advanced production voice-agent requirements |
| Voice Agent API — Advanced, BYO TTS | $0.122/min PAYG; $0.110/min Growth | N/A | Advanced agents using your own TTS |
Nova-3 is Deepgram's general-purpose STT family, while Flux is designed specifically for conversational voice agents and includes model-integrated turn detection. Aura-1 and Aura-2 handle text-to-speech and are priced by characters rather than audio minutes.
The Voice Agent API has Standard and Advanced tiers, with a Custom tier for bringing your own LLM. BYO options materially change the rate: Deepgram lists separate pricing for BYO TTS and BYO LLM + TTS, so you should not treat the headline Voice Agent API rate as the price for every configuration.
What the Deepgram Add-ons Cost
Deepgram charges separately for several STT features, including redaction, keyterm prompting, entity detection, and speaker diarization. Smart Formatting is included.

The exact add-on rates depend on whether you are using streaming or pre-recorded audio, but the current published rates for the main streaming add-ons on Pay As You Go are:
- $0.0020/min for redaction
- $0.0013/min for keyterm prompting
- $0.0017/min for entity detection
- $0.0020/min for speaker diarization
For example, suppose you process 10,000 minutes of Nova-3 Monolingual streaming and need redaction, keyterm prompting, and speaker diarization:
| Cost Component | Rate | 10,000-minute cost |
|---|---|---|
| Nova-3 Monolingual | $0.0048/min | $48.00 |
| Redaction | $0.0020/min | $20.00 |
| Keyterm Prompting | $0.0013/min | $13.00 |
| Speaker Diarization | $0.0020/min | $20.00 |
| Total | $0.0101/min effective | $101.00 |
That makes the effective STT cost $0.0101 per minute, more than twice the base Nova-3 rate. The same pattern applies at larger volumes: an apparently small add-on can materially change the effective per-minute price when it is used on every call.
Free Credits and What they Actually Get You
New Deepgram accounts receive $200 in free credits. At the current $0.0048/min Nova-3 Monolingual streaming rate, that equals approximately 41,667 minutes:
$200 ÷ $0.0048 = 41,667 minutes, or about 694 hours of transcription.
At the $0.0043/min pre-recorded rate, the same credit covers approximately 46,512 minutes, or 775 hours. Add-ons reduce those totals because they consume additional usage credit.
Deepgram Pricing Compared to Alternatives
There is no single apples-to-apples rate across speech APIs because providers bundle features and offer different streaming and batch products. The table provides a rate-card comparison rather than a quality ranking.
| Provider | Streaming Rate | Batch Rate | Free Tier | Notable Trade-off |
|---|---|---|---|---|
| Deepgram | From $0.0048/min for Nova-3 Monolingual | From $0.0043/min | $200 credit | Add-ons such as diarization, redaction, and keyterm prompting can increase the effective STT cost. |
| AssemblyAI | $0.0075/min for Universal-3.5 Pro Realtime | $0.0035/min for Universal-3.5 Pro | $50 credit | Streaming and feature pricing differ from async; realtime add-ons are charged only when used. |
| OpenAI Whisper API | N/A | $0.006/min | No dedicated free tier | Whisper is not a streaming STT model |
| Google Cloud Speech-to-Text | From $0.016/min for standard V2 recognition | From $0.003/min for Dynamic Batch | 60 min/month | Cloud billing, tiers, and request-level pricing can complicate comparisons. cloud |
| ElevenLabs Scribe | $0.0065/min for Scribe v2 Realtime | $0.00367/min for Scribe v2 | 10,000 credits | Credits are shared across ElevenLabs products, so TTS or other usage can reduce available STT capacity |
The rates above are normalized from each provider's published pricing and are not a claim that the products provide identical functionality.
- AssemblyAI publishes Universal-3.5 Pro Realtime at $0.45/hour and async at $0.21/hour
- OpenAI lists Whisper at $0.006/minute
- Google lists standard V2 recognition from $0.016/minute and dynamic batch from $0.003/minute
- ElevenLabs lists Scribe v2 at $0.22/hour and Scribe v2 Realtime at $0.39/hour.
What a Full Voice Deployment Actually Costs
A production voice agent costs more than its STT provider because STT is only one layer of the stack. A realistic cost model includes speech recognition, LLM inference, text-to-speech, telephony, orchestration, and the engineering required to make those pieces work reliably together.
For example, imagine 10,000 minutes of monthly voice traffic using Nova-3 at $0.0048/min. The STT line is only $48 per month before add-ons. If you add redaction, keyterm prompting, and diarization, the STT total rises to $101.
The rest of the bill comes from other layers:
- LLM inference: Every conversation requires model processing, with cost depending on the model, token volume, context length, and response pattern.
- TTS: Generated speech is separately priced by characters when you use Deepgram's Aura models.
- Telephony: Phone numbers, inbound and outbound calling, carrier charges, and recording can sit outside the AI API bill.
- Orchestration: Your application needs to manage state, tools, routing, authentication, retries, and integrations.
- Engineering: Turn detection, interruption handling, failover, latency monitoring, testing, evaluation, logging, and production debugging all require engineering time.
This is why comparing voice deployments using only an STT price per minute can give a misleading picture. The API may be inexpensive while the system built around it is not.
When to Buy the Components and When to Buy the Platform
Build directly with APIs when you already have a voice engineering team, have a relatively narrow use case, and want control over individual models and infrastructure.
Building makes sense when you:
- Have engineers who can own the voice stack.
- Need deep control over models and routing.
- Have a use case that does not change frequently.
- Are comfortable operating the infrastructure yourself.
A platform makes more sense when you:
- Need production voice agents across different accents, audio conditions, and workflows.
- Need reliability, monitoring, compliance, and integrations without building each layer yourself.
- Want to move from pilot to production in weeks rather than months.
- Care about the cost and operational complexity of the whole deployment, not just the STT line item.
For teams that want a managed workforce rather than assembling the individual components, HappyRobot's Voice AI platform provides the broader deployment layer around voice agents.




