OfficeBooks
Voice AI

Deepgram Review 2026: Developer-grade speech infrastructure with a fast Voice Agent API on its own STT and TTS models.

Voice AISpeech-to-TextDeveloper API

Affiliate disclosure: this review contains affiliate links — we may earn a commission if you sign up, at no cost to you. Ratings are our own editorial scores.

Deepgram screenshot
Our verdict

Deepgram

4.4
out of 5 · our rating

Pros

  • Owns its full speech stack (Nova-3 STT, Aura-2 TTS), giving strong accuracy and low latency
  • Transparent, granular usage-based pricing with a genuinely generous $200 no-card free credit
  • Voice Agent API unifies STT, LLM orchestration, and TTS with barge-in and turn-taking built in
  • Flexible deployment: cloud, single-tenant, in-VPC, self-hosted, and on-prem with HIPAA/GDPR support

Cons

  • Developer-first and API-only — no no-code builder, dashboard flows, or prebuilt agent templates
  • Add-ons (diarization, redaction, summarization) and per-channel stereo billing can multiply effective cost
  • Growth plan's $4,000/year floor risks overpayment for seasonal or low-volume users
  • Advanced Voice Agent tier ($0.163/min) is pricey versus the standard tier and BYO options

Best for: Developers building custom voice agents on a fast STT/TTS stack, Contact centers needing high-accuracy real-time transcription, Teams requiring self-hosted or in-VPC deployment for compliance.

What is Deepgram?

Deepgram is voice AI infrastructure aimed at developers. The homepage leads with the idea of voice AI built for real conversations, and the framing is deliberate: rather than selling transcription alone, Deepgram packages speech-to-text, text-to-speech and LLM orchestration behind one set of APIs so a team can ship a talking application without stitching three vendors together.

Deepgram builds its own models - Flux for conversational speech, Nova-3 for transcription and the Aura family for synthesis - alongside a hosted Whisper Large option. Flux is the conversation-aware layer, Nova-3 handles general recognition, and the Aura family covers synthesis. Owning that stack is what lets Deepgram argue on latency and cost at once, and it shows in the product menu: Speech-to-Text API, Text-to-Speech API, Voice Agent API and Audio Intelligence API, each usable alone or together.

Transcription built for production audio

Deepgram Speech-to-Text API product page showing Nova-3 and Flux STT models

The Speech-to-Text API is the foundation of the lineup, powering everything from transcription and analytics to real-time voice agents. Deepgram exposes real-time streaming and pre-recorded processing, with Nova-3 offered in monolingual and multilingual variants and Whisper Large available for batch work. The vendor reports more than 50 supported languages and streaming latency under 300ms.

What makes the service practical rather than merely accurate is the option list around each model. Smart formatting is included at no extra cost, while redaction, keyterm prompting, entity detection and speaker diarization are metered add-ons billed per minute on top of the base model rate. A contact centre team can enable speaker labels and PII redaction; a medical scribe product can push domain vocabulary through keyterm prompting instead of retraining. Custom model training and industry-tuned models cover audio that never resembles a benchmark set.

Voice agents on a single unified API

Deepgram Voice Agent API page describing unified conversational AI building blocks

The Voice Agent API is where Deepgram is pushing hardest, bundling speech-to-text, LLM orchestration and text-to-speech into one real-time connection. The vendor describes it as a unified conversational AI API for building enterprise-ready voice agents, and the features that matter on a live call are there: barge-in detection so a caller can interrupt, turn-taking prediction so the agent knows when a sentence has actually ended, function calling for tool use, and mid-session control for changing behaviour partway through.

Deepgram also lets you keep pieces you already like. BYO LLM and BYO TTS are supported configurations, and agent settings expose model choices directly, with Flux-General-EN as the listen model and Flux-Kit-EN as the speak model alongside third-party reasoning options such as Gemini-3.1-Flash-Lite.

Synthesis and audio intelligence

On the output side, Aura-1 and Aura-2 are the established text-to-speech models, joined more recently by Flux TTS for conversational work. The Audio Intelligence API layers understanding on top of transcripts through summarization, topic detection, sentiment analysis and intent recognition, covering post-call analytics without a second vendor.

Deployment options and plan structure

Deepgram runs in the cloud or self-hosted, and lists SOC 2 Type 1 and Type 2 certification, HIPAA compliance with BAAs signed for Enterprise customers handling ePHI, GDPR readiness with a dedicated EU endpoint at api.eu.deepgram.com, plus CCPA and PCI compliance among its controls. Buyers who cannot route audio through the shared cloud endpoint can use the dedicated EU endpoint or run the models self-hosted.

Commercially the ladder is short. Pay As You Go starts with a free credit and no card required, Growth is an annual pre-paid commitment that discounts per-unit rates across STT, TTS and Voice Agent usage, and Enterprise is quote-only. Speech-to-text and Voice Agent usage meters per minute, text-to-speech per 1,000 characters and Audio Intelligence per 1,000 tokens, so a pilot runs on the $200 free credit while Growth starts at $4K+ per year and only pays off at volume.

Who should choose Deepgram

Deepgram suits engineering teams embedding voice into a product - the vendor names contact centres, medical transcription, conversational AI, speech analytics and media transcription - and especially anyone assembling a real-time agent who would rather call one Voice Agent API than orchestrate three services. The deployment options and per-request add-ons make it particularly strong for regulated industries and for products with unusual vocabulary.

It is a poor fit for non-technical buyers wanting a finished application, because what ships here is an API and a key, not a dashboard you hand to an operations team. Businesses needing occasional transcription of a few files will find a consumer app simpler and cheaper than a metered developer platform.

Key features

FeatureWhat it does
Voice Agent APISingle API that orchestrates speech-to-text, an LLM, and text-to-speech in real time, with barge-in detection, turn-taking prediction, and function calling.
Nova-3 speech-to-textDeepgram's flagship STT model for streaming and pre-recorded audio, with monolingual and multilingual options tuned for accuracy and speed.
Aura-2 text-to-speechLow-latency neural voices priced per 1k characters, designed for natural, real-time conversational output.
Bring-your-own LLM/TTSUse Deepgram's orchestration while swapping in your own language model or TTS provider to cut cost or meet specific requirements.

Deepgram pricing

PlanPriceIncluded
Free credit$200 creditNo credit card required; credit does not expire. Good for meaningful evaluation.
Pay As You GoPOPULARUsage-basedNo minimums. Voice Agent from ~$0.075/min ($4.50/hr) full-stack; STT (Nova-3) from ~$0.0048/min streaming; TTS (Aura-2) $0.030/1k chars. BYO-LLM/TTS agent tiers drop to ~$0.041–0.065/min.
Growth$4,000+/yearPrepaid annual credits for ~16–20% lower rates, higher concurrency, and email support.
EnterpriseCustom (quote-only)Negotiated volume rates, unlimited concurrency, dedicated SLAs, single-tenant/VPC/on-prem, HIPAA/GDPR. Pricing not published.

How Deepgram compares

AlternativeHow it differs
VapiNo-code/low-code voice agent orchestration that stitches third-party STT/LLM/TTS; faster to launch but not a first-party model owner.
Retell AITurnkey conversational voice agent platform aimed at contact centers, with more built-in call tooling but less model-level control.
ElevenLabsBest-in-class TTS with a Conversational AI product; stronger voice quality, weaker STT and self-hosting story than Deepgram.

Deepgram ratings on other platforms

Independent user ratings from third-party review sites, linked here for transparency. These are not our editorial score, are captured on the date shown, and may have changed since.

Frequently asked questions

Does Deepgram have a free plan?

There is no perpetual free tier, but new accounts get a $200 credit with no credit card required and no expiration — enough for substantial testing before you pay.

How much does the Deepgram Voice Agent API cost?

The standard full-stack Voice Agent runs about $0.075/min ($4.50/hr) pay-as-you-go. Bring-your-own-LLM/TTS tiers fall to roughly $0.041–0.065/min, while the Advanced tier is $0.163/min.

Verdict

Deepgram is one of the strongest foundations in the Voice AI space because it owns its entire speech stack — Nova-3 STT and Aura-2 TTS — rather than reselling other vendors' models. Its Voice Agent API bundles that stack with LLM orchestration, barge-in, and turn-taking into a single fast, real-time interface, and pricing is transparent and usage-based with a generous $200 no-card credit. The trade-off is that it is unapologetically developer-first: you get an excellent API and flexible deployment (including self-hosted and in-VPC), but no no-code builder, and add-ons plus per-channel billing can push effective costs higher than the headline rate. For engineering teams building custom voice agents or high-accuracy transcription at scale, it is a top pick; teams wanting a plug-and-play agent will look elsewhere.

OB
OfficeBooks Editorial — Research desk

Our research desk checks every feature and price against the vendor’s own pricing page and dates each review when it was last checked. We do not run hands-on product tests — reviews are documentation-based, and third-party ratings are always attributed and dated.

Facts verified against: deepgram.com, deepgram.com, smallest.ai, www.ringlyn.com, deepgram.com, deepgram.com (as of August 2026).

Deepgram
Our rating 4.4/5 · $200 free credit, then usage-based
Visit →