Starter
Pay-as-you-go processing with 10 free hours per month. Audio intelligence features are included according to the current plan terms.
Check current pricing →Gladia is an audio AI API for real-time and pre-recorded transcription, translation, diarization, summarization and conversation analysis.
Gladia provides pre-recorded and real-time APIs, multilingual transcription, code switching, diarization, translation, summaries, entity extraction, sentiment analysis and SDKs. G2 reviewers generally praise accuracy, speed and API usability, while some note documentation, latency and high-volume co
Gladia is an audio AI infrastructure service for developers building voice-enabled products. Its API processes live streams and recorded audio or video for transcription and downstream language analysis.
The platform supports more than 100 languages and includes automatic language detection and code switching. Typical applications include meeting assistants, contact-center systems, sales tools, live captions and voice agents.
Gladia exposes REST endpoints for pre-recorded jobs and WebSocket connectivity for live transcription. It also provides JavaScript and Python SDKs, plus integrations with voice and meeting infrastructure providers.
Gladia’s pre-recorded API accepts audio URLs or uploaded files and returns structured transcription results. Developers can configure callbacks or webhooks instead of polling for completion.
The live API streams audio through WebSocket connections. It supports partial and final transcript events, configurable audio formats, language settings, code switching and real-time processing options.
Available processing features include speaker diarization, translation, custom vocabulary, custom spelling, subtitles, punctuation enhancement, named entity recognition, sentiment analysis, summarization, audio-to-LLM prompts and personally identifiable information redaction.
Documented integration partners include Pipecat, LiveKit, Vapi, Recall, MeetingBaaS, Twilio, VideoSDK, Composio, Zapier, Make and n8n.
Gladia charges according to audio duration. The Starter tier costs $0.61 per hour for asynchronous processing and $0.75 per hour for real-time processing. It includes 10 free hours each month.
The Growth tier uses upfront volume commitments. Published starting rates are as low as $0.20 per hour for asynchronous processing and $0.25 per hour for real-time processing.
Enterprise pricing is customized and can include annual commitments, dedicated support, custom models, data residency options and zero data retention. Buyers should confirm final rates and feature terms directly with Gladia.
Pros
Cons
Comparable products include Deepgram, AssemblyAI, Google Cloud Speech-to-Text, Amazon Transcribe, Microsoft Azure AI Speech and Rev AI. These services also provide developer-focused speech recognition APIs, although language coverage, analysis features, pricing and deployment options differ.
Deepgram and AssemblyAI are particularly direct comparisons for application developers seeking hosted speech recognition. Cloud Speech-to-Text services from Google, Amazon and Microsoft may suit organizations already standardized on those cloud ecosystems.
Does Gladia offer a free trial?
Yes. The Starter tier includes 10 free hours of audio processing per month.
Does Gladia provide an API?
Yes. Gladia provides REST APIs for pre-recorded transcription and a WebSocket API for real-time transcription.
Which programming languages have official SDKs?
Gladia documents JavaScript or TypeScript and Python SDKs.
Does Gladia support integrations?
Yes. Documented partners include LiveKit, Pipecat, Vapi, Recall, MeetingBaaS, Twilio, VideoSDK, Composio, Zapier, Make and n8n.
What file types does Gladia accept?
Supported formats include AAC, AC3, FLAC, M4A, MP3, OGG, Opus and WAV, plus common video formats such as MP4, MOV, AVI and WebM-compatible containers.
Can Gladia identify speakers?
Yes. Speaker diarization is available for pre-recorded transcription workflows.
Does Gladia support live transcription?
Yes. The live API uses WebSocket streaming and returns transcript events during the session.
| Primary delivery | Cloud API |
|---|---|
| API styles | REST and WebSocket |
| SDKs | JavaScript/TypeScript and Python |
| Processing modes | Pre-recorded and real-time |
| Language support | 100+ languages |
| Pre-recorded request limit | 135 minutes per request |
| Real-time session limit | 3 hours per WebSocket session |
| Free trial | 10 hours per month on Starter |
| Billing | Per audio hour |
| Headquarters | New York, United States |
| Company founded | 2022 |
A proprietary editorial score based on product capabilities, usability, value, performance, support and user sentiment evidence.
Gladia provides pre-recorded and real-time APIs, multilingual transcription, code switching, diarization, translation, summaries, entity extraction, sentiment analysis and SDKs. G2 reviewers generally praise accuracy, speed and API usability, while some note documentation, latency and high-volume cost concerns.
Methodology v1.0. This is a TechZella editorial assessment, not a direct user-review average.
Pay-as-you-go processing with 10 free hours per month. Audio intelligence features are included according to the current plan terms.
Check current pricing →Volume-based pricing with upfront usage commitments. Published rates vary by commitment level.
Check current pricing →Custom pricing with enterprise support, deployment, data retention and data residency options.
Check current pricing →Pricing may change. Verify current plans on the vendor's website.
No reviews yet. Be the first to share your experience.
Please log in to submit a review.