Software / Developer APIs / Speechmatics

Speechmatics

Speechmatics provides multilingual speech-to-text, text-to-speech and voice AI APIs for real-time, batch, cloud, on-premises and on-device deployments.

8.6/10 TechZella Score
Visit Website ↗
At a glance

Quick Verdict

Speechmatics offers batch and real-time transcription, 55-plus language support, diarization, translation, speech intelligence features, multiple deployment models and documented APIs. G2 lists a 4.7 out of 5 rating from 69 reviews, with reviewers citing integration speed and support.

Overview

Speechmatics is a business-focused voice AI platform founded in 2006 and headquartered in Cambridge, England. It provides speech-to-text, text-to-speech and voice agent technologies through cloud APIs and deployable software.

The platform supports real-time and batch transcription across more than 55 languages. Its documented capabilities include speaker and channel diarization, word-level timestamps, custom dictionaries, confidence scores, translation, summaries, sentiment analysis and topic detection.

Speechmatics supports cloud, on-premises and on-device deployments. On-premises options include Docker containers, Kubernetes and a virtual appliance. On-device deployment is documented for Mac and Windows environments.

Key Features

Speechmatics provides real-time transcription through streaming connections and batch transcription for uploaded audio or video files. Real-time transcripts can begin arriving within milliseconds, while final output improves as additional context becomes available.

  • Speech-to-text for live streams and prerecorded media
  • Support for more than 55 languages and dialects
  • Speaker and channel diarization
  • Word-level timestamps and confidence scores
  • Custom dictionaries for technical terms, acronyms and proper nouns
  • Automatic language identification and multilingual processing
  • Translation to and from English for more than 30 languages
  • Numeral formatting, punctuation, casing, profanity and disfluency handling
  • Summaries, sentiment analysis, topics and chapters
  • REST and WebSocket APIs
  • Python and Node.js SDKs
  • Integrations with LiveKit, Pipecat, Vapi and NVIDIA Holoscan

Pricing

Speechmatics offers a free entry tier with $100 in usage credit and no credit card requirement. The free tier includes speech-to-text access, 55-plus languages, two concurrent real-time sessions and multi-region cloud options.

The Pro tier starts at $0.129 per hour. It includes 50 concurrent real-time sessions, up to 10 file jobs per second, online email support and the same initial $100 credit. Speechmatics also advertises a 20% discount option.

Enterprise pricing is quote-based. It adds volume discounts, custom models, unlimited scale, custom language development, SaaS or on-premises deployment, higher concurrency and prioritized support. Pricing can vary by model, feature and deployment.

Pros & Cons

Pros

  • Supports both real-time and batch transcription workflows.
  • Offers broad multilingual coverage with speaker diarization.
  • Provides cloud, on-premises and on-device deployment choices.
  • Includes documented REST, WebSocket and SDK-based integration paths.
  • Supports technical customization through custom dictionaries and metadata.

Cons

  • Enterprise deployment and custom models require a sales process.
  • Feature and deployment costs can vary beyond the base hourly rate.
  • The product is primarily designed for developers, platforms and enterprises.
  • On-device deployment is documented for Mac and Windows, limiting stated desktop coverage.

Alternatives

Comparable speech-to-text and voice AI services include Deepgram, AssemblyAI, Google Cloud Speech-to-Text, Amazon Transcribe, Microsoft Azure AI Speech, OpenAI speech-to-text and ElevenLabs Scribe. Speechmatics is particularly relevant when multilingual support, diarization and controlled deployment are priorities.

FAQ

Does Speechmatics offer a free trial?

Yes. The free tier provides $100 in usage credit without requiring a credit card. The pricing page describes this as free access for developers and early exploration.

Does Speechmatics have an API?

Yes. Speechmatics documents REST and WebSocket APIs for batch and real-time transcription. It also provides Python and Node.js SDKs.

What integrations does Speechmatics support?

Official documentation lists native integrations with LiveKit, Pipecat, Vapi and NVIDIA Holoscan. The platform also provides SDKs and API references for custom integrations.

Which platforms support Speechmatics?

Speechmatics is available through cloud APIs and can be deployed on-premises using Docker, Kubernetes or a virtual appliance. Its on-device documentation covers Mac and Windows.

How many languages does Speechmatics support?

Speechmatics advertises support for more than 55 languages. Its Melia model is designed for multilingual audio and code-switching across supported languages.

Can Speechmatics identify speakers?

Yes. Speaker and channel diarization are available for batch and real-time transcription. The company states that real-time streams can identify up to 50 speakers by default, with higher limits available.

Capabilities

Features

Speech Processing

  • Real-time transcription
  • Batch transcription
  • 55-plus supported languages
  • Automatic language identification
  • Multilingual and code-switched speech processing

Transcript Intelligence

  • Speaker diarization
  • Channel diarization
  • Word-level timestamps
  • Confidence scores
  • Custom dictionaries
  • Numeral formatting
  • Punctuation and casing
  • Profanity and disfluency detection

Speech Understanding

  • Translation
  • Summarization
  • Sentiment analysis
  • Topic detection
  • Chapters

Developer Access

  • REST API
  • WebSocket API
  • Python SDK
  • Node.js SDK
  • API keys and portal management

Integrations

  • LiveKit
  • Pipecat
  • Vapi
  • NVIDIA Holoscan

Deployment

  • Cloud SaaS
  • On-premises deployment
  • Docker containers
  • Kubernetes
  • Virtual appliance
  • On-device deployment for Mac and Windows
Product details

Specifications

Product typeB2B voice AI software and APIs
Primary capabilityAutomatic speech recognition
ProcessingReal-time and batch
Language coverage55+ languages
API stylesREST and WebSocket
SDKsPython and Node.js
DeploymentCloud, on-premises and on-device
On-device systemsMac and Windows
Free trial$100 credit, no card required
Founded2006
HeadquartersCambridge, England, United Kingdom
Visual preview

Demo & Screenshots

Screenshots are not available yet.TechZella will add verified product images when suitable official screenshots are found.
Our evidence-based assessment

TechZella Score

A proprietary editorial score based on product capabilities, usability, value, performance, support and user sentiment evidence.

8.6/10Moderate confidence
Features9.0
Ease of Use8.4
Value for Money8.0
Performance8.8
Support8.5
User Sentiment8.6

Speechmatics offers batch and real-time transcription, 55-plus language support, diarization, translation, speech intelligence features, multiple deployment models and documented APIs. G2 lists a 4.7 out of 5 rating from 69 reviews, with reviewers citing integration speed and support.

Methodology v1.0. This is a TechZella editorial assessment, not a direct user-review average.

Independent review sources

Trusted Ratings

Ratings are published by the respective review platforms and may change over time.

Plans & pricing

Pricing

Free

$100 credit
usage credit

No credit card required. Includes speech-to-text access, 55-plus languages, two concurrent real-time sessions and multi-region cloud options.

Check current pricing →

Pro

From $0.129 per hour
usage-based

Includes 50 concurrent real-time sessions, 10 file jobs per second, online email support and an initial $100 credit.

Check current pricing →

Enterprise

Contact sales
custom

Volume discounts, custom models, flexible cloud or on-premises deployment, custom language development and prioritized support.

Check current pricing →

Pricing may change. Verify current plans on the vendor's website.

Editorial assessment

Pros & Cons

Pros

  • Real-time and batch transcription are available through the same product family.
  • Broad language coverage supports multilingual applications.
  • Speaker diarization, timestamps and confidence scores support downstream processing.
  • Cloud, on-premises and on-device deployment options address different privacy requirements.
  • REST, WebSocket and SDK access supports custom application development.

Cons

  • Enterprise and on-premises requirements may require contacting sales.
  • Total pricing can vary by model, add-on feature and deployment method.
  • The product is oriented toward developers, platforms and enterprise teams.
  • Documented on-device support currently focuses on Mac and Windows.
Similar software

Alternatives

Compare options

Top Competitors

Community feedback

Reviews

No reviews yet. Be the first to share your experience.

Write a Review