TTS API SELECTION · PRIMARY SOURCES

Choose the API that fits the app.

Six providers. Different meters, limits and workflows. Start with your requirements, compare the actual text cost, then verify the sound.

Compiled by CastReader. Competitor facts checked · Official documentation comparison, not an independent audio benchmark.

A useful CastReader fit

Prepared learning clips, product guidance and short voiceovers with authorized cloned voices. Current deployment: 10 languages, 500 characters/job, concurrency 1.

When to choose another service

If streaming conversation, native multi-speaker orchestration, emotion controls or a contracted latency SLA are mandatory, do not select this CastReader offering on price alone. We do not currently provide those capabilities.

What the published offering actually says.

Prices are USD list rates, not a binding quote. Trials, taxes, discounts, top-up minimums and account terms may change the amount you pay. Unknown means verify—not unsupported.

CastReader

clone-v1
$8 / 1M normalized characters

Short, prepared multilingual audio and authorized voice cloning. Saved MP3/WAV; recoverable queued jobs.

Access, limits and source

Access: self-service; 3,000 trial characters. Verify identity; one grant per account, validity and daily limits apply.

Current configuration: en, zh, de, ja, fr, es, ko, pt, ru, it; 500 characters/queued job, concurrency 1. No streaming, realtime, emotion parameter or latency SLA.

ElevenLabs

Flash / Turbo
$50 / 1M characters

A candidate when faster delivery, more languages or longer individual inputs are required.

Access, limits and source

Pay-as-you-go is available; check free/startup allowances and account terms separately.

Official Flash/Turbo listing: 32 languages, 40,000-character input limit. v2/v3 have different prices and capabilities; not interchangeable benchmarks.

Fish Audio

s2.1-pro (paid)
$15 / 1M UTF-8 bytes

Official Python/TypeScript clients, cloning, REST and WebSocket streaming.

Access, limits and source

Pay-as-you-go. The official price list also includes s2.1-pro-free at $0; verify free-model access and limits before choosing paid usage.

Bytes are not characters. The pricing table lists starter concurrency 5; account/model-specific free limits and production terms require confirmation.

Deepgram

Aura-2
$30 / 1M characters (PAYG)

Consider alongside other speech-service options; evaluate the chosen voice and language in your own workload.

Access, limits and source

PAYG and Growth tiers have different rates. Free credits and other models are separate.

Aura-1 is also listed at $15/1M characters. Flux has a separate price/promotion; this row does not mix these models or establish clone parity.

Cartesia

Pro plan
$5/month · 100K credits

Evaluate when its voice/cloning workflow matches the application; compare plan utilization, not only the headline fee.

Access, limits and source

A Free plan exists. Pro includes instant voice cloning and commercial use in the published plan.

No character-equivalent estimate here: confirm model credit conversion, overage, term and account limits first.

OpenAI

gpt-4o-mini-tts
$0.60/1M text input tokens + $12/1M audio output tokens

Streaming speech and delivery instructions; a candidate when an existing application already uses OpenAI.

Access, limits and source

API access and rate limits depend on the account. Custom voices have a separate eligibility and consent process.

2,000 input tokens maximum on the model page. Output audio tokens are not known from character count alone; no fixed $/1M-character conversion.

SAME TEXT · DIFFERENT METERS

Price your text, not an imaginary million.

Runs in your browser. No text is uploaded and no audio is generated.

1–1,000,000 whole requests. This models repeating the text, not one oversized request.

All estimates use the same explicitly preprocessed text: CRLF → LF, NFC normalization, outer whitespace trimmed. Internal spaces and punctuation remain.

58Unicode code points / request
58UTF-8 bytes / request
CastReaderclone-v1$0.464
ElevenLabsFlash / Turbo$2.90
Fish Audios2.1-pro (paid)$0.87
DeepgramAura-2$1.74
CartesiaPro planNot comparable from text alone
OpenAIgpt-4o-mini-ttsNot comparable from text alone

Paid-rate illustration, not “lowest price.” Fish also lists s2.1-pro-free at $0; check its access and limits. Trial credits, subscriptions, discounts, taxes and minimum top-ups are not included.

ElevenLabs/Deepgram estimates assume one billed character per Unicode code point; verify their exact metering. Credits and output audio tokens require their own usage data. This is not equal-quality or equal-speed testing.

A lower rate is only one part of the decision.

Connect and recover

Install the exact published client, check credentials and the ready voice, then save the request and idempotency key before submitting. A timeout is not permission to create another billed job.

Run the connection check

Hear evidence, including failures

Our expanded first-party run records 30 attempts, 29 saved outputs and one failure. It is not a competing-model benchmark or a native-speaker verdict. Inspect the original reference, generated audio, input and usage.

Inspect the actual report
Total cost · not just synthesis
total_cost = successful_audio_usage
           + deliberate_regenerations
           + subscription_or_unused_minimums
           + app_hosting_storage_and_delivery
           + integration_and_recovery_work

// Measure complete-job time separately from first-audio time.
// Do not infer output audio tokens from text character count.
Build a complete learning-audio appReview real cross-language samplesCheck data and voice permissions