Every small step opens a new door. Listen, learn, and tell your own story.
5.136s · 41,709 bytes · $0.000592FIRST-PARTY EVIDENCE · 14 SEP 2026
Hear the result. Inspect the evidence.
A reproducible look at one short-audio workflow: what we generated, what it cost, and what this test does not establish.
clone-v1 · Rowan · MP3 · speed 1
23successful stored generations
$0.004984actual generation charge
623billed characters
4sampled languages
Internal example, not a customer case or independent benchmark. These costs are the historical API charges, not our GPU costs.
Original text stays beside the audio.
每一个小小的进步,都会打开一扇新的门。认真聆听,不断学习,讲述属于你自己的故事。
7.896s · 63,213 bytes · $0.000320小さな一歩が、新しい扉を開きます。耳を傾け、学び、自分だけの物語を語りましょう。
7.416s · 59,373 bytes · $0.000320Method and reproducibility
- Use one operator-published Rowan preset: three comparison texts and 20 cards across en, zh, ja and es. Execute ordinary paid API jobs; no trial characters used.
- Save outputs with the original input, returned usage, size, duration and SHA-256. Decode all 23 MP3 files and check the hash-to-file association.
- Inspect 14 measured English word intervals. The player uses actual audio time, preserves gaps, and does not invent word boundaries.
- Run the complete saved 20-card project again. Matching local checksums skip API calls. Download and run the starter to inspect recovery behavior yourself.
What this report does not prove
- It does not measure service-wide success rate: this manifest records completed outputs, not every attempt or live traffic.
- One voice and four sampled languages do not verify every voice or all configured languages.
- Offline ASR detected the expected language, but some Chinese and Japanese transcripts differed from the input. No native-speaker review or universal naturalness/similarity score is claimed.
- No latency percentiles, load test, independent competitor comparison or SLA. A new generation may sound different and will not match these file hashes.
A practical acceptance checklist for your app
Use your own representative scripts and authorized voice. Compare pronunciation, names, numbers and similarity with your target users. Record request-to-download time separately from voice preparation and queue waiting. Save actual billed characters and repeat-generation counts when assessing cost.