LISTENING WORKBENCH · 14 SEP 2026

Listen first. Reveal the voice later.

Compare each candidate’s manifest reference with its saved generated audio. Choose a language, declare your fluency, and rate pronunciation, naturalness and speaker similarity separately.

30 existing cases · 3 voices · 10 languages · 29 successful outputs · 1 retained failure. These technical outcomes are not quality scores.

Start a local listening session

Candidate labels stay consistent across languages. Starting again clears the current scores and audio. Save your export first. Reveal history remains recorded for this page visit.

Observed timing, successes only

Successful sample size
n = 29
Range
58.990133.997 s
Median
71.457 s
p95
129.797 s

Nearest rank for both median and p95: sort ascending; choose the value at one-based rank ceil(p × n), with p=0.50 and p=0.95. No interpolation or averaging. For n=29, median is rank 15 and p95 is rank 28.

Only succeeded rows from the 2026-09-14 manifest; n=29, China region, short prompts, serial requests. Wall time includes queueing, polling, network transport and download. Not pure inference time, not an SLA, not a load test or US-region performance evidence. The failed case includes a recovery pause and is reported separately.

Retained failure: FR · execution_recovered · 3150.496 s including recovery pause · $0.000000 charged. No generated audio or listening score exists for this failed case.

Reproduce and interpret a review

Save the local export. It includes the numeric seed, exact presentation order, input text hashes, expected generated-audio hashes, actual loaded-audio hashes, language declarations and whether identities were exposed when each score was entered. Null means “not rated”; a failure is never converted into a low score or omitted.

The manifest supplies same-preset English referenceUrl preview endpoints, not immutable original-reference files, transcripts or hashes. Current reference bytes are hashed locally when loaded; their identity with the historical original cannot be verified. Cross-language similarity may be affected by the different text and language.

The same seed recreates the candidate mapping and order; it does not prove the same person listened or that a live reference endpoint still returns the same recording. Playback events are browser observations, not proof of complete or attentive listening.

Randomization and checksum methods

Unsigned 32-bit seed; Mulberry32 PRNG; Fisher-Yates shuffle of manifest.plan.voices followed by Fisher-Yates shuffle of rows for each language in manifest.plan.languages order. One additional PRNG draw per row selects reference-first when < 0.5. No case is dropped. Seed and complete presentation order are exported.

SHA-256. Input text: exact UTF-8 without normalization. Dataset: UTF-8 JSON.stringify(parsed manifest), preserving parsed key and array order, without whitespace; not the raw manifest file hash. Generated audio: manifest hash checked against downloaded bytes before playback. Reference audio: hash of the exact downloaded bytes played, with no expected historical hash.

Existing recovery evidence

Existing first-party release report documents fencing before admission, reservation release and no replacement generation for the original failed case. Its separate later smoke test does not replace the failure or enter this dataset or timing summary. No recovery-duration or reliability metric is inferred.

Repository source: docs/voice-api-execution-recovery-release-2026-09-14.md. Its checksum and the existing public report link are included in the machine evidence and local export after reveal.

All evaluation inputs and outcomes come exclusively from /examples/evaluation-20260914/manifest.json. The separate release report provides recovery context only.