Definition

What Is Transcription Accuracy?

VideoText editorial team. Updated September 13, 2026.

Direct definition

Transcription accuracy is how closely a transcript matches what was said, under a stated definition of “match.” Researchers usually report word error rate. Clients may mean names, numbers, and style-guide compliance instead.

The sections below explain how transcription accuracy is used in transcription and subtitle work, what people often mix up, and which VideoText workflow applies when the next step is a file or review pass.

Key takeaways

  • Ask for the metric, the audio condition, and the sample.
  • Names and numbers can be “wrong” even when WER looks low.
  • A product accuracy figure is only useful when it names the dataset, sample, and metric — not as a standalone percentage.

How accuracy is judged

Research: align to a reference and compute WER or CER. Delivery work: a reviewer checks speakers, proper nouns, and the requested verbatim level.

Why condition matters

The Whisper paper’s Table 2 is the clearest public illustration: the same model’s WER ranges from a few percent on clean read speech to the mid-twenties or worse on noisy multi-speaker sets.

Related terms

Sources

  1. OpenAI: Robust Speech Recognition via Large-Scale Weak Supervision

Related VideoText workflow