Definition
VideoText editorial team. Updated September 13, 2026.
A speech recognition model is the trained statistical or neural system that maps audio features to tokens or words. Examples include the Whisper family of checkpoints. The model is not the same thing as the product UI around it.
The sections below explain how speech recognition model is used in transcription and subtitle work, what people often mix up, and which VideoText workflow applies when the next step is a file or review pass.
Fix the dataset, the decoding settings, and the text normalizer. Then compare WER. Changing any of those three invalidates a head-to-head claim.