Definition

What Is Transcription?

VideoText editorial team. Updated September 13, 2026.

Direct definition

Transcription is the conversion of spoken words in audio or video into written text. A transcript may be plain text or include timestamps, speaker labels, and notes about inaudible or overlapping speech.

The sections below explain how transcription is used in transcription and subtitle work, what people often mix up, and which VideoText workflow applies when the next step is a file or review pass.

Key takeaways

  • Transcription produces a readable text record of speech — not necessarily timed captions.
  • Verbatim and clean-verbatim transcripts follow different style rules.
  • Automatic speech recognition can draft a transcript; humans still review names, overlap, and formatting.

How transcription works

A transcriber or a speech-recognition system listens to an audio track, decides what words were spoken, and writes them in order. Optional extras include speaker labels, timestamps, and tags such as [inaudible].

In a video workflow the audio is usually extracted first, then recognized, then formatted. VideoText’s video-to-transcript path follows that order: extract audio, run speech recognition, then offer a transcript plus optional subtitle files.

Why transcription matters

A transcript makes speech searchable, quotable, and readable without playing the media. It is also the source file many teams later turn into captions or translations.

Accessibility is one reason, not the only one. Journalists, researchers, students, and editors use transcripts to review what was said without scrubbing a timeline.

Example

A 40-minute interview exported as a timestamped transcript lets an editor jump to “00:18:12” instead of relistening to the whole file.

The thing people often get wrong

A transcript is not automatically a caption file. Captions and subtitles add timing and display rules so the text can appear on screen. You can make captions from a transcript, but the two deliverables are not the same.

Transcription vs speech-to-text

Speech-to-text is the technical process of mapping audio to words. Transcription is the resulting document and the editorial rules applied to it. Everyday usage overlaps; the distinction matters when a client specifies verbatim style or speaker labels.

Frequently asked questions

Is transcription the same as captioning?
No. Transcription is the text record. Captioning adds timing and on-screen presentation for viewers who cannot or do not hear the audio.
Can software replace a human transcript?
Automatic speech recognition can produce a strong first draft on clean audio. Names, overlap, punctuation, and style-guide rules still need a review pass for delivery-quality work.

Related terms

Sources

  1. W3C: Making Audio and Video Media Accessible
  2. WebAIM: Captions, Transcripts, and Audio Descriptions

Related VideoText workflow