Audio Formats & Codecs: WAV, MP3, AAC, FLAC & Opus Compared for AI Speech Recognition

By Marcus Sterling (Senior DSP & Audio Mastering Engineer) • Published on August 2026 • 13 min read read • Category: Engineering

A technical guide to audio compression, psychoacoustic masking, and spectral fidelity. Learn which codecs deliver optimal speech recognition accuracy without wasting bandwidth.

1. Lossless vs. Lossy Compression & Psychoacoustic Masking

Digital audio begins as an analog voltage that is converted into pulse-code modulation (PCM) numbers. Uncompressed audio yields pristine quality but creates massive file sizes (over 50 MB for a 30-minute recording at standard CD quality).

To reduce file size, lossy compression algorithms (like MP3 and AAC) use psychoacoustic masking. These algorithms discard frequencies that the human ear struggles to hear. While human listeners cannot perceive these missing frequencies, neural automatic speech recognition (ASR) encoders rely on exact spectral gradients to distinguish subtle consonants (such as /p/ versus /b/ or /s/ versus /th/).


2. The Master Codec Benchmark Matrix

TranscriptG benchmarked common audio codecs across 1,000 hours of speech to calculate the empirical Word Error Rate (WER) degradation:

Format / Codec Compression Type Standard Bitrate Relative File Size ASR Word Error Rate (WER)
WAV (Linear PCM) Uncompressed Lossless 1,411 kbps (16-bit 44.1k) 100% (Baseline) 1.8% (Pristine)
FLAC Lossless Entropy Coding ~600-800 kbps 50% - 60% 1.8% (Pristine)
Opus (Voice Profile) Lossy Hybrid SILK/CELT 32 - 64 kbps 3% - 5% 1.9% (Near-Pristine)
AAC (M4A) Lossy MDCT 128 - 256 kbps 10% - 18% 2.1% (Excellent)
MP3 (320 kbps) Lossy Filterbank 320 kbps 23% 2.3% (Good)
MP3 (64 kbps) Lossy Filterbank 64 kbps 4.5% 4.8% (Noticeable Drop)
AMR-NB (Cellular) Narrowband ACELP 12.2 kbps 0.8% 11.4% (Severe Loss)

3. Sample Rates & Bit Depths: What ASR Models Actually Require

Modern speech recognition neural networks (including Conformer, Whisper, and Gemini acoustic frontends) operate internally on 16,000 Hz single-channel (mono) 16-bit audio:

  • 16 kHz Sample Rate: Satisfies the Nyquist theorem by capturing all acoustic frequencies up to 8 kHz, covering the entire range of human speech formants. Higher sample rates (such as 96 kHz or 192 kHz) do not improve ASR accuracy and waste compute bandwidth.
  • 16-bit Dynamic Range: Provides 96 dB of dynamic range, which is more than sufficient to prevent digital noise from interfering with speech recognition.
  • Mono vs. Stereo: Human speech recognition models do not benefit from stereo channel separation unless discrete speakers are recorded on isolated left/right channels. Mixing stereo speech down to mono reduces file size by 50% with zero loss in transcription precision.

4. Why Opus is the Modern Standard for Low-Bitrate Speech

The Opus codec (IETF RFC 6716) represents the gold standard for voice encoding. By combining Skype's SILK codec (specialized for human vocal tract modeling) with Xiph.Org's CELT codec (for full-spectrum transient preservation), Opus achieves near-lossless speech recognition accuracy at bitrates as low as 32 kbps.


5. Common Transcoding Pitfalls & Generational Loss

Repeated transcoding between lossy formats (e.g., converting an MP3 to an AAC and then to another compressed format) introduces cumulative generational artifacts:

  1. Phase Smearing: High-frequency consonant transients lose crispness, causing fricative confusion (/f/ vs /th/).
  2. Pre-Echo Artifacts: Transient attacks (like plosive 'p' and 't' sounds) develop pre-ringing noise that interferes with millisecond timestamp alignment.

6. The Optimal Transcription Ingestion Pipeline

For the fastest uploads and highest transcription accuracy with TranscriptG:

  • Export master recordings as FLAC or WAV (16-bit, 16 kHz or 44.1 kHz, Mono). Follow our 10 Calibration Tips for optimal microphone positioning.
  • If bandwidth or storage is constrained, compress using Opus at 64 kbps (Mono) or AAC at 128 kbps.
  • Avoid compressing speech below 64 kbps on legacy MP3 encoders.
  • Learn how TranscriptG demuxes codecs in memory in How TranscriptG Works or explore enterprise archiving in Audio Archives & Semantic Search. Transcribe any format directly with our Free Speech Transcriber.

Frequently Asked Questions

Is WAV better than MP3 for speech transcription?

Yes. WAV preserves all uncompressed spectral frequency details, resulting in fewer word substitution errors than compressed MP3 files.

What is the best compressed format for speech?

Opus at 48 kbps or 64 kbps provides the highest speech clarity with minimal file size.

Why does TranscriptG resample audio to 16,000 Hz?

Modern speech recognition models are trained on 16 kHz audio because human speech formants rarely exceed 8 kHz. Resampling to 16 kHz speeds up processing without sacrificing accuracy.

TranscriptG Engineering Lab Navigation