Developer's Guide: Parsing & Manipulating SRT, WebVTT & JSON Subtitles in TypeScript & Python

By Akash Singh Solanki (Founder & Lead Systems Architect) • Published on August 2026 • 14 min read read • Category: Engineering

A practical guide for software engineers building subtitle parsers, video editors, and audio synchronization tools. Includes production-ready TypeScript and Python parsers, timecode converters, and regex patterns.

1. The Universal Subtitle Data Model

Whether parsing SubRip (.SRT), WebVTT (.VTT), or modern JSON transcripts, all time-synchronized caption structures share a foundational data model:

export interface SubtitleCue {
  id?: string | number;
  startTime: number; // Start timestamp in milliseconds
  endTime: number;   // End timestamp in milliseconds
  text: string;      // Rendered caption text
  speaker?: string;  // Optional speaker attribution
}

2. Building a Robust TypeScript SRT Parser

Here is a complete, zero-dependency TypeScript implementation for parsing SRT subtitle blocks into structured cue objects (compare syntax differences in our SRT vs. WebVTT Format Guide):

export function parseSrt(srtContent: string): SubtitleCue[] {
  const cues: SubtitleCue[] = [];
  const normalized = srtContent.replace(/\r\n|\r/g, '\n').trim();
  const blocks = normalized.split(/\n\n+/);

  const timecodeRegex = /^(\d{2}):(\d{2}):(\d{2})[,.](\d{3})\s*-->\s*(\d{2}):(\d{2}):(\d{2})[,.](\d{3})/;

  for (const block of blocks) {
    const lines = block.split('\n');
    if (lines.length < 2) continue;

    let timeLineIdx = 0;
    let cueId: string | undefined;

    // Check if first line is a numeric ID
    if (/^\d+$/.test(lines[0].trim())) {
      cueId = lines[0].trim();
      timeLineIdx = 1;
    }

    const timeLine = lines[timeLineIdx];
    const match = timeLine.match(timecodeRegex);
    if (!match) continue;

    const startMs = parseTimeToMs(match[1], match[2], match[3], match[4]);
    const endMs = parseTimeToMs(match[5], match[6], match[7], match[8]);
    const text = lines.slice(timeLineIdx + 1).join('\n').trim();

    cues.push({ id: cueId, startTime: startMs, endTime: endMs, text });
  }

  return cues;
}

function parseTimeToMs(hh: string, mm: string, ss: string, ms: string): number {
  return (
    parseInt(hh, 10) * 3600000 +
    parseInt(mm, 10) * 60000 +
    parseInt(ss, 10) * 1000 +
    parseInt(ms, 10)
  );
}

3. Building a WebVTT Parser & Validator in Python

Below is a clean Python 3.11 implementation for parsing and formatting WebVTT files:

import re
from typing import List, Dict, Any

def parse_webvtt(vtt_text: str) -> List[Dict[str, Any]]:
    lines = [l.strip() for l in vtt_text.strip().splitlines() if l.strip()]
    if not lines or not lines[0].startswith("WEBVTT"):
        raise ValueError("Invalid WebVTT: Missing WEBVTT header")

    cues = []
    tc_pattern = re.compile(r"(?:(\d{2}):)?(\d{2}):(\d{2})\.(\d{3})\s*-->\s*(?:(\d{2}):)?(\d{2}):(\d{2})\.(\d{3})")

    i = 1
    while i < len(lines):
        match = tc_pattern.match(lines[i])
        if match:
            start_ms = time_to_ms(*match.groups()[:4])
            end_ms = time_to_ms(*match.groups()[4:])
            text_lines = []
            i += 1
            while i < len(lines) and not tc_pattern.match(lines[i]):
                text_lines.append(lines[i])
                i += 1
            cues.append({
                "start": start_ms,
                "end": end_ms,
                "text": " ".join(text_lines)
            })
        else:
            i += 1
    return cues

def time_to_ms(hh, mm, ss, ms) -> int:
    h = int(hh) if hh else 0
    return h * 3600000 + int(mm) * 60000 + int(ss) * 1000 + int(ms)

4. Millisecond Timecode Math & Drift Correction

When stitching audio chunks together or correcting drift caused by frame rate conversions (e.g. 23.976 fps to 29.97 fps), working with pure integer milliseconds avoids floating-point precision loss. For audio format container details, consult our Audio Formats, Codecs & Containers Guide.


5. Standardizing Word-Level JSON Transcripts

Modern web video applications (like custom audio waveforms and interactive karaoke players) require word-level precision. TranscriptG exports structured JSON following this standard (learn how to build vector indexes with this schema in our Audio Archives & Semantic Search Guide):

{
  "durationSeconds": 142.5,
  "language": "en",
  "segments": [
    {
      "id": 1,
      "start": 0.12,
      "end": 3.48,
      "speaker": "Speaker 1",
      "text": "Welcome to TranscriptG's neural transcription platform.",
      "words": [
        { "word": "Welcome", "start": 0.12, "end": 0.65, "confidence": 0.99 },
        { "word": "to", "start": 0.68, "end": 0.82, "confidence": 0.99 },
        { "word": "TranscriptG", "start": 0.85, "end": 1.45, "confidence": 0.98 }
      ]
    }
  ]
}

6. Integrating with TranscriptG's Ephemeral API

Developers can integrate with TranscriptG to generate frame-accurate SRT, VTT, and JSON exports instantly with zero data persistence overhead. Try out instantaneous format conversions on our interactive Subtitle Converter Tool or test our speech models via the Free Speech Transcriber.

Frequently Asked Questions

Why use integer milliseconds instead of float seconds?

Floating-point numbers suffer from rounding errors during subtitle math. Storing timestamps as integer milliseconds guarantees exact synchronization.

How do I handle both comma and period millisecond delimiters?

Use a regular expression like '[,.]' that matches both commas (SRT standard) and periods (VTT standard).

Can I export word-level timestamps in TranscriptG?

Yes. TranscriptG exports structured JSON containing start and end timestamps for every individual spoken word.

TranscriptG Engineering Lab Navigation