1. The Universal Subtitle Data Model
Whether parsing SubRip (.SRT), WebVTT (.VTT), or modern JSON transcripts, all time-synchronized caption structures share a foundational data model:
export interface SubtitleCue {
id?: string | number;
startTime: number; // Start timestamp in milliseconds
endTime: number; // End timestamp in milliseconds
text: string; // Rendered caption text
speaker?: string; // Optional speaker attribution
}
2. Building a Robust TypeScript SRT Parser
Here is a complete, zero-dependency TypeScript implementation for parsing SRT subtitle blocks into structured cue objects (compare syntax differences in our SRT vs. WebVTT Format Guide):
export function parseSrt(srtContent: string): SubtitleCue[] {
const cues: SubtitleCue[] = [];
const normalized = srtContent.replace(/\r\n|\r/g, '\n').trim();
const blocks = normalized.split(/\n\n+/);
const timecodeRegex = /^(\d{2}):(\d{2}):(\d{2})[,.](\d{3})\s*-->\s*(\d{2}):(\d{2}):(\d{2})[,.](\d{3})/;
for (const block of blocks) {
const lines = block.split('\n');
if (lines.length < 2) continue;
let timeLineIdx = 0;
let cueId: string | undefined;
// Check if first line is a numeric ID
if (/^\d+$/.test(lines[0].trim())) {
cueId = lines[0].trim();
timeLineIdx = 1;
}
const timeLine = lines[timeLineIdx];
const match = timeLine.match(timecodeRegex);
if (!match) continue;
const startMs = parseTimeToMs(match[1], match[2], match[3], match[4]);
const endMs = parseTimeToMs(match[5], match[6], match[7], match[8]);
const text = lines.slice(timeLineIdx + 1).join('\n').trim();
cues.push({ id: cueId, startTime: startMs, endTime: endMs, text });
}
return cues;
}
function parseTimeToMs(hh: string, mm: string, ss: string, ms: string): number {
return (
parseInt(hh, 10) * 3600000 +
parseInt(mm, 10) * 60000 +
parseInt(ss, 10) * 1000 +
parseInt(ms, 10)
);
}
3. Building a WebVTT Parser & Validator in Python
Below is a clean Python 3.11 implementation for parsing and formatting WebVTT files:
import re
from typing import List, Dict, Any
def parse_webvtt(vtt_text: str) -> List[Dict[str, Any]]:
lines = [l.strip() for l in vtt_text.strip().splitlines() if l.strip()]
if not lines or not lines[0].startswith("WEBVTT"):
raise ValueError("Invalid WebVTT: Missing WEBVTT header")
cues = []
tc_pattern = re.compile(r"(?:(\d{2}):)?(\d{2}):(\d{2})\.(\d{3})\s*-->\s*(?:(\d{2}):)?(\d{2}):(\d{2})\.(\d{3})")
i = 1
while i < len(lines):
match = tc_pattern.match(lines[i])
if match:
start_ms = time_to_ms(*match.groups()[:4])
end_ms = time_to_ms(*match.groups()[4:])
text_lines = []
i += 1
while i < len(lines) and not tc_pattern.match(lines[i]):
text_lines.append(lines[i])
i += 1
cues.append({
"start": start_ms,
"end": end_ms,
"text": " ".join(text_lines)
})
else:
i += 1
return cues
def time_to_ms(hh, mm, ss, ms) -> int:
h = int(hh) if hh else 0
return h * 3600000 + int(mm) * 60000 + int(ss) * 1000 + int(ms)
4. Millisecond Timecode Math & Drift Correction
When stitching audio chunks together or correcting drift caused by frame rate conversions (e.g. 23.976 fps to 29.97 fps), working with pure integer milliseconds avoids floating-point precision loss. For audio format container details, consult our Audio Formats, Codecs & Containers Guide.
5. Standardizing Word-Level JSON Transcripts
Modern web video applications (like custom audio waveforms and interactive karaoke players) require word-level precision. TranscriptG exports structured JSON following this standard (learn how to build vector indexes with this schema in our Audio Archives & Semantic Search Guide):
{
"durationSeconds": 142.5,
"language": "en",
"segments": [
{
"id": 1,
"start": 0.12,
"end": 3.48,
"speaker": "Speaker 1",
"text": "Welcome to TranscriptG's neural transcription platform.",
"words": [
{ "word": "Welcome", "start": 0.12, "end": 0.65, "confidence": 0.99 },
{ "word": "to", "start": 0.68, "end": 0.82, "confidence": 0.99 },
{ "word": "TranscriptG", "start": 0.85, "end": 1.45, "confidence": 0.98 }
]
}
]
}
6. Integrating with TranscriptG's Ephemeral API
Developers can integrate with TranscriptG to generate frame-accurate SRT, VTT, and JSON exports instantly with zero data persistence overhead. Try out instantaneous format conversions on our interactive Subtitle Converter Tool or test our speech models via the Free Speech Transcriber.