There are two completely different ways to end up with a transcript in a language other than the one being spoken, and people constantly reach for the wrong one. Getting this right is the difference between a clean, accurate document and a machine translation of a machine transcription.
The two paths
| Switching the caption track | Translating | |
|---|---|---|
| What it does | Loads a different caption track that already exists on the video | Takes the transcript you have and rewrites it in your target language |
| Quality | As good as whatever the creator uploaded — often excellent | Depends on the quality of the source transcript |
| Availability | Only languages the video actually has tracks for | 125+ languages, any video with any transcript |
| When to use it | Always check this first | When no track exists in the language you need |
Translating, step by step
Get the transcript
Paste the video URL and extract as normal. If several caption tracks exist, start from the best one available — a creator-uploaded track gives a noticeably better translation than an auto-generated one.
Choose your target language
Pick from the dropdown of common languages, or type any language name into the field if yours isn't listed.
Run the translation
The whole transcript is translated segment by segment. Timestamps are preserved, so the translated text stays aligned to the video.
Read, adjust, and export
Switch to paragraph view for readability, drop any lines you don't need, then copy or download.
Why timestamps survive
A transcript isn't one block of text — it's a list of segments, each carrying its own start time and duration. Translation is applied to the segments, not to a flattened string, so every translated line inherits the timing of the line it came from.
The practical consequence is that click-to-seek keeps working after translation. You can read a Japanese lecture in Portuguese and still click any line to jump to that exact moment in the video. If you're curious about how the timing data is structured, how transcript timestamps work goes into it.
Getting better translations
- Start from the best source track. Punctuation and sentence boundaries give a translator enormous amounts of context. Translating an unpunctuated auto-generated track is the hardest possible starting point.
- Name the language plainly. Where a language has meaningful regional variants, being specific — Brazilian Portuguese rather than just Portuguese, Simplified rather than Traditional Chinese — gets you closer to what you actually want.
- Expect proper nouns to wobble. Names, places, product names and technical jargon are where machine translation is weakest, especially if they were already mangled by speech recognition upstream.
- Spot-check against the video. Click a few translated lines and listen. It takes a minute and catches the worst errors.
A note on subtitle files
Translating a transcript gives you readable text. It doesn't give you a subtitle file you can upload back to a video — that's a different format with its own structure and timing rules. The distinction between transcripts, subtitles and closed captions is covered in transcript vs. subtitles vs. closed captions.