One pipeline, source master to delivered package.

Traditional localisation is five vendors, five handoffs and five chances to lose the timeline. Voxglot runs the whole chain on one set of assets, so a change to line 214 in the script lands in the mix without a re-conform.

One pipeline, source to master.

Five stages, one system. No vendor handoffs, no re-conform, no waiting on a studio calendar. Every stage writes to the same timeline, so a change in the script lands in the mix.

  1. 01

    Transcribe

    Word-level transcript with speaker turns, timecode and emotion tags.

    Diarisation separates every speaker on the stem, including overlapping dialogue. Music and effects are split out before a single word is transcribed, so the M&E bed survives untouched into the final mix.

    • Speaker-tagged transcript
    • Frame-accurate timecode
    • Isolated M&E stem
  2. 02

    Adapt

    human in the loop

    Cultural adaptation, not literal translation. Register, idiom, humour, names.

    A joke that lands in Lagos is not the same sentence that landed in Los Angeles. Adaptation rewrites for meaning and rhythm under a duration constraint, so the line still fits the shot it has to fit.

    • Adapted script
    • Idiom and honorific map
    • Duration-fitted lines
  3. 03

    Voice

    The original performer's voice, carried across the language barrier.

    Timbre, breath, age and attack are cloned from the source performance and re-performed in the target language. Consent and likeness rights are recorded per talent before a voice is ever built.

    • Cloned voice model
    • Per-line performance takes
    • Talent consent record
  4. 04

    Sync

    Phoneme-aware alignment, with optional visual lip-sync on the plate.

    Audio-first alignment fits the take to the shot. Where the licence allows it, visual lip-sync reshapes the mouth on the plate so the language stops fighting the picture.

    • Aligned dialogue track
    • Optional visual lip-sync
    • Conform report
  5. 05

    Mix

    human in the loop

    Dialogue seated back into the original M&E, delivered to broadcast spec.

    Loudness normalised per territory, stems delivered alongside the master, IMF and broadcast packages assembled to the deliverable spec of the platform you are shipping to.

    • 5.1 and stereo masters
    • Dialogue / M&E stems
    • IMF or broadcast package

Automation does the work. People sign it off.

What the machine decides

Speaker separation, timecode, phoneme alignment, loudness, and the thousand mechanical judgements that used to eat a studio week. These are measured, not debated, and they run at catalogue scale.

  • Diarisation and M&E separation
  • Duration-constrained line fitting
  • Alignment and conform
  • Loudness and delivery spec

What a person decides

Whether the joke lands. Whether the honorific is right for the relationship. Whether the performance reads as the same character. Native linguists and a dubbing engineer own these calls on Broadcast and Premium titles.

  • Register, idiom and humour
  • Names, honorifics and glossary lock
  • Performance direction and retakes
  • Final sign-off against the source

Delivered to the spec you actually ship against.

Masters
5.1 and stereo, per-territory loudness (EBU R128 / ATSC A/85)
Stems
Dialogue, music and effects delivered alongside every master
Packages
IMF, ProRes, or the broadcaster's own delivery spec
Text
Adapted script, dubbing script, and conformed subtitle files
Reports
Conform report, QC log, and per-line confidence scoring
Rights
Talent consent records and likeness grants, per title

Pre-release assets, handled like pre-release assets.

Encrypted at rest and in transit, per-title access control, watermarked review copies, and a full audit trail of who opened what. Regional processing available where a licence or a regulator requires it.

Send us one episode.

We'll dub it into two languages of your choosing and send back the master, the stems and the conform report. If it isn't broadcast-ready, you owe us nothing.