An inflight digitalservices provider delivering connectivity, entertainment, and retail for over50 airlines
A provider of inflight digital services for airlines held its entire crew training library in a single source language. As the business expanded into new regions, every crew needed the same standardized training, and the established route to get there was re-recording each video with new voice talent in a studio. That route is cost-intensive, slow, and operationally complex, and it gets harder with every language added.
Marlabs proactively designed and delivered a working AI-powered dubbing solution that produces professional-quality voice-over versions of the client's existing videos without recreating the original content. The solution runs a seven-stage automated pipeline: it segments each video into 15-second scenes, transcribes and translates each segment, synthesizes a natural AI voice, time-aligns the new audio, dubs it over the original track, and assembles the finished video. Re-syncing at every segment boundary keeps long-form training aligned end to end.
The client now has a production-ready capability rather than a concept. A new language is a pipeline run instead of a production, and the same capability extends to product, compliance, and onboarding videos across the business.
The client's crew training content existed only in its original source language. Crew training is not optional content. It carries safety, service, and operating procedures that every crew member has to receive in a form they fully understand, and regulators and airline partners expect that training to be consistent wherever it is delivered.
Global expansion turned that single-language library into a constraint. Each new region meant each new crew needed the same standardized training, and the only established way to deliver it was to re-record every video in the new language. That demanded studio time, voice talent, and production scheduling the business could not absorb, and the cost and lead time scaled with every language added rather than flattening out.
The alternatives were no better. Subtitling changes how training lands for crew members who are working through the material rather than watching it. Manual localization of the audio drifts out of sync across a long video, and the further into a recording the drift accumulates, the more the narration stops matching what is on screen. Quality also varied by region and by vendor, which undercut the standardization the client needed in the first place.
Marlabs designed and delivered a working AI-powered multilingual dubbing solution, built on Azure AI services and delivered proactively at no cost to the client. The design principle was that the original footage is reused exactly as it is. Nothing is re-shot, no presenter is re-recorded, and the finished video keeps the visuals, pacing, and background audio of the original. Adding a language becomes a pipeline run rather than a production.
Marlabs assessed the client's crew training library to understand its structure, typical length, and audio characteristics. The team designed a segment-based architecture around a deliberate constraint: rather than processing a video as one long piece of audio, the pipeline breaks it into short scenes and re-establishes timing at every boundary. That decision is what keeps a long training video aligned from the first minute to the last. The team also defined the source and target language scope and confirmed that the approach supports any source and target language pair.
The team used PySceneDetect to split each video into 15-second scenes, giving the pipeline a consistent unit of work and a fixed set of timing reference points. Azure AI Speech then transcribed every segment individually, producing a speech-to-text output tied to a known position in the video. Working at segment level rather than across the whole file keeps each transcript short and its timing unambiguous. It also means an issue in one scene stays contained in that scene instead of propagating through everything that follows.
Our team translated each segment transcript into the target language with Azure AI Translator, keeping the translation aligned to the same segment boundaries established during transcription. Azure AI Speech then generated a natural AI voice track for each translated segment using neural text-to-speech. The resulting narration is professional-quality and consistent in tone across the whole video, because every segment is voiced by the same synthesized speaker rather than by different talent on different recording days. A human reviewer signs off on the translated script and the finished audio before anything is published to crew.
The team fit each synthesized audio segment to the duration of the scene it belongs to so that the narration lands where the original speech landed. The new voice track was then dubbed over the original audio using broadcast-style ducking, which lowers the original track under the new voice rather than removing it. That keeps background context, ambient sound, and on-screen cues intact, and it removes any need for lip-sync work. FFmpeg stitched the processed segments back together into a finished video, and a Streamlit interface gave the team a working front end for running the pipeline and reviewing output.
The client received a working, production-ready solution rather than a proof of concept that stops at a demonstration. The pipeline localizes single-language crew training into new languages with automated voice-over, and it is built to scale across multiple regional languages.
Rolling standardized crew training into a new region became a pipeline run rather than a studio production. Reshoots, voice talent, and studio time came out of the localization process entirely, because the original footage is reused as-is. Every language version comes from the same source material so that crews in different regions receive the same training message rather than regional variants that drifted apart through separate production efforts.
Long-form content holds up. Because the pipeline re-syncs timing at every 15-second boundary, a long training video stays aligned from start to finish instead of drifting further out of sync the longer it runs. Broadcast-style ducking keeps the original audio present underneath the new narration so that background context is preserved and no lip-sync work is required.
The capability is not specific to one video library or one language pair. It supports any source and target language combination, which means adding a language takes minimal incremental effort. The same pipeline applies to product, compliance, and onboarding videos, and to any organization with training or product content that has to reach people in more than one language.
The delivered solution gave the client: