Face-Translator Pipeline for Voice-Preserving Lip Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing dubbing technologies fail to accurately synchronize lip movements and preserve voice characteristics in translated audio-visual content, leading to unnatural results and high production costs, especially in dynamic and low-cost content like videoconferencing and lectures.
Innovation Solution
An end-to-end system for lip-synchronous, voice-preserving video translation using machine learning modules for audio and video processing, including face detection, lip generation, and voice conversion, to generate translated videos with synchronized lip movements and preserved voice characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If human voice talents are used for dubbing, then voice characteristics and lip synchronization quality are improved, but production cost and time consumption increase
Solution Approach 1:
The patent replaces the mechanical process of human voice talents recording and lip-syncing with an automated deep learning system. The system uses neural networks to generate synthesized voice tracks and synchronize lip movements automatically, eliminating the need for human performers while maintaining high synchronization precision and reducing production time significantly.
Solution Approach 2:
The patent creates synthetic copies of voice characteristics and lip movements through deep learning models. The system learns from training data to generate fake but realistic voice tracks and lip-synced video frames that replicate human performance, allowing unlimited scalability without the constraints of human availability and cost.
2Productivity
If machine translation with isometric temporal match is used, then translation speed is improved, but lip movement synchronization and voice characteristic preservation deteriorate
Solution Approach 1:
The patent merges multiple previously separate functions into a unified end-to-end system: speech recognition, machine translation, text-to-speech synthesis, voice conversion, and lip synchronization are combined into a single integrated pipeline. This allows the system to optimize all parameters simultaneously, achieving both fast translation and precise lip-syncing that were mutually exclusive in earlier approaches.
Solution Approach 2:
The patent implements feedback mechanisms where the system iteratively adjusts lip-synced video frames and voice tracks based on alignment between audio and visual features. The deep learning models continuously refine their outputs to ensure temporal consistency and natural synchronization, correcting mismatches that would occur in sequential processing approaches.
3Adaptability or versatility
If dubbing is used to translate video content, then audio-visual translation is achieved, but naturalness and conviction of the result deteriorate due to phonetic mismatch
Solution Approach 1:
The patent applies dynamic adjustments to lip movements and facial expressions based on the synthesized audio track. The system analyzes the phonetic content of the translated speech and dynamically modifies the corresponding lip-synced frames to match the actual sounds produced, creating a coherent and natural appearance that adapts to the specific linguistic characteristics of the target language.
Solution Approach 2:
The patent changes multiple parameters simultaneously including lip shape, facial expression, timing, and prosody to match the synthesized audio. The deep learning models adjust these parameters in a coordinated manner to ensure that the visual and audio channels are perfectly aligned, transforming the unnatural dubbing effect into a convincing and natural translation experience.
Data Source
AI summary
A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.


