Multilingual Voice-Over Generation with Lip-Synced Speaker Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing solutions for dubbing media content into multiple languages face challenges such as high costs, time consumption, and poor user experience due to mismatched vocal characteristics and lip movements, especially when multiple speakers are involved.

Innovation Solution

A system that uses machine learning models to analyze speaker attributes, generate synchronized audio and video tracks in target languages, and adjust lip movements to match the new audio, employing audio and video generation models trained on speaker characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If voice-over artists are hired for dubbing audio tracks in target languages, then language accessibility is improved, but cost and time consumption increase

Engineering Contradiction:
Improvelanguage accessibilityVSAvoiddubbing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent uses deepfake technology to create synthetic copies of original speakers' voices and appearances. The system generates virtual versions of speakers that can speak multiple languages while maintaining the original speaker's vocal characteristics and appearance, eliminating the need to hire multiple voice-over artists for different languages.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system changes the language parameter of the synthetic speakers while maintaining other parameters such as vocal characteristics, tone, and appearance. By manipulating audio and video parameters through machine learning models, the same virtual speaker can produce different language versions of the content without physical re-recording.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If voice-over artists are hired for dubbing audio tracks in target languages, then language accessibility is improved, but production cost increases

Engineering Contradiction:
Improvelanguage accessibilityVSAvoidproduction cost
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent creates digital copies of speakers using deepfake technology. Instead of hiring multiple human artists, the system generates synthetic copies that can be replicated indefinitely at minimal cost. These virtual speakers can be used across multiple languages and platforms without additional compensation to human artists.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system enables content to dub itself by using the original speaker's own voice and appearance synthesized in the target language. The machine learning models automatically generate the dubbed content without requiring human intervention, making the process self-service and significantly reducing production costs.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If traditional dubbing is used, then language translation is achieved, but vocal characteristics and lip movements become mismatched

Engineering Contradiction:
Improvelanguage translationVSAvoidsynchronization accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent merges audio and video processing into a unified deepfake system. The audio generation model and video generation model work together to ensure that the synthesized audio and corresponding video frames are perfectly synchronized. The lip movements in the video are generated to match the phonemes in the synthesized audio, creating seamless synchronization.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system uses feedback loops where the audio generation model processes phoneme sequences and the video generation model processes phoneme frames to generate corresponding video frames. The feedback mechanism ensures that the lip movements and facial expressions in the video accurately reflect the synthesized audio, maintaining synchronization precision.

Inventive Principle:
Principle #23Feedback

4Adaptability or versatility

If multiple speakers are involved in media content, then content complexity is increased, but dubbing difficulty increases significantly

Engineering Contradiction:
Improvemulti-language supportVSAvoiddubbing process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the audio and video content by speaker, assigning unique identifiers to distinguish different speakers. The system processes each speaker's segments separately through the machine learning models, generating language-specific versions for each speaker. This segmentation approach maintains the complexity of multi-speaker content while simplifying the dubbing process through automated processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250210065A1Voiced-over multimedia track generation
Publication Date: 2025.06.26 GAN STUDIO INC
  • US20250210065A1 patent drawing
  • US20250210065A1 patent drawing
  • US20250210065A1 patent drawing

AI summary

Examples approaches for generating a final media track in a final language by altering an initial media track in an initial language, are described. In an example, an audio generation model is used to convert or translate an initial audio track of an initial language into a final audio track of a final language. Further, a video generation model is used to manipulate or alter movement of lips of a speaker in an initial video track based on the final audio track and a final text corresponding to each individual sentences. Once generated, the final audio track and the final video track are merged to generate a final audio-visual track or final media file.