Translated Video Dubbing With Audio-Driven Lip Synchronization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video translation and dubbing systems face issues such as separate processing of audio and video elements leading to unnatural synthesis, speech length discrepancies, and mismatched mouth shapes, resulting in suboptimal translation and dubbing experiences.

Innovation Solution

An end-to-end converter framework generates audio tokens, translates text content, adjusts visual frames based on audio features, and synchronizes mouth shapes to produce natural and efficient translated and dubbed video content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If separate processing of audio and video elements is used, then processing flexibility is improved, but synthesis naturalness deteriorates

Engineering Contradiction:
Improveprocessing flexibilityVSAvoidsynthesis naturalness
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent merges separate audio and video processing into a unified end-to-end converter framework that jointly processes both modalities, enabling synchronized generation of translated audio and corresponding video frames with consistent mouth shapes and temporal alignment, thereby improving synthesis naturalness while maintaining processing flexibility through modular architecture

Inventive Principle:
Principle #5Merging (Combining)

2Adaptability or versatility

If traditional translation and dubbing pipelines are used, then language translation capability is improved, but time synchronization accuracy deteriorates

Engineering Contradiction:
Improvelanguage translation capabilityVSAvoidtime synchronization accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary extraction and encoding of audio features and visual features from the input video before translation, creating synchronized feature representations that are then processed through the translation model, ensuring temporal alignment is established early in the pipeline and maintained through subsequent generation stages

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If manual translation and dubbing processes are used, then quality control is improved, but productivity deteriorates

Engineering Contradiction:
Improvequality controlVSAvoidprocessing efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent implements an automated end-to-end converter framework that performs translation, audio generation, and video frame modification without manual intervention, using trained neural network models to automatically ensure quality through consistent application of learned patterns, thereby achieving both high productivity and maintained quality control

Inventive Principle:
Principle #25Self-service

4Speed

If speech length adjustment is not performed, then processing speed is improved, but mouth shape synchronization deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidmouth shape synchronization
Core Design Contradiction:
SpeedVSManufacturing precision

Solution Approach 1:

The patent dynamically adjusts the length and timing of translated speech output based on the original audio duration and speech rate characteristics, using conditional processing that modifies generation parameters according to input properties, thereby achieving both efficient processing and accurate mouth shape synchronization in the generated video frames

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250358486A1Video content processing
Publication Date: 2025.11.20 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250358486A1 patent drawing
  • US20250358486A1 patent drawing
  • US20250358486A1 patent drawing

AI summary

The embodiment of the disclosure relates to methods, apparatuses, devices, and storage media for processing video content. The method provided herein includes: generating a set of audio tokens corresponding to first audio content of first video content, the first audio content corresponding to first text content of a first language; generating, based on an audio feature representation corresponding to the set of audio tokens, second audio content corresponding to second text content, the second text content being generated by translating the first text content into a second language; generating a second set of video frames based on a set of visual features corresponding to a first set of video frames of the first video content and the audio feature representation; and generating second video content based on the second set of video frames and the second audio content.