Automated Video Dubbing Timing Synchronization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for translating speech in videos, such as using voice actors or automatic speech generation, often result in out-of-sync dubbed audio due to inaccuracies in timing synchronization, leading to a poor viewer experience.

Innovation Solution

A method that uses speech recognition data and original caption data to align translated audio with the video by mapping caption character strings to generated character strings based on semantic similarities, assigning timing information, and overlaying translated audio speech segments onto corresponding video segments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional voice dubbing methods are used to translate speech in videos, then the translated audio can be generated, but the dubbed speech does not align with the original speech timing and appears out of sync

Engineering Contradiction:
Improvetiming synchronization accuracyVSAvoiddubbing process complexity
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent applies preliminary action by extracting and storing timing information from the original video audio before translation. The system pre-processes the original audio to capture precise start and end times of speech segments, then uses this pre-extracted timing data to synchronize the translated audio, ensuring accurate alignment without requiring complex post-processing adjustments.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual voice dubbing methods with an automated system that uses speech recognition and text-to-speech generation. Instead of relying on human voice actors and manual timing adjustments, the system automatically generates translated audio and synchronizes it using algorithmic timing extraction and mapping, substituting mechanical automation for manual processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated speech generation is used for translation, then productivity increases, but the timing synchronization accuracy deteriorates causing out-of-sync audio

Engineering Contradiction:
Improvetranslation processing speedVSAvoidtiming alignment accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements feedback by continuously comparing the timing information of generated translated audio with the extracted original audio timing. The system uses this feedback to adjust and refine the synchronization of translated speech segments, ensuring that automated generation maintains precise timing alignment with the original video content.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces timing information extraction and mapping as an intermediary process between automated speech generation and final audio output. This intermediary layer captures precise timing data from the original audio and uses it to guide the synchronization of generated translated audio, bridging the gap between automated processing and timing accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12114048B2Automated voice translation dubbing for prerecorded videos
Publication Date: 2024.10.08 GOOGLE LLC
  • US12114048B2 patent drawing
  • US12114048B2 patent drawing
  • US12114048B2 patent drawing

AI summary

A method for aligning a translation of original caption data with an audio portion of a video is provided. The method involves identifying original caption data for the video that includes caption character strings, identifying translated language caption data for the video that includes translated character strings associated with audio portion of the video, and mapping caption sentence fragments generated from the caption character strings to corresponding translated sentence fragments generated from the translated character strings based on timing associated with the original caption data and the translated language caption data. The method further involves estimating time intervals for individual caption sentence fragments using timing information corresponding to individual caption character strings, assigning time intervals to individual translated sentence fragments based on estimated time intervals of the individual caption sentence fragments, generating a set of translated sentences using consecutive translated sentence fragments, and aligning the set of translated sentences with the audio portion of the video using assigned time intervals of individual translated sentence fragments from corresponding translated sentences.