Dialogue Localization via Max Duration Audio Sync

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Localization of voice-overs in media such as video games and animations often requires significant reworking of dialogue to fit with existing video components, which can be time-consuming and inefficient, especially when multiple languages are involved.

Innovation Solution

A computer-implemented method and system that determines the required length of a scene with a speaking character by translating scripts into multiple languages, performing text-to-speech processing, and calculating the maximum spoken duration, allowing for advanced configuration of media components and generation of synchronized animations or voice recordings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If localization is performed after media completion, then the media can be produced efficiently in original language first, but the localized dialogue must fit with existing video components requiring time-consuming reworking

Engineering Contradiction:
Improvemedia production efficiencyVSAvoidlocalization reworking time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary text-to-speech processing to generate localized audio samples before final media production. By determining the maximum spoken duration across multiple languages in advance, the system enables early configuration of media components to accommodate localized dialogue, eliminating the need for time-consuming reworking after completion.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If dialogue is translated into multiple languages, then the media becomes accessible to international audiences, but the duration of spoken content varies across languages requiring complex timing adjustments

Engineering Contradiction:
Improvelanguage accessibilityVSAvoidtiming coordination complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system determines the maximum spoken duration across all language versions and uses this single parameter to configure media components. By focusing on the maximum duration rather than individual language durations, the system simplifies timing coordination while maintaining adaptability across multiple languages.

Inventive Principle:
Principle #35Parameter changes

3Loss of time

If the script is translated into faster languages with shorter duration, then the media production timeline is shortened, but the animation and video components cannot be properly synchronized

Engineering Contradiction:
Improveproduction timelineVSAvoidsynchronization accuracy
Core Design Contradiction:
Loss of timeVSManufacturing precision

Solution Approach 1:

The system performs preliminary determination of the maximum spoken duration across all languages before animation and video production. This advance knowledge allows precise configuration of timing parameters, ensuring accurate synchronization between dialogue and visual components while accommodating the longest language version.

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables early anticipation and accommodation of localized dialogue in media production, improving efficiency by allowing for better alignment of animations and voice performances across languages, reducing the need for extensive reworking and ensuring consistent duration across different language versions.

Implementation Method 1

performing text-to-speech processing to generate a localized audio sample for the script in each of the first language and the one or more second languages

Methodology Applied
Scientific EffectText-to-speech processing:

Data Source

PatentUS20240371359A1Dialogue Localisation
Publication Date: 2024.11.07 SONY COMP ENTERTAINMENT EURO LTD
  • US20240371359A1 patent drawing
  • US20240371359A1 patent drawing
  • US20240371359A1 patent drawing

AI summary

The invention includes a computer-implemented method of determining a required length of a scene comprising a speaking character, wherein the scene is localized in multiple spoken languages, the method comprising: obtaining a script for the speaking character in a first language; automatically translating the script into one or more second languages; performing text-to-speech processing to generate a localized audio sample for the script in each of the first language and the one or more second languages; determining a duration of the localized audio sample in each of the first language and the one or more second languages; and determining a maximum spoken duration of the script as the maximum of the respective durations of the localized audio sample in each of the first language and the one or more second languages.