Multilingual Speech Translation with Adaptive Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The high labor intensity and economic constraints of dubbing multimedia content for global diffusion, particularly due to the need for precise linguistic, acoustic, and emotional translation, create a barrier in making content accessible across languages.

Innovation Solution

A dubbing service that utilizes machine learning models to automatically translate and render audio in multimedia files, replicating the original audio's timbre, duration, prosody, and background noise, by integrating automatic speech recognition, machine translation, and text-to-speech synthesis, while also considering visual cues for synchronization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional dubbing methods are used to maintain linguistic, acoustic, and emotional translation quality, then translation accuracy is improved, but labor intensity and cost increase significantly

Engineering Contradiction:
Improvetranslation accuracyVSAvoiddubbing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system creates a digital copy of the original speaker's voice characteristics using machine learning models. The neural network learns and replicates the speaker's timbre, prosody, and speaking style, generating synthetic speech that closely mimics the original voice without requiring human dubbing artists to manually replicate these characteristics

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical process of human dubbing with an automated machine learning system. Instead of human actors recording new dialogue, the system uses neural networks to automatically generate speech that matches the original speaker's characteristics, substituting manual labor with computational processes

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If human dubbing artists are used to match timbre, duration, and emotion, then audio quality is improved, but time consumption and cost increase

Engineering Contradiction:
Improveaudio qualityVSAvoiddubbing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of the original speaker's voice characteristics before generating the dubbed speech. The machine learning model pre-learns the speaker's timbre, prosody, and speaking style from sample data, so that when generating new speech, it can quickly replicate these characteristics without time-consuming manual adjustment by dubbing artists

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adjusts multiple audio parameters including pitch, timbre, duration, and prosody to match the original speaker's characteristics. The neural network continuously optimizes these parameters during speech generation to maintain audio quality comparable to human dubbing while significantly reducing the time required

Inventive Principle:
Principle #35Parameter changes

3Productivity

If automated speech translation is used to reduce cost, then productivity is improved, but audio naturalness and emotional accuracy deteriorate

Engineering Contradiction:
Improvedubbing efficiencyVSAvoidaudio naturalness
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system creates a digital copy of the original speaker's voice characteristics using machine learning models. The neural network learns and replicates the speaker's timbre, prosody, and speaking style, generating synthetic speech that closely mimics the original voice without requiring human dubbing artists to manually replicate these characteristics

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system uses feedback mechanisms where the generated speech is compared against the original speaker's characteristics, and the model iteratively improves its replication accuracy. This allows the automated system to achieve natural-sounding audio by continuously refining its output based on learned patterns from the original speaker

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11545134B1Multilingual speech translation with adaptive speech synthesis and adaptive physiognomy
Publication Date: 2023.01.03 AMAZON TECH INC
  • US11545134B1 patent drawing
  • US11545134B1 patent drawing
  • US11545134B1 patent drawing

AI summary

Techniques for the generation of dubbed audio for an audio/video are described. An exemplary approach is to receive a request to generate dubbed speech for an audio/visual file; and in response to the request to: extract speech segments from an audio track of the audio/visual file associated with identified speakers; translate the extracted speech segments into a target language; determine a machine learning model per identified speaker, the trained machine learning models to be used to generate a spoken version of the translated, extracted speech segments based on the identified speaker; generate, per translated, extracted speech segment, a spoken version of the translated, extracted speech segments using a trained machine learning model that corresponds to the identified speaker of the translated, extracted speech segment and prosody information for the extracted speech segments; and replace the extracted speech segments from the audio track of the audio/visual file with the spoken versions spoken version of the translated, extracted speech segments to generate a modified audio track.