Neural Voice Synthesis for Automated Media Localization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for automatically generating localized dubbed video content face challenges due to the lack of diversity in available voices, accents, and speaking characteristics, resulting in suboptimal customer experience and increased time and cost for media content localization.

Innovation Solution

The development of a system using neural network-based voice synthesis models that learn and replicate the speech characteristics of actors across different languages, enabling rapid generation of dubbed content with similar voice characteristics, allowing for more efficient and effective media content localization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional manual dubbing processes are used, then voice diversity and accuracy are improved, but time consumption and cost increase significantly

Engineering Contradiction:
Improvevoice characteristics accuracyVSAvoiddubbing production time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent creates virtual voice copies of actors using neural networks. The system learns speech patterns from existing audio recordings and generates synthetic voice performances that replicate the original actor's characteristics without requiring the actor to physically record new languages. This copying approach enables rapid dubbing production while maintaining voice authenticity.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical process of manual dubbing recording sessions with an automated neural network-based speech synthesis system. Instead of coordinating actors, directors, and recording equipment across multiple sessions, the system uses machine learning models to automatically generate localized dialogue, dramatically reducing production time and complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated voice synthesis is used, then production speed increases, but voice diversity and naturalness deteriorate

Engineering Contradiction:
Improvedubbing generation speedVSAvoidvoice characteristics accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent employs neural networks that learn and replicate specific speech parameters including pitch, tone, rhythm, and pronunciation patterns from original actor recordings. By adjusting and optimizing these acoustic parameters, the system generates synthetic voices that maintain naturalness and accuracy while enabling rapid automated production across multiple languages.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If expert language specialists are used for dialogue translation, then translation accuracy improves, but cost and time requirements increase

Engineering Contradiction:
Improvedialogue translation accuracyVSAvoidlocalization process complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent integrates multiple functions into a unified automated system: speech recognition, translation, and voice synthesis are combined in a single pipeline. The neural network system simultaneously handles language conversion and voice replication, eliminating the need for separate expert interventions in each stage and reducing overall process complexity while maintaining quality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10930263B1Automatic voice dubbing for media content localization
Publication Date: 2021.02.23 AMAZON TECH INC
  • US10930263B1 patent drawing
  • US10930263B1 patent drawing
  • US10930263B1 patent drawing

AI summary

This disclosure describes techniques for replicating characteristics of an actor or actresses voice across different languages. The disclosed techniques have the practical application of enabling automatic generation of dubbed video content for multiple languages, with particular speakers in each dubbing having the same voice characteristics as the corresponding speakers in the original version of the video content.