Voice Generation Model for Authentic Multilingual Audio Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current Text to Speech (TTS) systems fail to accurately capture the tone and voice modulation of original audio tracks when converting text scripts into different languages, limiting the reach of multimedia content to regions where the original language is spoken, as they require costly and time-consuming dubbing processes.

Innovation Solution

A voice generation model that extracts and processes voice characteristic information from a reference voice sample, including phonemes, pitch, and energy, to generate an output voice track that mimics the vocal characteristics of the original audio, allowing for the conversion of text data into a desired language while maintaining the tone and modulation of the original voice.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional Text to Speech systems are used to convert text scripts into different languages, then language conversion is achieved, but the tone and voice modulation of the original audio track are lost

Engineering Contradiction:
Improvelanguage conversion capabilityVSAvoidvoice characteristic accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent creates a digital copy of the original voice characteristics by extracting phoneme-level features (pitch, energy, duration) from the reference audio track and applying them to the target language text. This copying approach preserves the original voice's tone and modulation while enabling language conversion, resolving the contradiction between adaptability and precision.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent segments the audio processing into phoneme-level units, analyzing and synthesizing voice characteristics at the smallest meaningful linguistic unit. This segmentation allows precise control over pitch, energy, and duration for each phoneme, maintaining voice authenticity while enabling language translation.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If voice actors are used to dub content in different languages, then vocal authenticity is maintained, but the process becomes costly and time-consuming

Engineering Contradiction:
Improvevocal authenticityVSAvoiddubbing efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent replaces the mechanical system of human voice actors with an automated computational system that extracts and applies voice characteristics algorithmically. This substitution maintains vocal authenticity through precise phoneme-level analysis while dramatically improving productivity by eliminating the need for manual dubbing processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables the audio content to serve itself by automatically extracting its own voice characteristics and applying them to translated text. This self-service capability eliminates the need for external voice actors, reducing costs and time while preserving the original speaker's vocal identity.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If coarse control TTS converters are used for language conversion, then any language can be generated, but tone and voice modulation cannot be captured

Engineering Contradiction:
Improvelanguage generation capabilityVSAvoidvoice characteristic precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent adds phoneme-level dimensionality to the TTS process, moving from word or sentence-level control to individual phoneme control. This dimensional change enables simultaneous achievement of multi-language capability and precise voice characteristic preservation, as each phoneme can be independently analyzed and synthesized with accurate pitch, energy, and duration parameters.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20240347038A1Output voice track generation
Publication Date: 2024.10.17 GAN STUDIO INC
  • US20240347038A1 patent drawing
  • US20240347038A1 patent drawing
  • US20240347038A1 patent drawing

AI summary

Approaches for generating an output voice track corresponding to an input text data using a voice generation system are described. In an example, by the voice generation system, a reference voice sample and the input text data is obtained. In an example, form the reference voice sample, a voice characteristic information and corresponding attribute values are extracted. The voice characteristic information may thus be processed based on a voice generation model. The voice generation model is to assign a weight for each of the voice characteristics based on their attribute values. Once a weighted voice characteristic information is generated, an output voice track corresponding to the input text data is generated.