Voice Generation Model for Authentic Multilingual Audio Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current Text to Speech (TTS) systems fail to accurately capture the tone and voice modulation of original audio tracks when converting text scripts into different languages, limiting the reach of multimedia content to regions where the original language is spoken, as they require costly and time-consuming dubbing processes.
Innovation Solution
A voice generation model that extracts and processes voice characteristic information from a reference voice sample, including phonemes, pitch, and energy, to generate an output voice track that mimics the vocal characteristics of the original audio, allowing for the conversion of text data into a desired language while maintaining the tone and modulation of the original voice.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional Text to Speech systems are used to convert text scripts into different languages, then language conversion is achieved, but the tone and voice modulation of the original audio track are lost
Solution Approach 1:
The patent creates a digital copy of the original voice characteristics by extracting phoneme-level features (pitch, energy, duration) from the reference audio track and applying them to the target language text. This copying approach preserves the original voice's tone and modulation while enabling language conversion, resolving the contradiction between adaptability and precision.
Solution Approach 2:
The patent segments the audio processing into phoneme-level units, analyzing and synthesizing voice characteristics at the smallest meaningful linguistic unit. This segmentation allows precise control over pitch, energy, and duration for each phoneme, maintaining voice authenticity while enabling language translation.
2Manufacturing precision
If voice actors are used to dub content in different languages, then vocal authenticity is maintained, but the process becomes costly and time-consuming
Solution Approach 1:
The patent replaces the mechanical system of human voice actors with an automated computational system that extracts and applies voice characteristics algorithmically. This substitution maintains vocal authenticity through precise phoneme-level analysis while dramatically improving productivity by eliminating the need for manual dubbing processes.
Solution Approach 2:
The system enables the audio content to serve itself by automatically extracting its own voice characteristics and applying them to translated text. This self-service capability eliminates the need for external voice actors, reducing costs and time while preserving the original speaker's vocal identity.
3Adaptability or versatility
If coarse control TTS converters are used for language conversion, then any language can be generated, but tone and voice modulation cannot be captured
Solution Approach 1:
The patent adds phoneme-level dimensionality to the TTS process, moving from word or sentence-level control to individual phoneme control. This dimensional change enables simultaneous achievement of multi-language capability and precise voice characteristic preservation, as each phoneme can be independently analyzed and synthesized with accurate pitch, energy, and duration parameters.
Data Source
AI summary
Approaches for generating an output voice track corresponding to an input text data using a voice generation system are described. In an example, by the voice generation system, a reference voice sample and the input text data is obtained. In an example, form the reference voice sample, a voice characteristic information and corresponding attribute values are extracted. The voice characteristic information may thus be processed based on a voice generation model. The voice generation model is to assign a weight for each of the voice characteristics based on their attribute values. Once a weighted voice characteristic information is generated, an output voice track corresponding to the input text data is generated.


