Multilingual Text-to-Speech Control with Unified Pitch and Duration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text-to-speech (TTS) systems lack the necessary controllability and efficiency for altering aspects of human speech like pitch, speaking rates, and pause durations, especially in dubbing applications, requiring additional sub-networks that complicate training and computation.
Innovation Solution
A system utilizing a residual vector quantization (RVQ) based aligner and TTS model trained on common representations, with integrated pitch and duration predictors, and leveraging large language models (LLMs) for linguistic context alignment, enabling efficient control over speech characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If additional sub-networks are added to control pitch, speaking rates, and pause durations, then speech characteristic controllability is improved, but device complexity increases
Solution Approach 1:
The patent merges pitch prediction and duration prediction into a unified TTS framework. The duration predictor network and pitch predictor network are integrated components that work together within the same model architecture, sharing common representations and training data, thereby achieving speech characteristic controllability without proportionally increasing system complexity
Solution Approach 2:
The TTS model is designed with multi-functionality to handle multiple speech characteristics simultaneously. The same model architecture predicts both pitch and duration values, and can control various aspects of speech including speaking rate, pause durations, and pitch contours through a unified framework that serves multiple control functions
2Reliability
If separate training processes are used for aligner and TTS model, then model specialization is improved, but training efficiency decreases
Solution Approach 1:
The patent combines the training processes of the aligner and TTS model into a joint training framework. Both components are trained together using shared training data, allowing them to learn complementary representations simultaneously. This unified training approach improves training efficiency while maintaining the specialized functions of each component through their interconnected architecture
3Measurement precision
If multiple independent predictors are used for pitch and duration, then prediction accuracy is improved, but computational overhead increases
Solution Approach 1:
The patent merges the computational processes of pitch and duration prediction into a unified framework. Both predictors share common input representations and are trained jointly, allowing them to leverage shared computational resources and intermediate representations. This reduces redundant computations while maintaining the accuracy benefits of having separate prediction capabilities for pitch and duration
Data Source
AI summary
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for performing text-to-speech modeling. In some implementations, a computer device receives input text in a first language to convert to a desired speech in a second language. The computer device receives one or more criteria for modifying the desired speech and converts the input text to a desired text in the second language. The computer device generates audio representations of the desired text in the second language and predicts, for each of the audio representations of the desired text and using the one or more received criteria, a pitch value and a duration value. The computer device generates the desired output speech using the predicted pitch value and the predicted duration value for each of the audio representations and provides the desired speech for output.


