Speaking-Rate Normalized Prosodic Model for Multi-Dialect TTS
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for building prosodic models in Chinese text-to-speech systems are limited in their ability to handle various dialects, cross-language conversions, and speaker variations, and lack effective speaking-rate controlled prosodic parameters.
Innovation Solution
A speaking-rate controlled prosodic information generation device and method that uses a Maximum a Posteriori Linear Regression (MAPLR) algorithm to normalize and generate prosodic parameters for different languages and speakers, allowing for the synthesis of speech with varying speaking rates and styles using a combination of adaptive processing techniques and statistical models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a prosodic model is built using existing prosodic models or pattern recognition tools, then the model creation process is simplified, but the model lacks sufficient training data and systematic framework for various Chinese dialects
Solution Approach 1:
The patent pre-processes and normalizes prosodic parameters from large amounts of speech data before model training, creating a standardized training dataset in advance. This preliminary normalization of pitch contours, syllable durations, and pause durations based on speaking rate enables the model to effectively learn from diverse Chinese dialects without requiring extensive manual data preparation during model creation
Solution Approach 2:
The patent creates a universal prosodic model framework that can handle multiple Chinese dialects and languages simultaneously. The model uses language-independent prosodic features and speaking-rate normalization techniques that apply across different linguistic contexts, enabling a single model to serve multiple dialects including Mandarin, Cantonese, and other Chinese languages
2Ease of operation
If speech conversion is limited to one language and its sub-dialects, then the conversion process remains manageable, but the system cannot convert among seven dialects of Chinese
Solution Approach 1:
The patent implements a universal prosodic conversion framework that handles multiple Chinese dialects and languages through language-independent features. The system uses common prosodic parameters such as pitch contours, syllable durations, and pause durations that are consistent across different languages, enabling conversion among seven Chinese dialects while maintaining manageable operational complexity
Solution Approach 2:
The patent transforms prosodic parameters into speaking-rate normalized forms that are language-agnostic. By converting raw prosodic data into standardized parameters normalized for speaking rate effects, the system enables seamless conversion across different languages and dialects without requiring language-specific processing pipelines
3Device complexity
If prosodic model refinement is not performed across languages and speakers, then the model development process is simpler, but the prosody estimation remains limited
Solution Approach 1:
The patent segments the prosodic model development into distinct components: pitch contour modeling, syllable duration modeling, and pause duration modeling. Each component is independently refined across multiple languages and speakers, then integrated into a unified model. This segmented approach improves estimation precision while keeping the development process organized and manageable
Solution Approach 2:
The patent implements iterative refinement of prosodic parameters using feedback from multiple languages and speakers. The model continuously adjusts its parameters based on performance metrics across different dialects, improving prosody estimation accuracy through repeated cycles of evaluation and optimization
4Device complexity
If speaking-rate controlled prosodic parameters are not implemented, then the system is simpler, but it cannot generate speech with varying speaking rates and styles
Solution Approach 1:
The patent introduces speaking-rate normalized prosodic parameters that explicitly model the relationship between speaking rate and prosodic features. By transforming raw prosodic data into rate-normalized parameters, the system can generate speech at any desired speaking rate while maintaining natural prosodic patterns, adding versatility without excessive complexity
5Device complexity
If only a single language from a single speaker is learned, then the training process is simpler, but the system cannot mimic various speakers' speaking styles
Solution Approach 1:
The patent segments speaker-specific characteristics into separate learnable parameters distinct from language-level prosodic patterns. By separating speaker identity features from linguistic prosody, the system can independently adapt to multiple speakers while maintaining efficient training processes
Solution Approach 2:
The patent creates a unified training framework that simultaneously learns from multiple speakers and languages using consistent processing pipelines. The system uses universal feature representations that work across different speakers, enabling multi-speaker style imitation without requiring separate training processes for each speaker
Data Source
AI summary
A speaking-rate dependent prosodic model builder and a related method are disclosed. The proposed builder includes a first input terminal for receiving a first information of a first language spoken by a first speaker, a second input terminal for receiving a second information of a second language spoken by a second speaker and a functional information unit having a function, wherein the function includes a first plurality of parameters simultaneously relevant to the first language and the second language or a plurality of sub-parameters in a second plurality of parameters relevant to the second language alone, and the functional information unit under a maximum a posteriori condition and based on the first information, the second information and the first plurality of parameters or the plurality of sub-parameters produces speaking-rate dependent reference information and constructs a speaking-rate dependent prosodic model of the second language.


