Speaking-Rate Normalized Prosodic Model for Multi-Dialect TTS

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for building prosodic models in Chinese text-to-speech systems are limited in their ability to handle various dialects, cross-language conversions, and speaker variations, and lack effective speaking-rate controlled prosodic parameters.

Innovation Solution

A speaking-rate controlled prosodic information generation device and method that uses a Maximum a Posteriori Linear Regression (MAPLR) algorithm to normalize and generate prosodic parameters for different languages and speakers, allowing for the synthesis of speech with varying speaking rates and styles using a combination of adaptive processing techniques and statistical models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If a prosodic model is built using existing prosodic models or pattern recognition tools, then the model creation process is simplified, but the model lacks sufficient training data and systematic framework for various Chinese dialects

Engineering Contradiction:
Improvemodel creation processVSAvoidtraining data
Core Design Contradiction:
Ease of manufactureVSQuantity of substance

Solution Approach 1:

The patent pre-processes and normalizes prosodic parameters from large amounts of speech data before model training, creating a standardized training dataset in advance. This preliminary normalization of pitch contours, syllable durations, and pause durations based on speaking rate enables the model to effectively learn from diverse Chinese dialects without requiring extensive manual data preparation during model creation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a universal prosodic model framework that can handle multiple Chinese dialects and languages simultaneously. The model uses language-independent prosodic features and speaking-rate normalization techniques that apply across different linguistic contexts, enabling a single model to serve multiple dialects including Mandarin, Cantonese, and other Chinese languages

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If speech conversion is limited to one language and its sub-dialects, then the conversion process remains manageable, but the system cannot convert among seven dialects of Chinese

Engineering Contradiction:
Improveconversion processVSAvoidlanguage conversion capability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal prosodic conversion framework that handles multiple Chinese dialects and languages through language-independent features. The system uses common prosodic parameters such as pitch contours, syllable durations, and pause durations that are consistent across different languages, enabling conversion among seven Chinese dialects while maintaining manageable operational complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent transforms prosodic parameters into speaking-rate normalized forms that are language-agnostic. By converting raw prosodic data into standardized parameters normalized for speaking rate effects, the system enables seamless conversion across different languages and dialects without requiring language-specific processing pipelines

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If prosodic model refinement is not performed across languages and speakers, then the model development process is simpler, but the prosody estimation remains limited

Engineering Contradiction:
Improvemodel development processVSAvoidprosody estimation
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent segments the prosodic model development into distinct components: pitch contour modeling, syllable duration modeling, and pause duration modeling. Each component is independently refined across multiple languages and speakers, then integrated into a unified model. This segmented approach improves estimation precision while keeping the development process organized and manageable

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements iterative refinement of prosodic parameters using feedback from multiple languages and speakers. The model continuously adjusts its parameters based on performance metrics across different dialects, improving prosody estimation accuracy through repeated cycles of evaluation and optimization

Inventive Principle:
Principle #23Feedback

4Device complexity

If speaking-rate controlled prosodic parameters are not implemented, then the system is simpler, but it cannot generate speech with varying speaking rates and styles

Engineering Contradiction:
Improvesystem structureVSAvoidspeaking rate control
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent introduces speaking-rate normalized prosodic parameters that explicitly model the relationship between speaking rate and prosodic features. By transforming raw prosodic data into rate-normalized parameters, the system can generate speech at any desired speaking rate while maintaining natural prosodic patterns, adding versatility without excessive complexity

Inventive Principle:
Principle #35Parameter changes

5Device complexity

If only a single language from a single speaker is learned, then the training process is simpler, but the system cannot mimic various speakers' speaking styles

Engineering Contradiction:
Improvetraining processVSAvoidspeaker style imitation
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent segments speaker-specific characteristics into separate learnable parameters distinct from language-level prosodic patterns. By separating speaker identity features from linguistic prosody, the system can independently adapt to multiple speakers while maintaining efficient training processes

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a unified training framework that simultaneously learns from multiple speakers and languages using consistent processing pipelines. The system uses universal feature representations that work across different speakers, enabling multi-speaker style imitation without requiring separate training processes for each speaker

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10192542B2Speaking-rate normalized prosodic parameter builder, speaking-rate dependent prosodic model builder, speaking-rate controlled prosodic-information generation device and prosodic-information generation method able to learn different languages and mimic various speakers' speaking styles
Publication Date: 2019.01.29 NATIONAL TAIPEI UNIVERSITY
  • US10192542B2 patent drawing
  • US10192542B2 patent drawing
  • US10192542B2 patent drawing

AI summary

A speaking-rate dependent prosodic model builder and a related method are disclosed. The proposed builder includes a first input terminal for receiving a first information of a first language spoken by a first speaker, a second input terminal for receiving a second information of a second language spoken by a second speaker and a functional information unit having a function, wherein the function includes a first plurality of parameters simultaneously relevant to the first language and the second language or a plurality of sub-parameters in a second plurality of parameters relevant to the second language alone, and the functional information unit under a maximum a posteriori condition and based on the first information, the second information and the first plurality of parameters or the plurality of sub-parameters produces speaking-rate dependent reference information and constructs a speaking-rate dependent prosodic model of the second language.