Local Inverse Speaking Rate Estimation via Hierarchical Prosodic Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for estimating speaking rate in speech recognition and text-to-speech systems fail to accurately estimate local speaking rate due to the lack of consideration for prosodic structure and text content, resulting in synthesized speech sounding boring and lacking local variation.
Innovation Solution
A method using a hierarchical structure to combine a prosodic module with a prosodic structure, employing maximum a posteriori (MAP) conditions to estimate local inverse speaking rate (ISR) by analyzing syllable duration, tone, and break types, allowing for more accurate and varied speech synthesis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional methods use average syllable duration or phoneme duration of the whole utterance to estimate speaking rate, then the estimation process is simple, but the estimation precision is poor and cannot capture local variation
Solution Approach 1:
The patent divides the utterance into multiple local regions (e.g., phoneme-level, syllable-level, word-level segments) and estimates speaking rate for each segment separately. This segmentation allows the system to capture local variations in speaking rate while maintaining a systematic estimation framework, resolving the contradiction between precision and complexity.
Solution Approach 2:
The patent introduces a hierarchical structure with multiple levels (utterance-level, phrase-level, word-level, phoneme-level) to estimate speaking rate. By adding this dimensional hierarchy, the system achieves both local precision and global consistency, transforming a one-dimensional average estimation into a multi-dimensional localized estimation system.
2Adaptability or versatility
If the prosodic generation scheme uses the whole utterance to estimate SR in the training stage, then the training process is straightforward, but the synthesized utterance cannot present local variation in SR
Solution Approach 1:
The patent applies different speaking rate parameters to different local regions of the utterance based on prosodic structure analysis. Each phoneme or syllable segment receives a locally optimized SR value derived from training data, enabling the synthesized speech to exhibit natural local variations while maintaining overall prosodic coherence.
Solution Approach 2:
The system performs preliminary analysis of prosodic structure and estimates local speaking rate parameters during the training stage before actual synthesis. This preliminary preparation of local SR estimates for each segment allows the synthesis process to directly apply these pre-computed values, reducing real-time complexity while achieving adaptability.
3Measurement precision
If ISR estimation considers prosodic structure and text content factors, then the estimation accuracy improves, but the estimation process becomes more complex and computationally intensive
Solution Approach 1:
The patent segments the ISR estimation process into multiple independent components: prosodic structure analysis, text content feature extraction, and local region identification. Each component processes specific features separately and combines results hierarchically, improving accuracy through comprehensive feature consideration while managing complexity through modular processing.
Data Source
AI summary
A method is disclosed. The proposed method includes: providing an initial speech corpus including plural utterances; based on a condition of maximum a posteriori (MAP), according to respective sequences of syllable duration, syllable duration prosodic state, syllable tone, base-syllable type, and break type of the kth utterance, using a probability of an ISR of the kth utterance xk to estimate an estimated value {circumflex over (x)}k of the xk; and through the MAP condition, according to respective sequences of syllable duration, syllable duration prosodic state, syllable tone, base-syllable type, and break type of the given lth breath group/prosodic phrase group (BG/PG) of the kth utterance, using a probability of an ISR of the lth BG/PG of the kth utterance xk,l to estimate an estimated value {circumflex over (x)}k,l of the xk,l wherein the {circumflex over (x)}k,l is the estimated value of local ISR, and a mean of a prior probability model of the {circumflex over (x)}k,l is the {circumflex over (x)}k.


