Local Inverse Speaking Rate Estimation via Hierarchical Prosodic Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for estimating speaking rate in speech recognition and text-to-speech systems fail to accurately estimate local speaking rate due to the lack of consideration for prosodic structure and text content, resulting in synthesized speech sounding boring and lacking local variation.

Innovation Solution

A method using a hierarchical structure to combine a prosodic module with a prosodic structure, employing maximum a posteriori (MAP) conditions to estimate local inverse speaking rate (ISR) by analyzing syllable duration, tone, and break types, allowing for more accurate and varied speech synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional methods use average syllable duration or phoneme duration of the whole utterance to estimate speaking rate, then the estimation process is simple, but the estimation precision is poor and cannot capture local variation

Engineering Contradiction:
Improvelocal speaking rate estimation precisionVSAvoidestimation system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the utterance into multiple local regions (e.g., phoneme-level, syllable-level, word-level segments) and estimates speaking rate for each segment separately. This segmentation allows the system to capture local variations in speaking rate while maintaining a systematic estimation framework, resolving the contradiction between precision and complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a hierarchical structure with multiple levels (utterance-level, phrase-level, word-level, phoneme-level) to estimate speaking rate. By adding this dimensional hierarchy, the system achieves both local precision and global consistency, transforming a one-dimensional average estimation into a multi-dimensional localized estimation system.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If the prosodic generation scheme uses the whole utterance to estimate SR in the training stage, then the training process is straightforward, but the synthesized utterance cannot present local variation in SR

Engineering Contradiction:
Improvelocal SR variation capabilityVSAvoidprosody generation complexity
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent applies different speaking rate parameters to different local regions of the utterance based on prosodic structure analysis. Each phoneme or syllable segment receives a locally optimized SR value derived from training data, enabling the synthesized speech to exhibit natural local variations while maintaining overall prosodic coherence.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system performs preliminary analysis of prosodic structure and estimates local speaking rate parameters during the training stage before actual synthesis. This preliminary preparation of local SR estimates for each segment allows the synthesis process to directly apply these pre-computed values, reducing real-time complexity while achieving adaptability.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If ISR estimation considers prosodic structure and text content factors, then the estimation accuracy improves, but the estimation process becomes more complex and computationally intensive

Engineering Contradiction:
ImproveISR estimation accuracyVSAvoidestimation process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the ISR estimation process into multiple independent components: prosodic structure analysis, text content feature extraction, and local region identification. Each component processes specific features separately and combines results hierarchically, improving accuracy through comprehensive feature consideration while managing complexity through modular processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11200909B2Method of generating estimated value of local inverse speaking rate (ISR) and device and method of generating predicted value of local ISR accordingly
Publication Date: 2021.12.14 NAT YANG MING CHIAO TUNG UNIV
  • US11200909B2 patent drawing
  • US11200909B2 patent drawing
  • US11200909B2 patent drawing

AI summary

A method is disclosed. The proposed method includes: providing an initial speech corpus including plural utterances; based on a condition of maximum a posteriori (MAP), according to respective sequences of syllable duration, syllable duration prosodic state, syllable tone, base-syllable type, and break type of the kth utterance, using a probability of an ISR of the kth utterance xk to estimate an estimated value {circumflex over (x)}k of the xk; and through the MAP condition, according to respective sequences of syllable duration, syllable duration prosodic state, syllable tone, base-syllable type, and break type of the given lth breath group/prosodic phrase group (BG/PG) of the kth utterance, using a probability of an ISR of the lth BG/PG of the kth utterance xk,l to estimate an estimated value {circumflex over (x)}k,l of the xk,l wherein the {circumflex over (x)}k,l is the estimated value of local ISR, and a mean of a prior probability model of the {circumflex over (x)}k,l is the {circumflex over (x)}k.