Hierarchical Prosodic Boundary Prediction for Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-to-speech conversion systems suffer from poor naturalness and flexibility due to the inflexibility of prosodic structure prediction models, which fail to accurately capture the varied prosodic structures influenced by speaker characteristics, emotions, and sentence meanings.

Innovation Solution

A method and apparatus for speech synthesis based on a large corpus that utilizes a prosodic structure prediction model to provide multiple alternative prosodic boundary partitioning solutions, determining the most probable solution based on structure probability information from a speech corpus, and performing speech synthesis accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a uniform prosodic structure prediction model is used, then the system complexity is reduced, but the naturalness and flexibility of synthesized speech deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidnaturalness of synthesized speech
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments the prosodic structure prediction into multiple hierarchical levels (intonation phrase level, prosodic phrase level, and prosodic word level), with each level having its own prediction model. This segmentation allows the system to capture different aspects of prosodic structure independently, improving naturalness without requiring a single overly complex uniform model.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces dynamic adjustment mechanisms where prosodic boundary information from higher levels (e.g., intonation phrase boundaries) is used to constrain and guide predictions at lower levels (e.g., prosodic phrase boundaries). This dynamic interaction between hierarchical levels enables the system to adapt to varying speech contexts while maintaining manageable model complexity.

Inventive Principle:
Principle #15Dynamics

2Reliability

If multiple alternative prosodic boundary partitioning solutions are provided, then the flexibility and naturalness of synthesized speech is improved, but the computational complexity and processing time increases

Engineering Contradiction:
Improveflexibility of synthesized speechVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent employs dynamic programming algorithms that efficiently explore multiple prosodic boundary partitioning solutions by breaking down the problem into sub-problems at different hierarchical levels. The dynamic programming approach stores intermediate results and uses them to construct optimal solutions, reducing computational complexity compared to exhaustive search methods while maintaining flexibility.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies different prediction strategies and models to different local regions of the prosodic structure (e.g., different boundary types, different hierarchical levels). By tailoring the prediction approach to local characteristics rather than applying a uniform method throughout, the system achieves high flexibility without proportionally increasing overall computational complexity.

Inventive Principle:
Principle #3Local quality

3Productivity

If prosodic structure prediction is performed without considering speaker characteristics and emotions, then the processing speed is maintained, but the naturalness of synthesized speech deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidnaturalness of synthesized speech
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the prosodic prediction task into hierarchical levels, with each level handling specific aspects of prosodic structure. This segmentation allows the system to process different aspects in parallel or in an optimized sequence, maintaining processing speed while incorporating rich contextual information including speaker characteristics and emotions at appropriate levels.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary processing of speaker characteristics and contextual information during the training phase to create speaker-specific models and contextual embeddings. During actual synthesis, these pre-processed representations are quickly applied, maintaining processing speed while still capturing the influence of speaker characteristics and emotions on prosodic structure.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP2958105B1Method and apparatus for speech synthesis based on large corpus
Publication Date: 2018.04.04 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • EP2958105B1 patent drawingFigure 1~2
  • EP2958105B1 patent drawingFigure 3~4
  • EP2958105B1 patent drawingFigure 5~6

AI summary

The present invention discloses a method and apparatus for speech synthesis based on a large corpus. The method for speech synthesis based on a large corpus comprises: utilizing a prosodic structure prediction model to carry out prosodic structure prediction processing on input text to provide at least one alternative prosodic boundary partitioning solution; determining a prosodic boundary partitioning solution according to structure probability information about a prosodic unit in a speech corpus in the at least one alternative prosodic boundary partitioning solution; and carrying out speech synthesis according to the determined prosodic boundary partitioning solution. The method and apparatus for speech synthesis based on a large corpus provided by the embodiments of the present invention improve the naturalness and flexibility of speech synthesis.