Prosody Statistic Model Training for Text-to-Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice synthesis systems require manual labeling of corpora for prosody parsing, which is labor-intensive and difficult to control quality, limiting their ability to automatically insert pauses in text-to-speech synthesis without punctuation.

Innovation Solution

A method and apparatus for training a Chinese prosody statistic model using a raw corpus with punctuation, transforming sentences into token sequences, counting token pair frequencies, calculating pause probabilities, and constructing a prosody statistic model to automatically determine pause positions for insertion in text-to-speech synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labeling is used for prosody parsing, then parsing accuracy can be controlled, but the process becomes labor-intensive and time-consuming

Engineering Contradiction:
Improveprosody parsing accuracyVSAvoidmanual labeling time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system uses punctuation marks in the text itself to automatically determine pause positions, eliminating the need for manual labeling. The prosody parsing module extracts pause information directly from punctuation patterns, allowing the system to serve itself without external manual intervention while maintaining reasonable parsing accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The invention creates a prosody statistic model that copies pause patterns from punctuated text. Instead of manually labeling each text, the system learns from the distribution of punctuation marks in training data and applies these learned patterns to new texts, significantly reducing manual labeling requirements

Inventive Principle:
Principle #26Copying

2Reliability

If manual prosody labeling is performed on corpus, then training data quality can be ensured, but the work becomes arduous and hard to control quality

Engineering Contradiction:
Improvetraining data qualityVSAvoidcorpus processing ease
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The corpus processing system automatically extracts pause information from punctuation marks without requiring manual annotation. The prosody statistic model learns directly from the punctuated corpus, making the system self-sufficient and eliminating the arduous manual labeling process while maintaining data quality through statistical learning

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the approach from manual quality control to statistical parameter-based quality assurance. By using pause probabilities and statistical models derived from punctuation patterns, the system ensures consistent training data quality without relying on manual annotation consistency

Inventive Principle:
Principle #35Parameter changes

3Extent of automation

If rule-learning algorithms are used for prosody prediction, then automatic parsing is achieved, but large amounts of manually labeled corpus are required

Engineering Contradiction:
Improveprosody parsing automationVSAvoidlabeled corpus volume
Core Design Contradiction:
Extent of automationVSQuantity of substance

Solution Approach 1:

Instead of requiring large amounts of manually labeled corpus for rule-learning, the system copies pause patterns directly from punctuation marks in unpunctuated or lightly-punctuated text. The prosody statistic model learns from the natural distribution of punctuation in the corpus, dramatically reducing the need for manually labeled training data while maintaining automation

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The punctuation marks in the text serve multiple functions: they indicate both grammatical structure and prosodic pause positions. By leveraging this universal property of punctuation, the system achieves automatic prosody parsing without requiring separate manual prosody labeling, reducing the quantity of labeled corpus needed

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8024174B2Method and apparatus for training a prosody statistic model and prosody parsing, method and system for text to speech synthesis
Publication Date: 2011.09.20 TOSHIBA DIGITAL SOLUTIONS CORP
  • US8024174B2 patent drawing
  • US8024174B2 patent drawing
  • US8024174B2 patent drawing

AI summary

The present invention provides a method and apparatus for training a prosody statistic model and prosody parsing, a method and system for text to speech synthesis. Said method for training a prosody statistic model with a raw corpus that includes a plurality of sentences with punctuation, comprising: transforming said plurality of sentences in said raw corpus into a plurality of token sequences respectively; counting a frequency for each adjacent token pair occurring in said plurality of token sequences and frequencies of punctuation that represents a pause occurring at associated positions of said each token pair; calculating pause probabilities at said associated positions of said each token pair; and constructing said prosody statistic model based on said token pairs and said pause probabilities at associated positions thereof. With the present invention a prosody statistic model can be trained from a raw corpus without manually prosody parsing tags. And the prosody statistic model can be used in the prosody parsing and further voice synthesis.