Prosody Pattern Normalization for Natural Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesizing technologies using hidden Markov models (HMMs) face high computational costs and inability to generate prosody patterns sequentially, leading to delayed speech output and reduced naturalness due to reliance on optimal parameter searching and global fundamental frequency distributions.

Innovation Solution

A prosody-pattern generating apparatus that includes an initial-prosody-pattern generating unit, normalization-parameter generating unit, and prosody-pattern normalizing unit, which uses language information and prosody models to generate and normalize prosody patterns in units of phonemes, syllables, and words, allowing for sequential output and improved naturalness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If optimal parameter strings are searched for by repeatedly using algorithms, then the naturalness of fundamental frequency pattern is improved, but the amount of calculation increases

Engineering Contradiction:
Improvenaturalness of fundamental frequency patternVSAvoidamount of calculation
Core Design Contradiction:
Manufacturing precisionVSPower

Solution Approach 1:

The patent pre-calculates and stores normalization parameters (mean values and standard deviations) for different phoneme types during an offline training phase. These parameters are then directly applied during online speech synthesis without requiring repeated optimization algorithms, thus maintaining naturalness while reducing computational burden.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms the fundamental frequency values by applying normalization using pre-computed mean and standard deviation parameters specific to each phoneme type. This parameter transformation approach replaces complex iterative optimization with simple statistical normalization, achieving natural prosody with reduced calculation.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If the distribution of fundamental frequencies of the entire text sentence is employed, then the naturalness is improved, but the speech cannot be output until the fundamental frequency pattern of the entire text is completed

Engineering Contradiction:
ImprovenaturalnessVSAvoidoutput delay
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent divides the speech synthesis process into independent phoneme-level segments. Each phoneme's fundamental frequency is normalized using pre-computed parameters specific to that phoneme type, allowing sequential processing and immediate output as each phoneme is synthesized, rather than waiting for the entire sentence.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The normalization parameters for each phoneme type are pre-computed during offline training based on the distribution of fundamental frequencies in the training corpus. This preliminary preparation enables online synthesis to proceed phoneme-by-phoneme without requiring global sentence-level optimization, thus enabling real-time output.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8046225B2Prosody-pattern generating apparatus, speech synthesizing apparatus, and computer program product and method thereof
Publication Date: 2011.10.25 TOSHIBA DIGITAL SOLUTIONS CORP
  • US8046225B2 patent drawing
  • US8046225B2 patent drawing
  • US8046225B2 patent drawing

AI summary

Normalization parameters are generated at a normalization-parameter generating unit by calculating the mean values and the standard deviations of an initial prosody pattern and a prosody pattern of a training sentence of a speech corpus. Then, the variance range or variance width of the initial prosody pattern is normalized at the prosody-pattern normalizing unit in accordance with the normalization parameters. As a result, a prosody pattern similar to speech of human beings and improved in naturalness can be generated with a small amount of calculation.