Prosody Pattern Normalization for Natural Speech Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech synthesizing technologies using hidden Markov models (HMMs) face high computational costs and inability to generate prosody patterns sequentially, leading to delayed speech output and reduced naturalness due to reliance on optimal parameter searching and global fundamental frequency distributions.
Innovation Solution
A prosody-pattern generating apparatus that includes an initial-prosody-pattern generating unit, normalization-parameter generating unit, and prosody-pattern normalizing unit, which uses language information and prosody models to generate and normalize prosody patterns in units of phonemes, syllables, and words, allowing for sequential output and improved naturalness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If optimal parameter strings are searched for by repeatedly using algorithms, then the naturalness of fundamental frequency pattern is improved, but the amount of calculation increases
Solution Approach 1:
The patent pre-calculates and stores normalization parameters (mean values and standard deviations) for different phoneme types during an offline training phase. These parameters are then directly applied during online speech synthesis without requiring repeated optimization algorithms, thus maintaining naturalness while reducing computational burden.
Solution Approach 2:
The patent transforms the fundamental frequency values by applying normalization using pre-computed mean and standard deviation parameters specific to each phoneme type. This parameter transformation approach replaces complex iterative optimization with simple statistical normalization, achieving natural prosody with reduced calculation.
2Manufacturing precision
If the distribution of fundamental frequencies of the entire text sentence is employed, then the naturalness is improved, but the speech cannot be output until the fundamental frequency pattern of the entire text is completed
Solution Approach 1:
The patent divides the speech synthesis process into independent phoneme-level segments. Each phoneme's fundamental frequency is normalized using pre-computed parameters specific to that phoneme type, allowing sequential processing and immediate output as each phoneme is synthesized, rather than waiting for the entire sentence.
Solution Approach 2:
The normalization parameters for each phoneme type are pre-computed during offline training based on the distribution of fundamental frequencies in the training corpus. This preliminary preparation enables online synthesis to proceed phoneme-by-phoneme without requiring global sentence-level optimization, thus enabling real-time output.
Data Source
AI summary
Normalization parameters are generated at a normalization-parameter generating unit by calculating the mean values and the standard deviations of an initial prosody pattern and a prosody pattern of a training sentence of a speech corpus. Then, the variance range or variance width of the initial prosody pattern is normalized at the prosody-pattern normalizing unit in accordance with the normalization parameters. As a result, a prosody pattern similar to speech of human beings and improved in naturalness can be generated with a small amount of calculation.


