Speech Synthesis Prosodic Parameter Normalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text-to-speech synthesis systems struggle to maintain the personality of a speaker when interpolating prosodic parameters, particularly when interpolating between parameters of different speakers, leading to inappropriate results due to large differences in feature amounts.
Innovation Solution
A speech synthesis apparatus that includes a text analysis unit, a dictionary storage unit, a prosodic parameter generation unit, a normalization unit, and a prosodic parameter interpolation unit, which normalizes and interpolates prosodic parameters based on weight information to generate synthesized speech that maintains the personality of the target speaker.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If interpolation processing is performed between prosodic parameters of different speakers (e.g., male and female), then synthesized speech can achieve varied prosodic features, but the personality of the target speaker is lost due to large differences in feature amounts
Solution Approach 1:
The patent applies parameter changes by normalizing prosodic parameters (F0, phoneme duration, power) to a common statistical distribution before interpolation. This involves transforming parameters using z-score normalization or similar techniques to eliminate speaker-specific scale differences, enabling meaningful interpolation while preserving target speaker characteristics.
Solution Approach 2:
The patent introduces a normalization process as an intermediary step between parameter extraction and interpolation. This intermediary transformation aligns the statistical properties of prosodic parameters across different speakers, creating a common reference frame that enables accurate interpolation without losing speaker identity.
2Productivity
If prosodic parameters are directly interpolated without normalization, then the synthesis process is simple and fast, but the interpolation result becomes inappropriate when feature amounts differ significantly
Solution Approach 1:
The patent implements preliminary action by performing normalization of prosodic parameters before the interpolation step. This pre-processing aligns the statistical characteristics of parameters from different speakers, ensuring that subsequent interpolation produces accurate and appropriate results without requiring complex iterative adjustments.
Data Source
AI summary
According to one embodiment, a speech synthesis apparatus is provided with generation, normalization, interpolation and synthesis units. The generation unit generates a first parameter using a prosodic control dictionary of a target speaker and one or more second parameters using a prosodic control dictionary of one or more standard speakers based on language information for an input text. The normalization unit normalizes the one or more second parameters based a normalization parameter. The interpolation unit interpolates the first parameter and the one or more normalized second parameters based on weight information to generate a third parameter and the synthesis unit generates synthesized speech using the third parameter.


