Speech Synthesis Model Training Using Pitch Synchronous Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech synthesis technologies face issues with acoustic quality deterioration due to fixed frame rate analysis and unnatural phoneme duration caused by pitch mismatch during training and synthesis.
Innovation Solution
A speech synthesis device and method that employs pitch synchronous analysis to generate acoustic feature parameters, using Hidden Semi-Markov Models (HSMM) to model duration distribution based on timing parameters, ensuring accurate representation of speech waveforms and preventing phoneme duration mismatch.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If fixed frame rate speech analysis is used, then the analysis process is simple, but acoustic quality deteriorates
Solution Approach 1:
The patent changes the fundamental parameter of speech analysis from fixed frame rate to pitch-synchronous analysis. By aligning analysis frames with pitch periods rather than using uniform time intervals, the system achieves both improved acoustic quality and maintained computational efficiency. The analysis frame rate dynamically adapts to the speech signal's pitch characteristics.
2Measurement precision
If pitch synchronous analysis is used, then acoustic quality improves, but phoneme duration becomes unnatural due to pitch mismatch
Solution Approach 1:
The patent introduces a pitch-cycle waveform count as an intermediary parameter that bridges pitch-synchronous analysis and natural phoneme duration. This count represents the number of pitch cycles within each phoneme, allowing the system to maintain pitch-synchronous analysis benefits while ensuring that synthesized phonemes have natural durations by controlling the number of pitch cycles rather than fixed time duration.
3Manufacturing precision
If pitch-cycle waveform count is decided using duration distribution and pitch information, then phoneme duration accuracy improves, but computational complexity increases
Solution Approach 1:
The patent applies preliminary action by pre-training duration distributions from speech corpus data before synthesis. The system learns typical duration patterns for different phonemes and contexts during training, then uses this pre-learned knowledge during synthesis to quickly determine appropriate pitch-cycle waveform counts without complex real-time calculations.
Solution Approach 2:
The system uses feedback mechanisms where the decided pitch-cycle waveform count is evaluated against the output distribution of pitch feature parameters. This feedback loop ensures that the chosen waveform count is consistent with the synthesized pitch information, allowing the system to adjust and refine duration decisions based on the actual pitch characteristics of the generated speech.
Data Source
AI summary
A speech synthesis model training device includes one or more hardware processors configured to perform the following. Storing, in a speech corpus storing unit, speech data, and pitch mark information and context information of the speech data. From the speech data, analyzing acoustic feature parameters at each pitch mark timing in pitch mark information. From the acoustic feature parameters analyzed, training a statistical model which has a plurality of states and which includes an output distribution of acoustic feature parameters including pitch feature parameters and a duration distribution based on timing parameters.


