Speech Synthesis Prosody Modification via Two-Path Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Waveform concatenation speech synthesis technologies face challenges in producing accurate and natural prosody, particularly in Japanese, due to inconsistencies in pitch accent and context-dependent frequency changes, leading to unnatural synthesized speech.
Innovation Solution
A two-path search approach is implemented, combining speech segment selection and prosody modification using a statistical model of prosody variations to ensure consistent prosody, with priority given to continuous speech segments to maintain high sound quality and natural accent.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If speech segments are selected based on minimizing cost, then the synthesis process is efficient, but prosody consistency is lost
Solution Approach 1:
The patent introduces a prosody modification value as an additional parameter to the traditional cost minimization approach. By adjusting this modification value, the system can control the degree of prosody modification while maintaining synthesis efficiency, thus resolving the contradiction between productivity and prosody consistency
Solution Approach 2:
The system dynamically adjusts the prosody modification value based on the specific synthesis requirements and input text. This dynamic adjustment allows the system to adapt prosody consistency to different contexts while maintaining efficient synthesis, addressing the contradiction between fixed efficiency and flexible consistency
2Manufacturing precision
If continuous speech segments are used directly, then sound quality is high, but accent accuracy deteriorates due to context-dependent frequency changes
Solution Approach 1:
The patent applies different processing strategies to different parts of the speech segments. Continuous speech segments maintain their original high sound quality where appropriate, while discrete segments undergo prosody modification to ensure accent accuracy. This localized differentiation resolves the contradiction between sound quality and accent accuracy
Solution Approach 2:
The system segments the speech into continuous and discrete portions based on their characteristics. By treating these segments differently - preserving continuous segments for high sound quality while modifying discrete segments for accent accuracy - the system resolves the contradiction between overall sound quality and accent precision
3Measurement precision
If prosody is modified to improve accent accuracy, then accent precision improves, but sound quality deteriorates
Solution Approach 1:
The system applies prosody modification partially only to the extent necessary to achieve accent accuracy, rather than applying it excessively to all segments. By controlling the modification value to apply changes only where needed, the system improves accent accuracy while minimizing degradation of sound quality
Solution Approach 2:
The prosody modification value serves as a control parameter that adjusts the degree of prosody modification. By optimizing this parameter, the system achieves the minimum necessary accent accuracy while preserving sound quality, resolving the contradiction between precision and quality
Data Source
AI summary
Waveform concatenation speech synthesis with high sound quality. Prosody with both high accuracy and high sound quality is achieved by performing a two-path search including a speech segment search and a prosody modification value search. An accurate accent is secured by evaluating the consistency of the prosody by using a statistical model of prosody variations (the slope of fundamental frequency) for both of two paths of the speech segment selection and the modification value search. In the prosody modification value search, a prosody modification value sequence that minimizes a modified prosody cost is searched for. This allows a search for a modification value sequence that can increase the likelihood of absolute values or variations of the prosody to the statistical model as high as possible with minimum modification values.


