Syllable-Centered Polynomial Pitch Contour Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech synthesis systems struggle to generate natural-sounding prosody, as prior methods result in discontinuous and incomplete pitch signals due to the absence of pitch values in unvoiced consonants and silence, limiting the ability to produce human-like speech.
Innovation Solution
The use of polynomial expansion coefficients for pitch contour representation near syllable centers, combined with interpolation methods to generate continuous prosody parameters, allows for the creation of a parametrical representation of prosody that can be applied to input text, ensuring smooth pitch contours and accurate duration and intensity profiles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If pitch values are only extracted from voiced frames, then the pitch signal is accurate for voiced segments, but the pitch signal becomes discontinuous and incomplete due to absence of pitch values in unvoiced consonants and silence
Solution Approach 1:
The patent performs preliminary action by extracting pitch values not only from voiced frames but also from unvoiced frames and silence segments. This preliminary extraction of pitch information from previously excluded segments enables the construction of a continuous pitch contour that includes all temporal regions of speech, resolving the discontinuity problem while maintaining accuracy through the use of polynomial expansion coefficients fitted to the extracted pitch values.
2Reliability
If polynomial expansion coefficients are used to represent pitch contours near syllable centers, then the pitch contour becomes continuous and complete, but the system complexity increases due to the need for correlation database construction and interpolation calculations
Solution Approach 1:
The patent applies segmentation by dividing the pitch contour representation into syllable-centered polynomial expansion coefficients. Each syllable is processed independently to extract pitch values and compute polynomial coefficients, which are then stored in a correlation database. This segmentation enables systematic organization of complex pitch information and simplifies the synthesis process through database lookup and interpolation, managing system complexity while achieving continuous pitch contours.
3Quantity of substance
If pitch values are interpolated for unvoiced consonants and silence sections, then a complete pitch contour is generated, but the interpolation process requires additional computational steps and parameters
Solution Approach 1:
The patent applies parameter changes by transforming the pitch representation from raw pitch values to polynomial expansion coefficients. This parameter transformation enables efficient interpolation through the correlation database, where pre-computed polynomial coefficients are combined with context information to generate pitch values for unvoiced consonants and silence sections. This approach achieves complete pitch contours while managing computational complexity through pre-computation and efficient database querying.
Data Source
AI summary
The present invention discloses a parametrical representation of prosody based on polynomial expansion coefficients of the pitch contour near the center of each syllable. The said syllable pitch expansion coefficients are generated from a recorded speech database, read from a number of sentences by a reference speaker. By correlating the stress level and context information of each syllable in the text with the polynomial expansion coefficients of the corresponding spoken syllable, a correlation database is formed. To generate prosody for an input text, stress level and context information of each syllable in the text is identified. The prosody is generated by using the said correlation database to find the best set of pitch parameters for each syllable. By adding to global pitch contours and using interpolation formulas, complete pitch contour for the input text is generated. Duration and intensity profile are generated using a similar procedure.


