Variable Expansion Rate Phoneme Synthesis for Natural Voice
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice synthesis technologies struggle to produce aurally natural voices when expanding phonetic pieces, as they maintain a fixed expansion and contraction rate within phonetic piece ranges, leading to unnatural voice synthesis.
Innovation Solution
A voice synthesis apparatus that adjusts the expansion rate of phonetic pieces based on phoneme types, using different expansion processes for consonant phonemes, and interpolates envelope data for voiced consonants to create synthesized phonetic piece data with varying expansion rates, ensuring natural voice synthesis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a fixed expansion and contraction rate is maintained within a phonetic piece range, then the processing simplicity is improved, but the aural naturalness of the synthesized voice deteriorates
Solution Approach 1:
The patent divides the phonetic piece into multiple sections (normal part and transition part) and applies different expansion/contraction rates to each section. The transition part uses a smaller expansion rate than the normal part, creating local quality variations that match natural speech patterns while maintaining overall processing simplicity.
Solution Approach 2:
The patent introduces dynamic expansion rate adjustment based on the position within the phonetic piece. By making the expansion rate variable (different for normal vs. transition parts) rather than fixed, the system achieves more natural speech synthesis while keeping the adjustment rule-based and computationally manageable.
2Device complexity
If the expansion rate is uniformly applied across all sections of a phonetic piece, then the device complexity is reduced, but the synthesis accuracy deteriorates
Solution Approach 1:
The patent implements local quality by assigning different expansion rates to different sections (normal part vs. transition part) of the phonetic piece. This section-based differentiation improves synthesis accuracy by matching natural speech characteristics without requiring complex global adjustment mechanisms.
Solution Approach 2:
The patent segments the phonetic piece into distinct parts (normal part and transition part) with different expansion characteristics. This segmentation allows for more accurate synthesis by treating different sections differently, while the segmentation itself is simple and based on predefined boundaries.
Data Source
AI summary
A voice signal is synthesized using a plurality of phonetic piece data each indicating a phonetic piece containing at least two phoneme sections corresponding to different phonemes. In the apparatus, a phonetic piece adjustor forms a target section from first and second phonetic pieces so as to connect the first and second phonetic pieces to each other such that the target section includes a rear phoneme section of the first piece and a front phoneme section of the second piece, and expands the target section by a target time length to form an adjustment section such that a central part is expanded at an expansion rate higher than that of front and rear parts of the target section, to thereby create synthesized phonetic piece data having the target time length. A voice synthesizer creates a voice signal from the synthesized phonetic piece data.


