Voice Synthesis Device Morphing Intermediate Voice Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice synthesis devices have limited freedom in voice quality conversion, suffer from poor sound quality due to excessive dynamic range processing, and may inaccurately specify reference points for morphing, resulting in poor sound quality and increased calculation complexity.
Innovation Solution
A voice synthesis device that stores multiple voice element information for different qualities, uses a designating unit to set fixed points for morphing, and generates intermediate voice information by calculating intermediate values of characteristic parameters, allowing for greater voice quality freedom without excessive dynamic range processing and reducing calculation complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional voice synthesis devices use preset voice element databases with fixed voice qualities, then the device complexity is reduced, but the freedom of voice quality conversion is limited
Solution Approach 1:
The voice element information is segmented into multiple characteristic parameters (pitch, intensity, timbre, etc.) that can be independently adjusted. This allows the system to create intermediate voice qualities by combining parameters from different preset voice elements, thereby increasing voice quality conversion freedom without requiring a complete redesign of the voice synthesis system.
Solution Approach 2:
The invention changes the approach from selecting entire preset voice elements to adjusting individual characteristic parameters of voice elements. By modifying parameters such as pitch, intensity, and timbre independently, the system can generate intermediate voice qualities that were not previously available in the preset databases, resolving the contradiction between versatility and complexity.
2Adaptability or versatility
If the dynamic range in the spectrum is increased to improve voice quality adjustment, then the freedom of voice quality conversion is improved, but the sound quality deteriorates
Solution Approach 1:
Instead of uniformly increasing the dynamic range across the entire spectrum, the invention applies local adjustments to specific characteristic parameters (pitch, intensity, timbre) at different time points. This localized approach allows for flexible voice quality adjustment while maintaining the natural spectral characteristics and avoiding sound quality deterioration.
3Productivity
If a morphing process is carried out on waveform data with automatically specified reference points, then the processing speed is improved, but the sound quality becomes poor due to incorrect specification
Solution Approach 1:
The invention performs preliminary action by pre-specifying reference points in the waveform data before the morphing process begins. These reference points are carefully selected to correspond to important acoustic features such as peaks and transitions. By preparing these reference points in advance, the system ensures accurate morphing without sacrificing processing speed, as the reference specification does not need to be performed during the actual morphing operation.
4Reliability
If intermediate voice qualities are generated by calculating intermediate values of characteristic parameters, then the sound quality is maintained, but the calculation complexity increases
Solution Approach 1:
The invention simplifies the calculation by focusing changes on a limited set of key characteristic parameters (pitch, intensity, timbre) rather than processing the entire waveform or spectrum. By calculating intermediate values only for these essential parameters, the system maintains sound quality while significantly reducing calculation complexity compared to full waveform or spectral processing.
Data Source
AI summary
A voice synthesis device for generating synthetic voice having great freedom in voice quality and good sound quality from text data is provided.The voice synthesis device is provided with: voice synthesis DBs (110a and 101z); a voice synthesis unit (103) for acquiring a text (10) and generating a voice synthesis parameter value sequence (11) of voice quality A which corresponds to a character that is included in the text (10) from the voice synthesis DB (101a); a voice synthesis unit (103) generating a voice synthesis parameter value sequence (11) of voice quality Z which corresponds to a character that is included in the text (10) from the voice synthesis DB (101z); a voice morphing unit (105) for generating a intermediate voice synthesis parameter value sequence (13) for indicating synthetic voice of intermediate voice quality between voice quality A and voice quality Z which corresponds to a character that is included in the text (10) from the voice synthesis parameter value sequence (11) of voice quality A and voice quality Z; and a speaker (107) for converting the generated intermediate voice synthesis parameter value sequence (13) to the synthetic voice and outputting the resulting synthetic voice.


