Speech Synthesis System Speed and Pitch Adjustment via Spectrogram Transformation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis methods, such as Unit Selection Synthesis and HMM-based Speech Synthesis, struggle to produce natural-sounding speech that reflects the style and emotional expression of a speaker, limiting their ability to effectively synthesize speech from text.
Innovation Solution
An artificial intelligence-based speech synthesis technique that uses an artificial neural network to convert text into speech, allowing for the natural manipulation of speech speed and pitch by generating spectrograms through short-time Fourier transformation and inverse short-time Fourier transformation, and applying phase information estimation to adjust playback rate and pitch change rates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional speech synthesis methods (Unit Selection Synthesis, HMM-based Speech Synthesis) are used, then speech can be generated from text, but the speech lacks naturalness and cannot effectively reflect speaker style and emotional expression
Solution Approach 1:
The patent replaces traditional mechanical speech synthesis methods (Unit Selection, HMM-based approaches) with an artificial intelligence-based deep learning system. The system uses neural networks to learn speaker characteristics and emotional expressions from training data, enabling natural-sounding speech synthesis that adapts to different speakers and emotional states without relying on pre-defined phoneme units or statistical models.
Solution Approach 2:
The patent dynamically adjusts speech parameters including playback rate and pitch change rate based on learned speaker characteristics and emotional context. By modifying these parameters adaptively rather than using fixed synthesis rules, the system achieves more natural speech output that reflects the intended emotional expression and speaker style.
2Speed
If speech speed and pitch are adjusted using traditional methods, then playback rate can be changed, but the speech loses naturalness and emotional expression
Solution Approach 1:
The patent implements dynamic adjustment of playback rate and pitch change rate based on the emotional content and speaker characteristics learned during training. Rather than applying uniform speed changes, the system dynamically modifies these parameters to maintain natural speech patterns even when overall speed is adjusted, preserving emotional expression through adaptive parameter control.
Solution Approach 2:
The system uses learned speaker characteristics and emotional context as feedback to guide parameter adjustment. By continuously referencing the trained model's understanding of natural speech patterns for specific speakers and emotions, the system adjusts playback rate and pitch to maintain naturalness while achieving the desired speed changes.
Data Source
AI summary
This application relates to a method of synthesizing a speech of which a speed and a pitch are changed. In one aspect, the method includes a spectrogram may be generated by performing a short-time Fourier transformation on a first speech signal based on a first hop length and a first window length, and speech signals of sections having a second window length at the interval of a second hop length from the spectrogram. A ratio between the first hop length and the second hop length may be set to be equal to the value of a playback rate and a ratio between the first window length and the second window length may be set to be equal to the value of a pitch change rate, thereby generating a second speech signal of which the speed and the pitch are changed.


