Speech Synthesis System Speed and Pitch Adjustment via Spectrogram Transformation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis methods, such as Unit Selection Synthesis and HMM-based Speech Synthesis, struggle to produce natural-sounding speech that reflects the style and emotional expression of a speaker, limiting their ability to effectively synthesize speech from text.

Innovation Solution

An artificial intelligence-based speech synthesis technique that uses an artificial neural network to convert text into speech, allowing for the natural manipulation of speech speed and pitch by generating spectrograms through short-time Fourier transformation and inverse short-time Fourier transformation, and applying phase information estimation to adjust playback rate and pitch change rates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional speech synthesis methods (Unit Selection Synthesis, HMM-based Speech Synthesis) are used, then speech can be generated from text, but the speech lacks naturalness and cannot effectively reflect speaker style and emotional expression

Engineering Contradiction:
Improvenaturalness of synthesized speechVSAvoidability to reflect speaker style and emotional expression
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent replaces traditional mechanical speech synthesis methods (Unit Selection, HMM-based approaches) with an artificial intelligence-based deep learning system. The system uses neural networks to learn speaker characteristics and emotional expressions from training data, enabling natural-sounding speech synthesis that adapts to different speakers and emotional states without relying on pre-defined phoneme units or statistical models.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent dynamically adjusts speech parameters including playback rate and pitch change rate based on learned speaker characteristics and emotional context. By modifying these parameters adaptively rather than using fixed synthesis rules, the system achieves more natural speech output that reflects the intended emotional expression and speaker style.

Inventive Principle:
Principle #35Parameter changes

2Speed

If speech speed and pitch are adjusted using traditional methods, then playback rate can be changed, but the speech loses naturalness and emotional expression

Engineering Contradiction:
Improveplayback rate and pitch change rateVSAvoidnaturalness of speech
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The patent implements dynamic adjustment of playback rate and pitch change rate based on the emotional content and speaker characteristics learned during training. Rather than applying uniform speed changes, the system dynamically modifies these parameters to maintain natural speech patterns even when overall speed is adjusted, preserving emotional expression through adaptive parameter control.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses learned speaker characteristics and emotional context as feedback to guide parameter adjustment. By continuously referencing the trained model's understanding of natural speech patterns for specific speakers and emotions, the system adjusts playback rate and pitch to maintain naturalness while achieving the desired speed changes.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11776528B2Method for changing speed and pitch of speech and speech synthesis system
Publication Date: 2023.10.03 XINAPSE CO LTD
  • US11776528B2 patent drawing
  • US11776528B2 patent drawing
  • US11776528B2 patent drawing

AI summary

This application relates to a method of synthesizing a speech of which a speed and a pitch are changed. In one aspect, the method includes a spectrogram may be generated by performing a short-time Fourier transformation on a first speech signal based on a first hop length and a first window length, and speech signals of sections having a second window length at the interval of a second hop length from the spectrogram. A ratio between the first hop length and the second hop length may be set to be equal to the value of a playback rate and a ratio between the first window length and the second window length may be set to be equal to the value of a pitch change rate, thereby generating a second speech signal of which the speed and the pitch are changed.