Speech Speed Variation Using Pitch Period Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional language-learning devices with speed-varying functions distort speech signals when adjusting playback speed, causing echoes and machine noises due to the continuous analog nature of speech signals and differences in voiceprint frequencies.
Innovation Solution
The method involves calculating the pitch period of the original speech signal, dividing it into sections using the Sum of Magnitude Difference Function (SMDF) or Average of Magnitude Difference Function (AMDF), and applying a speed-varying algorithm to each section to facilitate deceleration or acceleration without distortion, using a weighting function to smooth out the signal.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If conventional speed-varying technology is used to adjust playback speed, then the playback speed can be varied, but the speech signal becomes distorted with echoes and machine noises
Solution Approach 1:
The speech signal is divided into multiple frames, with each frame further segmented into pitch periods based on pitch detection. This hierarchical segmentation allows the speed-varying operation to be applied to small, manageable units (pitch periods) rather than the entire continuous signal, thereby maintaining signal integrity while achieving speed variation.
Solution Approach 2:
The invention utilizes the periodic nature of speech signals by detecting pitch periods and using them as the basis for speed variation. By operating on periodic pitch periods rather than continuous analog signal, the method maintains the natural periodic structure of speech while enabling accurate speed control without distortion.
2Speed
If the speech signal is treated as continuous analog signal for speed variation, then speed can be adjusted, but voiceprint frequencies are decreased causing distortion
Solution Approach 1:
The invention replaces the mechanical approach of varying playback speed (which physically stretches or compresses the continuous signal and alters voiceprint frequencies) with a digital signal processing approach. By detecting pitch periods and manipulating discrete pitch period units, the system achieves speed variation while preserving the original voiceprint frequencies in the output signal.
3Manufacturing precision
If pitch period-based segmentation is used for speed variation, then speech quality is maintained, but the processing complexity increases
Solution Approach 1:
The invention performs pitch period detection and frame segmentation as preliminary steps before applying the speed-varying operation. By pre-organizing the speech signal into pitch period units and determining appropriate search ranges in advance, the actual speed variation process becomes simpler and more efficient, as it only needs to operate on the pre-segmented units rather than processing the entire continuous signal.
Data Source
AI summary
A method for varying speech speed is provided. The method includes the following steps: receive an original speech signal; calculate a pitch period of the original speech signal; define search ranges according to the pitch period; find a maximum within each of the search ranges of the original speech signal; divide the original speech signal into speech sections according to the maxima; obtain a speed-varied speech signal by applying a speed-varying algorithm to each speech section of the original speed signal according to a speed-varying command; and eventually, output the speed-varied speech signal.


