Pitch Modification Ratio Thresholds for Speech Synthesis Speed
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Concatenative speech synthesis techniques face challenges in maintaining high output quality while reducing computational complexity, particularly when modifying pitch at low sampling rates or large pitch changes, leading to inefficiencies in commercial speech synthesizers.
Innovation Solution
The approach involves determining the pitch modification ratio between the current and requested pitches and selectively applying either a complex frequency-domain pitch modification algorithm or a simpler windowing technique based on the pitch difference, with low and high ratio thresholds to optimize processing speed and quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If complex frequency-domain pitch modification algorithms are used, then speech output quality is improved, but computational complexity increases and processing speed decreases
Solution Approach 1:
The patent changes the parameter being processed from the speech signal itself to the pitch ratio parameter. By extracting pitch information and working with pitch ratios rather than raw speech waveforms, the system achieves pitch modification with reduced computational complexity while maintaining quality.
Solution Approach 2:
The patent extracts pitch information from the speech signal as a separate parameter. By taking out the pitch component and modifying it independently, then applying the modification to the speech segments, the system avoids the computational burden of processing the entire speech signal in the frequency domain.
2Speed
If time domain pitch modification techniques are used, then processing speed is improved, but speech output quality deteriorates when pitch changes are large
Solution Approach 1:
The patent changes the approach from modifying the speech signal directly in the time domain to modifying pitch parameters. By working with pitch ratios and applying them as scaling factors to speech segments, the system achieves both speed and quality.
Solution Approach 2:
The patent performs preliminary pitch extraction and ratio calculation before speech segment modification. By determining pitch ratios in advance and preparing modification factors beforehand, the system enables fast time-domain processing while ensuring quality through pre-computed accurate pitch parameters.
3Manufacturing precision
If pitch modification is applied to all speech segments, then speech naturalness is improved, but processing time increases
Solution Approach 1:
The patent applies partial action by selectively modifying pitch based on the pitch ratio. Instead of uniformly processing all speech segments, the system identifies segments requiring pitch modification and applies processing only to those, reducing overall processing time while maintaining naturalness where needed.
Solution Approach 2:
The patent applies local quality by using different processing approaches for different speech segments based on their pitch characteristics. Segments with significant pitch changes receive pitch modification, while segments with minimal pitch changes are processed differently or skipped, optimizing the balance between quality and speed.
Data Source
AI summary
When pitch of a speech segment is being modified from a current pitch to a requested pitch, and the difference between these is relatively large, a pitch modification algorithm is used to modify the pitch of the speech segment. When the difference between current and requested pitches is relatively small, the pitch of the speech segment is not modified. After one or the other speech modification techniques are used, then the resultant modified speech segment is overlapped and added to previously modified speech segments. A modification ratio is determined in order to quantify the difference between the current and requested pitches for a speech segment. The modification ratio is a ratio between the requested and current pitches. Low and high ratio thresholds are used to determine when pitch is being modified to a predetermined high degree, and whether pitch of the speech segment will or will not be modified.


