Pitch-Synchronous Speech Recognition Using Timbre Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current commercial speech recognition systems fail to cleanly separate pitch and timbre information due to unrelated processing window positions, leading to inaccurate phoneme identification as multiple frames often cross phoneme boundaries.
Innovation Solution
The system segments speech signals into pitch-synchronous frames using pitch-marks and Laguerre functions to generate timbre vectors, which are used for parametric representation, allowing for clean separation of timbre and pitch information and unique phoneme identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If traditional window-based speech signal processing is used, then the processing can be simplified and continuous, but pitch information and spectral information cannot be cleanly separated
Solution Approach 1:
The patent segments the speech signal into pitch-synchronous frames based on pitch period detection, where each frame corresponds to one pitch period. This segmentation allows the spectral content to be analyzed independently of pitch variations, achieving clean separation of timbre and pitch information while maintaining processing simplicity through automated pitch-based framing.
2Productivity
If fixed window duration and shift time are used, then the processing is regular and efficient, but multiple frames cross phoneme boundaries reducing recognition accuracy
Solution Approach 1:
The patent dynamically adjusts the frame timing based on the actual pitch period of the speech signal rather than using fixed window positions. By synchronizing frames with pitch periods, the system adapts to varying speech characteristics and ensures that phoneme boundaries align with frame boundaries, improving recognition accuracy without sacrificing processing efficiency.
3Loss of information
If pitch-synchronous framing is used, then timbre information is cleanly separated from pitch information, but the processing complexity increases
Solution Approach 1:
The system uses the pitch signal itself to generate the framing structure, where the pitch period detection automatically provides the timing information needed for frame segmentation. This self-service approach eliminates the need for complex external timing mechanisms or manual parameter adjustment, achieving clean timbre-pitch separation through the pitch signal's own characteristics.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach improves speech recognition accuracy by ensuring each frame represents a unique phoneme identity, enhancing the precision of phoneme sequence determination and subsequent text conversion.
Implementation Method 1
Using Fourier analysis, the speech signal in each frame is converted into a pitch-synchronous amplitude spectrum
Implementation Method 2
Laguerre functions are used to convert the said pitch-synchronous amplitude spectrum into a unit vector characteristic to the instantaneous timbre, referred to as the timbre vector
Data Source
AI summary
The present invention defines a pitch-synchronous parametrical representation of speech signals as the basis of speech recognition, and discloses methods of generating the said pitch-synchronous parametrical representation from speech signals. The speech signal is first going through a pitch-marks picking program to identify the pitch periods. The speech signal is then segmented into pitch-synchronous frames. An ends-matching program equalizes the values at the two ends of the waveform in each frame. Using Fourier analysis, the speech signal in each frame is converted into a pitch-synchronous amplitude spectrum. Using Laguerre functions, the said amplitude spectrum is converted into a unit vector, referred to as the timbre vector. By using a database of correlated phonemes and timbre vectors, the most likely phoneme sequence of an input speech signal can be decoded in the acoustic stage of a speech recognition system.


