Pitch Equalization Processor for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automatic speech recognition systems face challenges due to the speaker-dependence of Mel-frequency spectral coefficients and Mel-frequency cepstral coefficients, which vary with age, gender, and emotion, leading to complexity and the need for vast training data and complex algorithms to handle pitch variations.
Innovation Solution
A pre-processing system comprising a pitch estimation circuit and a pitch equalization processor that determines and equalizes the speech pitch, providing a pitch-equalized speech signal to improve recognition, reducing system complexity and enabling robustness against noise and variability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If Mel-frequency spectral coefficients (MFSC) or Mel-frequency cepstral coefficients (MFCC) are used for speech recognition, then speech recognition functionality is achieved, but the system becomes highly speaker-dependent requiring complex models and vast training data
Solution Approach 1:
The patent extracts and removes the pitch component from the speech signal before processing. By using a pitch detection algorithm to identify pitch contours and then subtracting the pitch component from the original signal, the system eliminates the primary source of speaker-dependency, allowing the remaining signal to be processed with simpler, more universal models
Solution Approach 2:
The patent transforms the speech signal by removing the pitch parameter, fundamentally changing the signal characteristics. This parameter change converts speaker-dependent spectral features into speaker-independent features, enabling the use of less complex acoustic models while maintaining recognition accuracy
2Adaptability or versatility
If complex algorithms with large number of layers and weights are used to handle speaker-dependency, then speech recognition performance is maintained across different speakers, but computational resources and training data requirements increase significantly
Solution Approach 1:
The patent performs pitch removal as a preliminary action before the main speech recognition processing. By pre-processing the speech signal to eliminate pitch variations, the system prepares the data in advance, allowing subsequent recognition algorithms to operate with reduced complexity and lower computational requirements
3Reliability
If vast amounts of training data are used to compensate for speaker-dependency in neural networks, then recognition accuracy across different speakers improves, but system training time and resource requirements increase
Solution Approach 1:
By extracting and removing the pitch component that causes speaker-dependency, the patent reduces the variability that neural networks would otherwise need to learn from extensive training data. This extraction approach allows the system to achieve good cross-speaker performance with less training data and shorter training time
Data Source
AI summary
Pre-processing systems, methods of pre-processing, and speech processing systems for improved Automated Speech Recognition are provided. Some pre-processing systems for improved speech recognition of a speech signal are provided, which systems comprise a pitch estimation circuit; and a pitch equalization processor. The pitch estimation circuit is configured to receive the speech signal to determine a pitch index of the speech signal, and the pitch equalization processor is configured to receive the speech signal and pitch information, to equalize a speech pitch of the speech signal using the pitch information, and to provide a pitch-equalized speech signal.


