Pitch Equalization Processor for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automatic speech recognition systems face challenges due to the speaker-dependence of Mel-frequency spectral coefficients and Mel-frequency cepstral coefficients, which vary with age, gender, and emotion, leading to complexity and the need for vast training data and complex algorithms to handle pitch variations.

Innovation Solution

A pre-processing system comprising a pitch estimation circuit and a pitch equalization processor that determines and equalizes the speech pitch, providing a pitch-equalized speech signal to improve recognition, reducing system complexity and enabling robustness against noise and variability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If Mel-frequency spectral coefficients (MFSC) or Mel-frequency cepstral coefficients (MFCC) are used for speech recognition, then speech recognition functionality is achieved, but the system becomes highly speaker-dependent requiring complex models and vast training data

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts and removes the pitch component from the speech signal before processing. By using a pitch detection algorithm to identify pitch contours and then subtracting the pitch component from the original signal, the system eliminates the primary source of speaker-dependency, allowing the remaining signal to be processed with simpler, more universal models

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms the speech signal by removing the pitch parameter, fundamentally changing the signal characteristics. This parameter change converts speaker-dependent spectral features into speaker-independent features, enabling the use of less complex acoustic models while maintaining recognition accuracy

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If complex algorithms with large number of layers and weights are used to handle speaker-dependency, then speech recognition performance is maintained across different speakers, but computational resources and training data requirements increase significantly

Engineering Contradiction:
Improvespeaker independenceVSAvoidcomputational resources
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent performs pitch removal as a preliminary action before the main speech recognition processing. By pre-processing the speech signal to eliminate pitch variations, the system prepares the data in advance, allowing subsequent recognition algorithms to operate with reduced complexity and lower computational requirements

Inventive Principle:
Principle #10Preliminary action

3Reliability

If vast amounts of training data are used to compensate for speaker-dependency in neural networks, then recognition accuracy across different speakers improves, but system training time and resource requirements increase

Engineering Contradiction:
Improvecross-speaker recognition accuracyVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

By extracting and removing the pitch component that causes speaker-dependency, the patent reduces the variability that neural networks would otherwise need to learn from extensive training data. This extraction approach allows the system to achieve good cross-speaker performance with less training data and shorter training time

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11270721B2Systems and methods of pre-processing of speech signals for improved speech recognition
Publication Date: 2022.03.08 HEWLETT PACKARD DEVELOPMENT COMPANY LP
  • US11270721B2 patent drawing
  • US11270721B2 patent drawing
  • US11270721B2 patent drawing

AI summary

Pre-processing systems, methods of pre-processing, and speech processing systems for improved Automated Speech Recognition are provided. Some pre-processing systems for improved speech recognition of a speech signal are provided, which systems comprise a pitch estimation circuit; and a pitch equalization processor. The pitch estimation circuit is configured to receive the speech signal to determine a pitch index of the speech signal, and the pitch equalization processor is configured to receive the speech signal and pitch information, to equalize a speech pitch of the speech signal using the pitch information, and to provide a pitch-equalized speech signal.