Speech Recognition Channel Normalization via Initial Utterance Statistics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automatic speech recognition systems face performance variability due to communication channel inconsistencies, such as speaker and microphone characteristics, and the challenge of estimating channel characteristics without introducing significant system delays or requiring explicit speech and non-speech differentiation.

Innovation Solution

A method that estimates feature normalization parameters by measuring energy statistics from an initial portion of a speech utterance, using a statistically derived mapping to reduce the amount of speech needed for channel estimation and minimize system delays, without explicitly distinguishing between speech and non-speech segments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If utterance-based normalization is used to estimate channel characteristics, then the estimation accuracy is improved, but system delay increases

Engineering Contradiction:
Improvechannel characteristics estimation accuracyVSAvoidsystem delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary channel estimation using only the initial portion of the speech utterance (first N frames) before the complete utterance is processed. This preliminary action provides early channel characteristics without waiting for the entire utterance to end, thereby reducing system delay while maintaining reasonable estimation accuracy.

Inventive Principle:
Principle #10Preliminary action

2Stability of the object's composition

If a long time window is used for channel estimation, then the estimation stability is improved, but the channel stationarity assumption is violated

Engineering Contradiction:
Improveestimation stabilityVSAvoidchannel stationarity assumption
Core Design Contradiction:
Stability of the object's compositionVSReliability

Solution Approach 1:

The system dynamically adjusts the time window size N based on the actual speech signal characteristics. The window size is determined by analyzing the energy distribution and statistical properties of the speech frames, allowing it to adapt to different speaking rates and channel conditions. This dynamic adjustment maintains the stationarity assumption while providing stable estimates.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If speech and non-speech segments are explicitly differentiated, then the channel estimation accuracy is improved, but the system complexity increases

Engineering Contradiction:
Improvechannel estimation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system uses the speech signal itself to automatically determine the appropriate time window size and processing parameters through statistical analysis of the energy distribution. The speech signal serves its own purpose in guiding the processing, eliminating the need for separate speech detection mechanisms and reducing overall system complexity.

Inventive Principle:
Principle #25Self-service

4Loss of time

If the time window for channel estimation is shortened, then system delay is reduced, but the estimation accuracy deteriorates

Engineering Contradiction:
Improvesystem delayVSAvoidchannel characteristics estimation accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The system changes the parameter N (time window size) based on the statistical properties of the speech signal. By analyzing the energy distribution and adjusting N accordingly, the system optimizes the balance between estimation accuracy and processing speed for each specific speech utterance, achieving both low delay and high accuracy.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS7797157B2Automatic speech recognition channel normalization based on measured statistics from initial portions of speech utterances
Publication Date: 2010.09.14 CERENCE OPERATING CO
  • US7797157B2 patent drawing
  • US7797157B2 patent drawing

AI summary

Channel normalization for automatic speech recognition is provided. Statistics are measured from an initial portion of a speech utterance. Feature normalization parameters are estimated based on the measured statistics and a statistically derived mapping relating measured statistics and feature normalization parameters. In some examples, the measured statistics comprise measures of an energy from the initial portion of the speech utterance. In some examples, measures of the energy comprise extreme values of the energy.