Joint Additive Convolutive Distortion Compensation ASR

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automatic speech recognition (ASR) systems face challenges in noisy environments due to additive and convolutive distortions, such as background noise and microphone variations, especially in mobile applications where computing resources are limited.

Innovation Solution

A system and method for joint additive and convolutive distortion adaptation, using an additive distortion factor estimator, an acoustic model compensator, an utterance aligner, and a convolutive distortion factor estimator to estimate and compensate for distortions in acoustic models, particularly suitable for digital signal processors in mobile applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional speech recognition systems are used in noisy environments, then they can process speech signals, but they suffer from high word error rates due to additive and convolutive distortions

Engineering Contradiction:
Improveword error rateVSAvoidadditive and convolutive distortions
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent segments the distortion compensation into two independent parts: additive distortion compensation and convolutive distortion compensation. This is achieved by separately estimating additive noise parameters and convolutive channel parameters, then applying compensation for each type of distortion independently in the acoustic model, thereby effectively addressing both harmful factors without mutual interference

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameters of the acoustic model by introducing distortion compensation parameters (additive noise parameters and convolutive channel parameters) that are estimated from the speech signal. These parameter changes allow the acoustic model to adapt to noisy and distorted conditions, significantly improving reliability in environments with additive and convolutive distortions

Inventive Principle:
Principle #35Parameter changes

2Reliability

If joint compensation of additive and convolutive distortions is implemented, then robust performance in noisy environments is achieved, but computing complexity increases

Engineering Contradiction:
Improverobust performance in noisy environmentsVSAvoidcomputing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the complex joint compensation task into separate additive distortion compensation and convolutive distortion compensation modules. Each module independently estimates and compensates for its specific type of distortion, reducing the overall computational complexity compared to a unified approach that would need to handle both distortions simultaneously

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary estimation of additive noise parameters and convolutive channel parameters from the speech signal before applying compensation in the acoustic model. This preliminary action separates the computationally intensive parameter estimation from the recognition process, allowing efficient implementation in resource-constrained mobile devices

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If distortion compensation techniques are applied, then recognition accuracy improves in noisy environments, but the system requires more computing resources

Engineering Contradiction:
Improverecognition accuracyVSAvoidcomputing resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary estimation of distortion parameters (additive noise and convolutive channel parameters) from the speech signal before the recognition process. This preliminary action captures the essential distortion characteristics without requiring excessive computing resources during actual recognition, maintaining accuracy while conserving energy in mobile applications

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the compensation process into independent additive and convolutive components, each with its own efficient parameter estimation method. This segmentation allows the system to apply only the necessary computational effort for each distortion type present in the environment, optimizing the balance between recognition accuracy and resource consumption

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS7584097B2System and method for noisy automatic speech recognition employing joint compensation of additive and convolutive distortions
Publication Date: 2009.09.01 TEXAS INSTRUMENTS INC
  • US7584097B2 patent drawing
  • US7584097B2 patent drawing
  • US7584097B2 patent drawing

AI summary

A system for, and method of, noisy automatic speech recognition employing joint compensation of additive and convolutive distortions and a digital signal processor incorporating the system or the method. In one embodiment, the system includes: (1) an additive distortion factor estimator configured to estimate an additive distortion factor, (2) an acoustic model compensator coupled to the additive distortion factor estimator and configured to use estimates of a convolutive distortion factor and the additive distortion factor to compensate acoustic models and recognize a current utterance, (3) an utterance aligner coupled to the acoustic model compensator and configured to align the current utterance using recognition output and (4) a convolutive distortion factor estimator coupled to the utterance aligner and configured to estimate an updated convolutive distortion factor based on the current utterance using differential terms but disregarding log-spectral domain variance terms.