Automated Distortion Classification for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automatic Speech Recognition (ASR) systems face challenges in accurately recognizing speech due to various types of distortions such as ambient noise, transient noise, and electronics noise, leading to rejection errors, misrecognition, or ineffective noise reduction techniques.

Innovation Solution

An automated distortion classification method and system that processes audio signals to generate acoustic feature vectors, decode them to produce hypotheses for distortions, and post-process these hypotheses to identify and address the specific distortions, improving speech recognition accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real-time noise removal techniques are used, then speech recognition can proceed during actual use, but the techniques are too ineffective or unreliable

Engineering Contradiction:
Improvenoise removal effectivenessVSAvoidspeech recognition speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary distortion classification by analyzing acoustic feature vectors before speech recognition processing. By identifying distortion types in advance and classifying them into categories (e.g., ambient noise, transient noise, electronics noise), the system can prepare appropriate mitigation strategies beforehand, improving reliability without significantly delaying speech recognition processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system continuously monitors the audio signal, classifies distortions in real-time, and feeds this information back to adjust speech recognition parameters dynamically. This feedback loop allows the system to adapt its processing based on the identified distortion types, enhancing both reliability of noise removal and maintenance of speech recognition speed.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If subjective listening techniques are used during development, then distortion scoring can be performed, but the process is too costly or complex and conducted off-line

Engineering Contradiction:
Improvedistortion scoring accuracyVSAvoidevaluation process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system enables self-service distortion classification by automatically analyzing acoustic feature vectors and identifying distortion types without requiring expert human listeners. The automated classifier processes audio signals independently, eliminating the need for costly and complex subjective evaluation processes while maintaining accurate distortion identification through computational analysis of acoustic characteristics.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The invention replaces the mechanical process of human listening and scoring with an automated computational system. By substituting human experts with algorithmic classification based on acoustic feature analysis, the system achieves the same measurement precision without the complexity and cost associated with subjective listening techniques, enabling off-line and real-time processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If automated distortion classification is implemented, then real-time processing is enabled, but additional processing steps are required

Engineering Contradiction:
Improveprocessing speedVSAvoidsystem processing steps
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system merges distortion classification with the existing speech recognition processing pipeline by integrating the classifier into the acoustic feature vector processing stream. The classification operation is combined with the decoding process, allowing both functions to be executed through unified processing steps, thereby enabling real-time operation without proportionally increasing system complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The acoustic feature vector processing system serves multiple functions: it performs speech recognition decoding, distortion classification, and provides inputs for both processes simultaneously. This multi-functionality allows the system to achieve real-time processing by having a single processing stream handle multiple tasks, reducing the overall complexity despite the additional classification capability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8438030B2Automated distortion classification
Publication Date: 2013.05.07 GENERAL MOTORS LLC
  • US8438030B2 patent drawing
  • US8438030B2 patent drawing
  • US8438030B2 patent drawing

AI summary

A method of and system for automated distortion classification. The method includes steps of (a) receiving audio including a user speech signal and at least some distortion associated with the signal; (b) pre-processing the received audio to generate acoustic feature vectors; (c) decoding the generated acoustic feature vectors to produce a plurality of hypotheses for the distortion; and (d) post-processing the plurality of hypotheses to identify at least one distortion hypothesis of the plurality of hypotheses as the received distortion. The system can include one or more distortion models including distortion-related acoustic features representative of various types of distortion and used by a decoder to compare the acoustic feature vectors with the distortion-related acoustic features to produce the plurality of hypotheses for the distortion.