Predictive Feature Extraction for Noisy Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional speech recognition systems face challenges in information extraction due to recognition errors and the need for extensive training data, especially in noisy environments, leading to reduced accuracy and coverage.

Innovation Solution

A novel predictive feature extraction method combining linguistic and statistical information, using Information Weighted Singular Value Decomposition (IWSVD) and hierarchical classification, along with a mixed approach of machine learning and rule-based systems to improve classification precision and recall.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional speech recognition systems are used for information extraction, then the system can process speech input, but recognition accuracy deteriorates under noisy environmental conditions and with scarce training data

Engineering Contradiction:
Improverecognition accuracyVSAvoidnoisy environmental conditions
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent segments the feature extraction process into multiple distinct feature types (acoustic features, linguistic features, phonetic features) rather than relying on a single feature set. This segmentation allows the system to process speech under noisy conditions by combining multiple complementary feature representations, where each feature type captures different aspects of the speech signal that remain robust to specific types of noise.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a composite feature representation by combining multiple different feature types (acoustic, linguistic, phonetic) into a unified model. This composite approach is analogous to using composite materials in engineering, where combining different materials with complementary properties creates a system that is more robust to various environmental conditions than any single material alone.

Inventive Principle:
Principle #40Composite materials

2Reliability

If traditional speech recognition systems are used for information extraction, then the system can operate with limited resources, but classification accuracy deteriorates when training data is scarce

Engineering Contradiction:
Improveclassification accuracyVSAvoidtraining data
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent performs preliminary feature extraction and representation learning to create rich, structured feature descriptors before the actual classification task. By pre-processing the speech data into comprehensive feature representations that encode acoustic, linguistic, and phonetic information, the system prepares the data in advance to maximize information content, allowing accurate classification even with limited training examples.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the parameter representation by transforming raw speech signals into multiple types of feature spaces (acoustic features, linguistic features, phonetic features) with different characteristics. This parameter transformation allows the system to capture speech information from multiple perspectives, improving classification accuracy by representing the same data in ways that are more discriminative and robust to variations.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If rule-based approaches are used for information extraction, then the system is robust to recognition error, but coverage and accuracy are diminished due to rule ambiguity and insufficient rules

Engineering Contradiction:
Improverobustness to recognition errorVSAvoidcoverage
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent introduces an intermediary layer of multiple feature types between the raw speech recognition output and the final classification decision. This intermediary feature representation layer acts as a mediator that translates recognition output into a form that is more robust to errors while expanding coverage. The multiple feature types serve as intermediaries that bridge the gap between noisy recognition output and accurate classification, allowing the system to handle both robustness and coverage requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8583416B2Robust information extraction from utterances
Publication Date: 2013.11.12 NANT HOLDINGS IP LLC
  • US8583416B2 patent drawing
  • US8583416B2 patent drawing
  • US8583416B2 patent drawing

AI summary

The performance of traditional speech recognition systems (as applied to information extraction or translation) decreases significantly with, larger domain size, scarce training data as well as under noisy environmental conditions. This invention mitigates these problems through the introduction of a novel predictive feature extraction method which combines linguistic and statistical information for representation of information embedded in a noisy source language. The predictive features are combined with text classifiers to map the noisy text to one of the semantically or functionally similar groups. The features used by the classifier can be syntactic, semantic, and statistical.