Unsupervised Speech Recognition Adaptation via Deep Autoencoder Bottleneck Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech recognition systems face significant performance degradation in unseen and noisy channel conditions due to acoustic condition mismatch between training and evaluation data, where traditional data-adaptation techniques require labeled data and are ineffective in real-world diverse acoustic environments.

Innovation Solution

The use of robust features extracted from deep autoencoders and deep convolutional networks, specifically bottleneck features and feature-space maximum likelihood linear regression transforms, to enhance speech recognition performance in unseen channel and noise conditions, enabling unsupervised model adaptation without labeled data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional data-adaptation techniques are used to improve speech recognition in unseen conditions, then recognition performance improves, but labeled or transcribed data is required which is often unavailable

Engineering Contradiction:
Improvespeech recognition performanceVSAvoidapplicability in unseen conditions
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system performs self-adaptation by automatically adjusting acoustic models to unseen channel conditions without requiring labeled data. The deep neural network extracts robust features and the system adapts to new conditions autonomously, making the service self-sufficient in environments where labeled data is unavailable.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes parameters by transforming acoustic features through deep neural networks and applying bottleneck features that capture essential speech characteristics while being invariant to channel conditions. This parameter transformation enables the model to generalize across different acoustic environments.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If supervised adaptation techniques are used to achieve substantial performance improvement, then recognition accuracy increases, but labeled data availability is required which limits practical application

Engineering Contradiction:
Improverecognition accuracyVSAvoidlabeled data availability
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent replaces the mechanical requirement for labeled data with a deep learning-based feature extraction system. Instead of relying on supervised adaptation that requires labeled data, the system uses deep neural networks to automatically learn robust acoustic features that work across different conditions without labeled examples.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If feature-transform methods are used for unsupervised adaptation, then applicability in unseen conditions improves, but recognition performance may be reduced compared to supervised techniques

Engineering Contradiction:
Improveunsupervised adaptation capabilityVSAvoidrecognition performance
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system combines multiple feature transformation techniques including deep neural networks, bottleneck features, and traditional feature transforms into a composite approach. This combination maintains the adaptability benefits of unsupervised methods while achieving recognition performance comparable to supervised techniques through the synergistic effect of multiple processing stages.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS11217228B2Systems and methods for speech recognition in unseen and noisy channel conditions
Publication Date: 2022.01.04 SRI INTERNATIONAL
  • US11217228B2 patent drawing
  • US11217228B2 patent drawing
  • US11217228B2 patent drawing

AI summary

Systems and methods for speech recognition are provided. In some aspects, the method comprises receiving, using an input, an audio signal. The method further comprises splitting the audio signal into auditory test segments. The method further comprises extracting, from each of the auditory test segments, a set of acoustic features. The method further comprises applying the set of acoustic features to a deep neural network to produce a hypothesis for the corresponding auditory test segment. The method further comprises selectively performing one or more of: indirect adaptation of the deep neural network and direct adaptation of the deep neural network.