Transform Domain Reconstruction for Noisy Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automatic speech recognition (ASR) systems face challenges in noisy environments, where noise reduction techniques can corrupt the acoustic signal, leading to poor recognition accuracy, as they struggle to adapt recognition models from quiet environments to noisy conditions.

Innovation Solution

The technique involves transform domain reconstruction of noise-corrupted portions of an acoustic signal to emulate speech, using replacement transform values based on speech features like cepstral coefficients or probabilistic models, ensuring the reconstructed signal closely resembles natural speech for improved recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If noise reduction techniques strongly attenuate noise-corrupted portions of the acoustic signal, then voice quality from the listener's perspective is improved, but the transform domain representation becomes dissimilar to natural speech, causing feature extraction to corrupt recognition accuracy

Engineering Contradiction:
Improvenoise qualityVSAvoidrecognition accuracy
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

Instead of suppressing noise-corrupted portions, the patent inverts the approach by reconstructing them to resemble natural speech. The system identifies noise-corrupted transform domain components and replaces them with synthesized components generated from speech portions, thereby improving both noise quality and preserving recognition accuracy.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent creates copies of speech portions to replace noise-corrupted portions. By extracting features from clean speech portions and synthesizing replacement transform values based on these speech characteristics, the system generates artificial speech segments that mimic natural speech, preventing feature corruption while maintaining voice quality.

Inventive Principle:
Principle #26Copying

2Measurement precision

If recognition models are retrained using new speech collected in noisy environments, then recognition accuracy in noisy conditions improves, but the process is time consuming and requires large amounts of new speech data

Engineering Contradiction:
Improverecognition accuracyVSAvoidmodel retraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by reconstructing the acoustic signal in advance before feature extraction and recognition. By pre-processing the noisy signal to replace noise-corrupted portions with synthesized speech portions, the system prepares a cleaned signal that can be processed by existing recognition models without requiring retraining, thus saving time while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the parameters of the acoustic signal in the transform domain by replacing transform values of noise-corrupted portions with synthesized values. This parameter transformation occurs in the frequency domain, allowing the system to modify signal characteristics without changing the underlying recognition models, thereby avoiding time-consuming retraining while improving recognition performance.

Inventive Principle:
Principle #35Parameter changes

3Object-affected harmful factors

If noise-corrupted portions are suppressed rather than reconstructed, then noise attenuation is achieved, but the reconstructed signal fails to resemble natural speech, worsening speech recognition performance

Engineering Contradiction:
Improvenoise levelVSAvoidspeech recognition performance
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent introduces an intermediary process between noise suppression and feature extraction. Instead of directly suppressing noise-corrupted portions, the system uses speech portions as intermediaries to generate replacement transform values. These synthesized components act as mediators that preserve speech-like characteristics while removing noise, ensuring both noise attenuation and reliable speech recognition.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8880396B1Spectrum reconstruction for automatic speech recognition
Publication Date: 2014.11.04 SAMSUNG ELECTRONICS CO LTD
  • US8880396B1 patent drawing
  • US8880396B1 patent drawing
  • US8880396B1 patent drawing

AI summary

The present technology provides techniques for transform domain reconstruction of noise-corrupted portions of an acoustic signal to emulate speech which is obscured by the noise. Replacement transform values for the noise-corrupted portions are determined utilizing the portions of the acoustic signal which contain speech.