Acoustic Environment Profile Estimation for ASR

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Real-world speech processing applications face challenges in estimating acoustic environment parameters due to factors like room reverberation and noise, which degrade speech quality and intelligibility, and existing solutions require a clean speech reference signal that is often unavailable.

Innovation Solution

The proposed solution involves extracting spectral and modulation features from audio signals, combining them to estimate an acoustic environment profile, and using this profile for improved automatic speech recognition (ASR) without requiring a clean speech reference signal, employing non-intrusive signal analysis (NISA) methods to enhance ASR accuracy and reduce error rates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If intrusive signal analysis (ISA) methods are used to estimate acoustic environment parameters, then measurement precision is improved, but device complexity increases due to the requirement for clean speech reference signals

Engineering Contradiction:
Improveacoustic environment parameter estimation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces non-intrusive signal analysis (NISA) as an intermediary approach that eliminates the need for clean speech reference signals. By using NISA, the system can still estimate acoustic environment parameters through the corrupted audio signal itself, combined with acoustic model predictions, thereby reducing system complexity while maintaining estimation accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a virtual copy of the clean speech signal by combining the corrupted audio signal with acoustic model predictions. This synthesized clean signal copy allows the system to perform accurate acoustic environment parameter estimation without requiring actual clean reference signals, thus reducing device complexity.

Inventive Principle:
Principle #26Copying

2Ease of operation

If non-intrusive signal analysis (NISA) methods are used to eliminate clean speech reference requirements, then ease of operation is improved, but measurement precision may deteriorate

Engineering Contradiction:
Improvedeployment flexibilityVSAvoidacoustic environment parameter estimation accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent merges NISA methods with acoustic model predictions to create a hybrid approach. By combining the ease of operation benefits of NISA (no clean reference needed) with the precision benefits of acoustic models, the system achieves both deployment flexibility and accurate parameter estimation simultaneously.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system uses feedback from acoustic model predictions to enhance the NISA estimation process. The acoustic model provides corrective information that refines the parameter estimates derived from the corrupted audio signal, thereby maintaining measurement precision while preserving the operational ease of NISA.

Inventive Principle:
Principle #23Feedback

3Productivity

If acoustic environment profile estimation is performed to improve ASR accuracy, then productivity is improved, but device complexity increases

Engineering Contradiction:
ImproveASR accuracyVSAvoidsignal processing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent performs preliminary acoustic environment profile estimation before the main ASR processing. By pre-characterizing the acoustic environment using NISA and acoustic models, the system prepares environment-specific parameters that can be directly utilized during ASR, thereby improving productivity without proportionally increasing processing complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The acoustic environment profile estimation serves multiple functions: it characterizes the acoustic environment, provides parameters for ASR, and enables adaptive speech processing. This multi-functionality improves ASR productivity while avoiding the need for separate dedicated systems, thus managing device complexity effectively.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240005908A1Acoustic environment profile estimation
Publication Date: 2024.01.04 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20240005908A1 patent drawing
  • US20240005908A1 patent drawing
  • US20240005908A1 patent drawing

AI summary

An acoustic environment profile estimation is provided for automatic speech recognition (ASR) to compensate for the acoustic behavior of an environment in which audio is collected. Examples receive an audio signal and extract spectral features and modulation features. Extracting spectral features involves determining Mel filter bank (MFB) coefficients, and extracting modulation features involves applying Fourier transforms. The spectral features and modulation features are combined, and an acoustic environment profile estimate is extracted and provided as an input to the ASR. In some examples, the acoustic environment profile estimate is realized as acoustic environment parameters, whereas in some other examples, the acoustic environment profile estimate is realized as an acoustic embedding vector. For versions using acoustic environment parameters, when the acoustic environment changes significantly, such as flooring changes and/or speakers or microphones changing position, a new set of acoustic environment parameters is determined.