Adaptive Multichannel Dereverberation for ASR Latency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for mitigating reverberation in audio signals for automated assistants suffer from latency and performance degradation, especially when the audio source is non-stationary, leading to poor automatic speech recognition and delayed responses.

Innovation Solution

An adaptive multichannel technique using an online and causal filter that updates and applies dereverberation to audio data in real-time, utilizing multiple microphones to mitigate reverberation and adapt to changes in the Room Impulse Response, while maximizing the source signal-to-reverberation ratio and compensating for spectral nulls.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional dereverberation techniques are used, then reverberation mitigation is achieved, but latency increases and response time is delayed

Engineering Contradiction:
Improvedereverberation qualityVSAvoidlatency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by continuously updating the Room Impulse Response (RIR) model during the utterance using causal filtering, so that when dereverberation is needed, the model is already prepared and updated, reducing the time required for processing

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adapts the RIR model in real-time during the utterance using an online causal filter that continuously updates as new audio data arrives, allowing the system to maintain high dereverberation quality without waiting for the entire utterance to be captured

Inventive Principle:
Principle #15Dynamics

2Reliability

If traditional dereverberation techniques are used, then reverberation is reduced, but performance degrades when the audio source is non-stationary

Engineering Contradiction:
Improvedereverberation qualityVSAvoidadaptation to non-stationary sources
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system employs a dynamic adaptive filter that continuously updates the RIR model during the utterance, allowing it to track and adapt to changes in the audio source characteristics in real-time, making it effective for non-stationary sources

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses feedback from the incoming audio signal to continuously update the RIR model through the causal filter, allowing the system to adapt to changing source characteristics by using the most recent audio information available

Inventive Principle:
Principle #23Feedback

3Loss of information

If the entire spoken utterance is captured before processing, then complete audio data is available for dereverberation, but latency in generating responses increases

Engineering Contradiction:
Improveaudio data completenessVSAvoidresponse time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary updates to the RIR model using audio data as it arrives, so that dereverberation processing can begin before the entire utterance is complete, reducing overall response time while maintaining processing quality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically processes audio data in real-time as it arrives, continuously updating the RIR model and performing dereverberation on available data, allowing response generation to begin before the utterance is fully captured

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11699453B2Adaptive multichannel dereverberation for automatic speech recognition
Publication Date: 2023.07.11 GOOGLE LLC
  • US11699453B2 patent drawing
  • US11699453B2 patent drawing
  • US11699453B2 patent drawing

AI summary

Utilizing an adaptive multichannel technique to mitigate reverberation present in received audio signals, prior to providing corresponding audio data to one or more additional component(s), such as automatic speech recognition (ASR) components. Implementations disclosed herein are “adaptive”, in that they utilize a filter, in the reverberation mitigation, that is online, causal and varies depending on characteristics of the input. Implementations disclosed herein are “multichannel”, in that a corresponding audio signal is received from each of multiple audio transducers (also referred to herein as “microphones”) of a client device, and the multiple audio signals (e.g., frequency domain representations thereof) are utilized in updating of the filter—and dereverberation occurs for audio data corresponding to each of the audio signals (e.g., frequency domain representations thereof) prior to the audio data being provided to ASR component(s) and/or other component(s).