Spectral-Spatial Mask Estimation for Multi-Channel Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional spectral-spatial methods for multi-channel speech enhancement in automatic speech recognition (ASR) are not extendable to other scenarios and rely on mask pooling, which can lead to inaccuracies in multi-channel mask estimation, especially in low signal-to-noise ratio (SNR) and high reverberation environments.

Innovation Solution

A spectral-spatial mask estimation method that concatenates channel features and cross-channel features without mask pooling, using a neural network-based approach to estimate spectral-spatial masks for beamforming in multiple channels, thereby integrating both channel and cross-channel information for improved ASR performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If mask pooling operations are used in conventional spectral-spatial methods, then the method is simpler to implement, but the multi-channel mask estimation accuracy deteriorates, especially in low-SNR and high reverberation environments

Engineering Contradiction:
Improveease of implementationVSAvoidmask estimation accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent extracts and removes the mask pooling operation from the conventional spectral-spatial method pipeline. By taking out the pooling step that causes accuracy degradation, the system directly feeds channel features and cross-channel features into the mask estimation network, preserving fine-grained spatial information and achieving superior mask estimation accuracy without the simplification of pooling operations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces a new dimensional approach by explicitly modeling cross-channel relationships as a separate feature dimension. Instead of pooling across channels, the system concatenates channel features with cross-channel features (such as inter-channel coherence and phase difference) to create an enhanced feature representation that captures spatial relationships without losing channel-specific information.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If conventional spectral-spatial methods are used, then the current ASR performance is achieved, but the method is not extendable to other multi-channel scenarios

Engineering Contradiction:
ImproveASR performanceVSAvoidextendability to other scenarios
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal mask estimation framework that can handle diverse multi-channel scenarios. By formulating the feature extraction and mask estimation in a generalizable manner that works with different microphone array configurations and acoustic environments, the system achieves both reliable ASR performance in current scenarios and adaptability to future scenarios without requiring method-specific modifications.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If channel features and cross-channel features are concatenated without mask pooling, then the robustness and accuracy of ASR is improved, but the computational complexity increases

Engineering Contradiction:
Improverobustness of ASRVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary feature extraction and organization before the mask estimation stage. By pre-computing channel features and cross-channel features in a structured manner and organizing them for efficient concatenation, the system reduces the computational burden during the main mask estimation process, making the more accurate approach computationally feasible.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11289109B2Systems and methods for audio signal processing using spectral-spatial mask estimation
Publication Date: 2022.03.29 BEIJING DIDI INFINITY TECH & DEV CO LTD
  • US11289109B2 patent drawing
  • US11289109B2 patent drawing
  • US11289109B2 patent drawing

AI summary

Embodiments of the disclosure provide systems and methods for audio signal processing. An exemplary system may include a communication interface configured to receiving a first audio signal acquired from an audio source through a first channel, and a second audio signal acquired from the same audio source through a second channel. The system may also include at least one processor coupled to the communication interface. The at least one processor may be configured to determine channel features based on the first audio signal and the second audio signal individually and determine a cross-channel feature based on the first audio signal and the second audio signal collectively. The at least one processor may further be configured to concatenate the channel features and the cross-channel feature and estimate spectral-spatial masks for the first channel and the second channel using the concatenated channel features and the cross-channel feature. The at least one processor may also be configured to perform beamforming based on the spectral-spatial masks for the first channel and the second channel.