Spectral-Spatial Mask Estimation for Multi-Channel Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional spectral-spatial methods for multi-channel speech enhancement in automatic speech recognition (ASR) are not extendable to other scenarios and rely on mask pooling, which can lead to inaccuracies in multi-channel mask estimation, especially in low signal-to-noise ratio (SNR) and high reverberation environments.
Innovation Solution
A spectral-spatial mask estimation method that concatenates channel features and cross-channel features without mask pooling, using a neural network-based approach to estimate spectral-spatial masks for beamforming in multiple channels, thereby integrating both channel and cross-channel information for improved ASR performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If mask pooling operations are used in conventional spectral-spatial methods, then the method is simpler to implement, but the multi-channel mask estimation accuracy deteriorates, especially in low-SNR and high reverberation environments
Solution Approach 1:
The patent extracts and removes the mask pooling operation from the conventional spectral-spatial method pipeline. By taking out the pooling step that causes accuracy degradation, the system directly feeds channel features and cross-channel features into the mask estimation network, preserving fine-grained spatial information and achieving superior mask estimation accuracy without the simplification of pooling operations.
Solution Approach 2:
The patent introduces a new dimensional approach by explicitly modeling cross-channel relationships as a separate feature dimension. Instead of pooling across channels, the system concatenates channel features with cross-channel features (such as inter-channel coherence and phase difference) to create an enhanced feature representation that captures spatial relationships without losing channel-specific information.
2Reliability
If conventional spectral-spatial methods are used, then the current ASR performance is achieved, but the method is not extendable to other multi-channel scenarios
Solution Approach 1:
The patent creates a universal mask estimation framework that can handle diverse multi-channel scenarios. By formulating the feature extraction and mask estimation in a generalizable manner that works with different microphone array configurations and acoustic environments, the system achieves both reliable ASR performance in current scenarios and adaptability to future scenarios without requiring method-specific modifications.
3Reliability
If channel features and cross-channel features are concatenated without mask pooling, then the robustness and accuracy of ASR is improved, but the computational complexity increases
Solution Approach 1:
The patent performs preliminary feature extraction and organization before the mask estimation stage. By pre-computing channel features and cross-channel features in a structured manner and organizing them for efficient concatenation, the system reduces the computational burden during the main mask estimation process, making the more accurate approach computationally feasible.
Data Source
AI summary
Embodiments of the disclosure provide systems and methods for audio signal processing. An exemplary system may include a communication interface configured to receiving a first audio signal acquired from an audio source through a first channel, and a second audio signal acquired from the same audio source through a second channel. The system may also include at least one processor coupled to the communication interface. The at least one processor may be configured to determine channel features based on the first audio signal and the second audio signal individually and determine a cross-channel feature based on the first audio signal and the second audio signal collectively. The at least one processor may further be configured to concatenate the channel features and the cross-channel feature and estimate spectral-spatial masks for the first channel and the second channel using the concatenated channel features and the cross-channel feature. The at least one processor may also be configured to perform beamforming based on the spectral-spatial masks for the first channel and the second channel.


