Time-Frequency Weight Masks for Far-Field Sound Enhancement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Far-field sound capture in hands-free devices suffers from reverberation and ambient noise, degrading voice intelligibility and requiring high computational resources for effective signal enhancement.
Innovation Solution
A method that estimates a weight mask in the time-frequency domain using the direction of arrival of sound sources, applying spatial filtering to enhance the desired signal without neural networks, utilizing techniques like Delay and Sum, MPDR, MVDR, and Multichannel Wiener filters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If deep neural networks are used for source separation and enhancement, then signal enhancement performance is improved, but computational cost and memory requirements increase significantly
Solution Approach 1:
The patent extracts and uses only the essential spatial information (direction of arrival) from the complex neural network approach, separating this key element from the computationally expensive full neural network processing. This allows achieving enhancement without the heavy computational burden of complete deep learning models.
Solution Approach 2:
The patent changes the processing parameters from complex neural network operations to simpler spatial filtering operations based on direction of arrival estimates. This parameter change maintains enhancement effectiveness while dramatically reducing computational requirements and memory usage.
2Reliability
If complex spatial filtering techniques (MVDR, Multichannel Wiener filters) are used to reduce noise and reverberation, then voice intelligibility is improved, but knowledge of spatial distribution requirements increases system complexity
Solution Approach 1:
The patent extracts only the direction of arrival information needed for spatial filtering, discarding the requirement for complete spatial distribution maps. This extraction approach enables effective filtering while avoiding the complexity of full spatial distribution estimation.
Solution Approach 2:
The patent applies partial spatial filtering by using only the direction of arrival component rather than complete spatial distribution information. This partial action is sufficient for achieving noise and reverberation reduction without requiring the excessive information needed by traditional MVDR and Multichannel Wiener filters.
Data Source
AI summary
A method and apparatus for processing sound data acquired by a plurality of microphones. The method includes: on the basis of the signals acquired by the plurality of microphones, determining a direction of arrival of a sound originating from at least one sound source of interest; applying spatial filtering to the sound data as a function of the direction of arrival of the sound; estimating ratios, in the time-frequency domain, in a magnitude representative of a signal amplitude, between the filtered sound data on the one hand and the acquired sound data on the other hand; and as a function of the estimated ratios, producing a weight mask to be applied in the time-frequency domain to the acquired sound data in order to construct an acoustic signal representing the sound originating from the source of interest but enhanced relative to the ambient noise.


