Direction-Based Speech Masking for Low-Latency Noise Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Distant sound recording introduces artifacts such as reverberation and ambient noise, degrading the intelligibility of the speaker's voice and complicating communication with speech recognition engines, despite the use of microphone antennas for signal enhancement.
Innovation Solution
A method that estimates a weighting mask in the time-frequency domain using the direction of arrival of the sound source, applying spatial filtering to enhance the desired signal without relying on neural networks, utilizing techniques like Delay and Sum, MPDR, and MWF filters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If deep neural networks are used for source separation and mask estimation, then speech enhancement performance is improved, but computational cost and memory requirements increase significantly
Solution Approach 1:
The patent replaces expensive deep neural networks with computationally inexpensive classical signal processing methods (spatial filtering, time-frequency masking) that can be executed with minimal computational resources, achieving acceptable enhancement performance without the high cost of AI models
Solution Approach 2:
The patent substitutes the 'mechanical system' of deep neural networks with a combination of spatial filtering techniques and time-frequency domain processing, replacing complex computational machinery with more efficient mathematical operations
2Measurement precision
If deep neural networks with many layers and parameters are deployed, then separation accuracy is improved, but memory consumption and training requirements increase
Solution Approach 1:
The patent uses lightweight computational methods that consume minimal memory, replacing large-scale neural network models with efficient classical algorithms that achieve separation accuracy without requiring substantial memory resources
3Object-affected harmful factors
If spatial filtering techniques are applied using direction of arrival information, then noise reduction is improved, but the system requires accurate spatial distribution knowledge which is difficult to obtain
Solution Approach 1:
The patent performs preliminary spatial filtering using direction of arrival information to separate speech from noise in the time-frequency domain, creating an initial estimate that is then refined through iterative processing, thereby obtaining spatial distribution knowledge progressively rather than requiring it a priori
Solution Approach 2:
The patent employs iterative refinement where the output of spatial filtering feeds back into mask estimation, and the estimated mask is used to improve subsequent spatial filtering operations, creating a closed-loop system that progressively improves noise reduction without requiring perfect initial spatial knowledge
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present description relates to processing sound data acquired by a plurality of microphones (MIC), in which: - on the basis of the signals acquired by the plurality of microphones, determining a direction of arrival of a sound originating from at least one sound source of interest (S4); - applying spatial filtering to the sound data as a function of the direction of arrival of the sound (S5); - estimating in the time-frequency domain ratios of a quantity representative of a signal amplitude, between the filtered sound data on the one hand and the acquired sound data on the other hand (S6); - as a function of the estimated ratios, producing a weight mask to be applied in the time-frequency domain to the acquired sound data (S7) with a view to constructing a sound signal representing the sound originating from the source of interest but boosted with respect to ambient noise (S10; S9- S10).