Deep Neural Network Sound Field Analysis for Direction of Arrival Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional digital signal processing algorithms for sound field analysis face limitations in resolving direction of arrival (DOA) above spatial aliasing frequencies and at low frequencies due to acoustic noise and low spatial resolution, especially in multichannel audio applications like speech enhancement and spatial sound reproduction.
Innovation Solution
A deep neural network (DNN) is trained using sub-band directional features such as Steered-Response Power Phase Transform (SRP-PHAT), inter-microphone phase differences, and diffuseness to accurately estimate DOA, addressing the aliasing issues and improving sound field analysis in various acoustic environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional DSP methods are used for DOA estimation, then the system is simple to implement, but DOA cannot be resolved above spatial aliasing frequencies and at low frequencies
Solution Approach 1:
The patent replaces conventional DSP algorithms with a deep neural network (DNN) based approach. The DNN is trained offline using simulated and recorded data, then deployed to estimate DOA across the full frequency range including regions where traditional methods fail (above spatial aliasing frequencies and at low frequencies). This substitution enables full-band DOA estimation while maintaining computational efficiency during operation.
2Reliability
If traditional multi-source localization is used, then the algorithm is computationally efficient, but it does not perform consistently well for arbitrary microphone arrays
Solution Approach 1:
The patent performs preliminary training of the DNN offline using extensive simulated data and real recorded data from various microphone arrays. This pre-training enables the model to learn robust features and adaptation strategies beforehand, allowing it to consistently perform well across different microphone array geometries and acoustic environments without requiring real-time adaptation.
Solution Approach 2:
The patent transforms the input acoustic signals into time-frequency domain representations and extracts specific features (inter-microphone phase differences, spectral characteristics) as inputs to the DNN. These parameter transformations enable the system to adapt to different microphone array configurations and acoustic conditions while maintaining consistent localization performance.
3Measurement precision
If conventional techniques are used for sound field analysis, then the processing is straightforward, but DOA cannot be resolved at low frequencies due to acoustic noise and low spatial resolution
Solution Approach 1:
The patent replaces conventional low-frequency DOA estimation techniques with a DNN-based approach that can effectively operate in the low-frequency region. The DNN learns to extract meaningful spatial information from acoustic signals even when traditional spatial resolution is insufficient, thereby overcoming the limitations imposed by acoustic noise and low spatial resolution at low frequencies.
Data Source
AI summary
Impulse responses of a device are measured. A database of sound files is generated by convolving source signals with the impulse responses of the device. The sound files from the database are transformed into time-frequency domain. One or more sub-band directional features is estimated at each sub-band of the time-frequency domain. A deep neural network (DNN) is trained for each sub-band based on the estimated one or more sub-band directional features and a target directional feature.


