Deep Neural Network Sound Field Analysis for Direction of Arrival Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional digital signal processing algorithms for sound field analysis face limitations in resolving direction of arrival (DOA) above spatial aliasing frequencies and at low frequencies due to acoustic noise and low spatial resolution, especially in multichannel audio applications like speech enhancement and spatial sound reproduction.

Innovation Solution

A deep neural network (DNN) is trained using sub-band directional features such as Steered-Response Power Phase Transform (SRP-PHAT), inter-microphone phase differences, and diffuseness to accurately estimate DOA, addressing the aliasing issues and improving sound field analysis in various acoustic environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional DSP methods are used for DOA estimation, then the system is simple to implement, but DOA cannot be resolved above spatial aliasing frequencies and at low frequencies

Engineering Contradiction:
ImproveDOA estimation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces conventional DSP algorithms with a deep neural network (DNN) based approach. The DNN is trained offline using simulated and recorded data, then deployed to estimate DOA across the full frequency range including regions where traditional methods fail (above spatial aliasing frequencies and at low frequencies). This substitution enables full-band DOA estimation while maintaining computational efficiency during operation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If traditional multi-source localization is used, then the algorithm is computationally efficient, but it does not perform consistently well for arbitrary microphone arrays

Engineering Contradiction:
Improvelocalization consistencyVSAvoidmicrophone array adaptability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent performs preliminary training of the DNN offline using extensive simulated data and real recorded data from various microphone arrays. This pre-training enables the model to learn robust features and adaptation strategies beforehand, allowing it to consistently perform well across different microphone array geometries and acoustic environments without requiring real-time adaptation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms the input acoustic signals into time-frequency domain representations and extracts specific features (inter-microphone phase differences, spectral characteristics) as inputs to the DNN. These parameter transformations enable the system to adapt to different microphone array configurations and acoustic conditions while maintaining consistent localization performance.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If conventional techniques are used for sound field analysis, then the processing is straightforward, but DOA cannot be resolved at low frequencies due to acoustic noise and low spatial resolution

Engineering Contradiction:
Improvelow frequency DOA resolutionVSAvoidacoustic noise impact
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent replaces conventional low-frequency DOA estimation techniques with a DNN-based approach that can effectively operate in the low-frequency region. The DNN learns to extract meaningful spatial information from acoustic signals even when traditional spatial resolution is insufficient, thereby overcoming the limitations imposed by acoustic noise and low spatial resolution at low frequencies.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10334357B2Machine learning based sound field analysis
Publication Date: 2019.06.25 APPLE INC
  • US10334357B2 patent drawing
  • US10334357B2 patent drawing
  • US10334357B2 patent drawing

AI summary

Impulse responses of a device are measured. A database of sound files is generated by convolving source signals with the impulse responses of the device. The sound files from the database are transformed into time-frequency domain. One or more sub-band directional features is estimated at each sub-band of the time-frequency domain. A deep neural network (DNN) is trained for each sub-band based on the estimated one or more sub-band directional features and a target directional feature.