DOA Estimation Using Energy Difference Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing DOA estimators fail to accurately determine the direction of arrival in highly reverberant environments with spatially-coherent noise sources, such as TVs and radios, due to interference issues.

Innovation Solution

A method that involves buffering audio samples, detecting a trigger point using known data, separating the samples into noise and signal-plus-noise segments, and computing energy differences across multiple directions to select the direction with the largest energy difference as the DOA, effectively filtering out spatially-coherent noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a histogram-based beamformer is used for DOA estimation, then computational efficiency is improved, but accuracy deteriorates in the presence of spatially-coherent noise

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidDOA estimation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The audio signal is segmented into multiple frequency bins using Fourier transform, allowing the algorithm to process different frequency components separately. This segmentation enables the system to identify and exclude frequency bins contaminated by spatially-coherent noise while retaining information from clean frequency bins, thus maintaining DOA estimation accuracy without sacrificing computational efficiency

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The algorithm dynamically changes parameters based on noise detection: it adjusts the weight or inclusion of different frequency bins based on their coherence properties. By changing the parameter of frequency bin selection and weighting according to noise characteristics, the system adapts to spatially-coherent interference while preserving computational efficiency through selective processing

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If conventional DOA estimators are used in highly reverberant environments, then they can handle diffuse noise, but they fail when spatially-coherent noise sources are present

Engineering Contradiction:
Improveability to handle diffuse noiseVSAvoidDOA estimation reliability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The algorithm introduces an intermediary coherence check mechanism that acts as a mediator between the raw audio signal and the DOA estimation process. This intermediary step analyzes the spatial coherence of different frequency bins and uses this information to selectively process or exclude bins, thereby protecting the final DOA estimate from spatially-coherent noise while maintaining the ability to handle diffuse noise

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The algorithm extracts and removes the harmful component by identifying frequency bins with high spatial coherence that indicate the presence of interfering sources like TVs or radios. By taking out these contaminated frequency bins from the processing stream and excluding them from the final DOA calculation, the system eliminates their detrimental effect while preserving the reliability of DOA estimation

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10811032B2Data aided method for robust direction of arrival (DOA) estimation in the presence of spatially-coherent noise interferers
Publication Date: 2020.10.20 CIRRUS LOGIC INC
  • US10811032B2 patent drawing
  • US10811032B2 patent drawing
  • US10811032B2 patent drawing

AI summary

A method and apparatus to determine a direction of arrival (DOA) of a talker in the presence of a source of spatially-coherent noise. A time sequence of audio samples that include the spatially-coherent noise is received and buffered. Aided by previously known data, a trigger point is detected in the time sequence of audio samples when the talker begins to talk. The buffered time sequence of audio samples is separated into a noise segment and a signal-plus-noise segment based on the trigger point. For each direction of a plurality of distinct directions: an energy difference is computed for the direction between the noise segment and the signal-plus-noise segment, and the DOA of the talker is selected as the direction of the plurality of distinct directions having a largest of the computed energy differences.