Microphone System DOA Estimation Using RTF Dictionary

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current microphone systems in hearing devices face challenges in accurately estimating the direction-of-arrival (DOA) of a target sound source in noisy environments, particularly due to the complexity of acoustic transfer functions and noise interference.

Innovation Solution

A microphone system employing a maximum likelihood (ML) method using a dictionary of relative transfer functions (RTFs) to estimate the DOA, where the signal processor determines the most likely RTF from the dictionary based on observed noisy signals, and utilizes this for beamforming and signal-to-noise ratio (SNR) estimation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a maximum likelihood method with a complete dictionary of acoustic transfer functions is used, then measurement precision of DOA estimation is improved, but device complexity increases due to computational requirements

Engineering Contradiction:
ImproveDOA estimation accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The acoustic environment is segmented into discrete spatial locations forming a grid, with transfer functions pre-calculated and stored in a dictionary for each location. This segmentation transforms the continuous acoustic space into discrete segments, allowing the system to evaluate only predefined transfer function candidates rather than computing all possible transfer functions in real-time, thus reducing computational complexity while maintaining estimation accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Transfer functions for all possible source locations are pre-calculated and stored in a dictionary before actual DOA estimation occurs. This preliminary action eliminates the need for real-time transfer function computation during signal processing, significantly reducing device complexity while preserving the ability to achieve high measurement precision through maximum likelihood estimation.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If a large dictionary of relative transfer functions is used to cover all possible directions, then adaptability of DOA estimation is improved, but loss of time increases due to evaluating more dictionary elements

Engineering Contradiction:
Improvecoverage of possible source directionsVSAvoidcomputation time for likelihood evaluation
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The spatial domain is segmented into a discrete grid of possible source locations, with each grid point having a pre-stored transfer function in the dictionary. This segmentation allows the system to maintain comprehensive directional coverage while evaluating only a finite, manageable number of discrete locations rather than continuously scanning all possible directions, thus reducing computation time.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of computing transfer functions for all possible directions during operation, the system creates a copy of commonly encountered transfer function patterns in a pre-computed dictionary. This copying approach allows rapid lookup and comparison during DOA estimation, maintaining adaptability across all directions while significantly reducing the time required for evaluating dictionary elements during actual signal processing.

Inventive Principle:
Principle #26Copying

3Ease of operation

If traditional beamforming methods are used without a dictionary constraint, then ease of operation is maintained, but measurement precision deteriorates in noisy environments due to own voice interference

Engineering Contradiction:
Improvesimplicity of implementationVSAvoidDOA estimation accuracy in noise
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The transfer function dictionary serves as an intermediary between the raw microphone signals and the DOA estimation process. By comparing observed signals against pre-stored transfer function candidates from the dictionary, the system can identify the most likely source direction even in noisy environments. This intermediary structure maintains ease of operation through straightforward dictionary lookup while dramatically improving measurement precision by eliminating own-voice interference through constrained transfer function matching.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP4184950A1A microphone system and a hearing device comprising a microphone system
Publication Date: 2023.05.24 OTICON
  • EP4184950A1 patent drawingFigure 1A~1C
  • EP4184950A1 patent drawingFigure 2A~2G
  • EP4184950A1 patent drawingFigure 3A~3C

AI summary

A hearing system comprising a hearing device comprising a multitude of microphones; a signal processor connected to said number of microphones, and being configured to estimate a direction- to and/or a position of the target sound source relative to the microphone system based on a maximum likelihood methodology; and a database Θ comprising a dictionary of relative transfer functions (RTF) in the form of an RTF-vector representing direction-dependent acoustic transfer functions from said target signal source to each of said microphones relative to a reference microphone among said microphones, wherein individual dictionary elements of said database Θ of relative transfer functions comprises relative transfer functions for a number of different directions and/or positions relative to the microphone system; and wherein the signal processor is configured to determine one or more of the most likely directions to or locations of said target sound source, and to provide own voice detection in that the database Θ of relative transfer functions comprises an RTF vector corresponding to the position of the mouth of the user, and wherein an indication that user's own voice is present is provided when the most likely look vector in the dictionary at a given point in time is the one that corresponds to the location of the user's mouth. The invention may e.g. be used for the hearing aids or other portable audio communication devices.