Time-Frequency Weight Masks for Far-Field Sound Enhancement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Far-field sound capture in hands-free devices suffers from reverberation and ambient noise, degrading voice intelligibility and requiring high computational resources for effective signal enhancement.

Innovation Solution

A method that estimates a weight mask in the time-frequency domain using the direction of arrival of sound sources, applying spatial filtering to enhance the desired signal without neural networks, utilizing techniques like Delay and Sum, MPDR, MVDR, and Multichannel Wiener filters.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If deep neural networks are used for source separation and enhancement, then signal enhancement performance is improved, but computational cost and memory requirements increase significantly

Engineering Contradiction:
Improvesignal enhancement performanceVSAvoidcomputational cost
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts and uses only the essential spatial information (direction of arrival) from the complex neural network approach, separating this key element from the computationally expensive full neural network processing. This allows achieving enhancement without the heavy computational burden of complete deep learning models.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the processing parameters from complex neural network operations to simpler spatial filtering operations based on direction of arrival estimates. This parameter change maintains enhancement effectiveness while dramatically reducing computational requirements and memory usage.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If complex spatial filtering techniques (MVDR, Multichannel Wiener filters) are used to reduce noise and reverberation, then voice intelligibility is improved, but knowledge of spatial distribution requirements increases system complexity

Engineering Contradiction:
Improvevoice intelligibilityVSAvoidspatial distribution knowledge requirements
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts only the direction of arrival information needed for spatial filtering, discarding the requirement for complete spatial distribution maps. This extraction approach enables effective filtering while avoiding the complexity of full spatial distribution estimation.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial spatial filtering by using only the direction of arrival component rather than complete spatial distribution information. This partial action is sufficient for achieving noise and reverberation reduction without requiring the excessive information needed by traditional MVDR and Multichannel Wiener filters.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12462825B2Estimating an optimized mask for processing acquired sound data
Publication Date: 2025.11.04 ORANGE SA
  • US12462825B2 patent drawing
  • US12462825B2 patent drawing
  • US12462825B2 patent drawing

AI summary

A method and apparatus for processing sound data acquired by a plurality of microphones. The method includes: on the basis of the signals acquired by the plurality of microphones, determining a direction of arrival of a sound originating from at least one sound source of interest; applying spatial filtering to the sound data as a function of the direction of arrival of the sound; estimating ratios, in the time-frequency domain, in a magnitude representative of a signal amplitude, between the filtered sound data on the one hand and the acquired sound data on the other hand; and as a function of the estimated ratios, producing a weight mask to be applied in the time-frequency domain to the acquired sound data in order to construct an acoustic signal representing the sound originating from the source of interest but enhanced relative to the ambient noise.