Audio Source Azimuth Positioning With Echo-Canceled Time-Frequency Weighting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio interaction devices face challenges in accurately determining the azimuth of a target voice due to sensitivity to the accuracy of the specified azimuth, which affects the quality and performance of voice interaction.

Innovation Solution

An audio recognition method that involves obtaining audio signals in multiple directions, performing echo cancellation, calculating weights of time-frequency points, and using a weighted covariance matrix to enhance the accuracy of sound source azimuth estimation, thereby reducing interference from noise and echoes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If beamforming algorithm is used to enhance target voice, then voice signal quality is improved, but azimuth accuracy requirement becomes more stringent

Engineering Contradiction:
Improvevoice signal qualityVSAvoidazimuth accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by performing echo cancellation before azimuth estimation. The echo-canceled audio signals are processed to obtain time-frequency domain expressions and calculate weights before the azimuth is determined. This preliminary processing removes echo interference that would otherwise distort the azimuth estimation, allowing the beamforming algorithm to operate with more accurate azimuth values without requiring extremely precise measurements.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary approach by using weighted covariance matrices as a mediator between the raw audio signals and the final azimuth estimation. The weights, calculated from time-frequency domain expressions of echo-canceled signals, serve as a mediator that adjusts the contribution of different signal components. This intermediary mechanism allows the system to achieve accurate azimuth estimation even in complex environments with echo and noise, resolving the contradiction between voice quality enhancement and azimuth measurement precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If azimuth estimation is performed in complex environments with echo and noise, then audio interaction capability is maintained, but azimuth accuracy deteriorates

Engineering Contradiction:
Improveaudio interaction capabilityVSAvoidazimuth accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by performing echo cancellation before azimuth estimation. The echo-canceled audio signals are processed to obtain time-frequency domain expressions and calculate weights before the azimuth is determined. This preliminary processing removes echo interference that would otherwise distort the azimuth estimation, allowing the beamforming algorithm to operate with more accurate azimuth values without requiring extremely precise measurements.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies parameter changes by transforming the audio signals from the time domain to the time-frequency domain using Short-Time Fourier Transform (STFT). This transformation changes the representation parameters of the signals, allowing for more accurate extraction of time-frequency domain expressions that can be used to calculate weights. These parameter changes enable the system to distinguish between echo, noise, and target voice more effectively, maintaining azimuth accuracy in complex environments.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12361935B2Audio recognition method, method, apparatus for positioning target audio, and device
Publication Date: 2025.07.15 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US12361935B2 patent drawing
  • US12361935B2 patent drawing
  • US12361935B2 patent drawing

AI summary

This application discloses a method for positioning a target audio signal by a computer device. The method includes: performing echo cancellation on the audio signals collected in a plurality of directions in a space, the audio signals comprising a target-audio direct signal; obtaining weights of a plurality of time-frequency points in the echo-canceled audio signals, a weight of each time-frequency point indicating a relative proportion of the target-audio direct signal in the echo-canceled audio signals at the time-frequency point; obtaining a weighted audio signal energy distribution of the audio signals in the plurality of directions by using the weights of the plurality of time-frequency points in the echo-canceled audio signals; and obtaining a sound source azimuth corresponding to the target-audio direct signal in the audio signals by using the weighted audio signal energy distribution of the audio signals in the plurality of directions.