Headset Sound Source Localization Using Frequency Band Discrimination

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Head-mounted terminal devices experience front/rear sound image confusion due to differences in earpiece playback, leading to incorrect determination of sound source direction, with a higher probability of mistaking a front sound image for a rear one.

Innovation Solution

A method involving at least three microphones positioned differently to receive sound signals, determining signal delay differences to determine the sound source's position and performing orientation enhancement processing to differentiate between front and rear characteristic frequency bands, thereby reducing confusion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If binaural sound signals are played back through earpieces of a head-mounted terminal device, then spatial information consistent with the original sound field is generated, but front/rear sound image confusion occurs due to loss of cognition information

Engineering Contradiction:
Improvesound source direction determination accuracyVSAvoidfront/rear orientation discrimination reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent introduces a new dimension of frequency band analysis to resolve the front/rear confusion problem. By dividing the sound signal into front characteristic frequency bands and rear characteristic frequency bands, the system creates an additional discriminative dimension beyond traditional spatial cues. This frequency-based dimensionality allows the system to distinguish front from rear sound images that would otherwise be confused in the spatial domain alone.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent applies local quality by selectively enhancing specific frequency bands rather than processing the entire sound signal uniformly. The front characteristic frequency bands are enhanced while rear characteristic frequency bands are suppressed, creating localized spectral modifications that correspond to directional information. This selective processing improves front/rear discrimination without affecting other aspects of the sound quality.

Inventive Principle:
Principle #3Local quality

2Device complexity

If only interaural time difference and interaural level difference are used for sound source localization, then a cone of confusion is determined, but the specific direction (front or rear) cannot be determined

Engineering Contradiction:
Improvelocalization system complexityVSAvoidsound source direction precision
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent introduces frequency band energy ratios as an intermediary parameter that mediates between the simple ITD/ILD measurements and the complex task of front/rear direction determination. By calculating the ratio of front characteristic frequency band energy to rear characteristic frequency band energy, the system creates an intermediate metric that bridges the gap between basic spatial cues and precise directional localization, enabling the system to resolve the cone of confusion without overly complicating the overall architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3249948B1Method and terminal device for processing voice signal
Publication Date: 2020.10.07 HUAWEI TECH CO LTD
  • EP3249948B1 patent drawingFigure 1
  • EP3249948B1 patent drawingFigure 2~3
  • EP3249948B1 patent drawingFigure 4

AI summary

Embodiments of the present invention provide a method for processing a sound signal and a terminal device. The method includes: receiving, by using channels located in different positions of a terminal device, at least three signals sent by a same sound source, where the at least three signals are in a one-to-one correspondence to the channels; determining, according to three signals in the at least three signals, a signal delay difference between every two of the three signals, where a position of the sound source relative to the terminal device can be determined according to the signal delay difference; determining, according to the signal delay difference, the position of the sound source relative to the terminal device; and when the sound source is located in front of the terminal device, performing orientation enhancement processing on a target signal in the at least three signals, and obtaining a first output signal and a second output signal of the terminal device according to a result of the orientation enhancement processing, where the orientation enhancement processing is used to increase a degree of discrimination between a front characteristic frequency band and a rear characteristic frequency band of the target signal. The embodiments of the present invention can enhance perception of a sound image orientation of an output signal, and reduce a probability of incorrectly determining a front sound image as a rear sound image.