Binaural HRTF High-Band Tuning to Reduce VR Audio Crosstalk

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing virtual reality audio systems using binaural playback with multi-channel headsets suffer from crosstalk between left and right channel signals due to the convolution of audio signals with head-related transfer functions (HRTFs), which affects the realism of the auditory experience.

Innovation Solution

Modify the high-band impulse responses of the head-related transfer functions (HRTFs) to reduce interference between left and right channel signals by applying specific modification factors to the HRTFs corresponding to virtual speakers positioned on different sides relative to the listener's ears, ensuring energy magnitude consistency between target audio signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If binaural playback is implemented using multi-channel headsets with HRTF convolution, then three-dimensional audio effect is achieved, but crosstalk occurs between left and right channel signals

Engineering Contradiction:
Improvethree-dimensional audio effectVSAvoidcrosstalk between channels
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent divides the M first HRTFs into a first group of a first HRTFs and a second group of c first HRTFs, and the M second HRTFs into a first group of b second HRTFs and a second group of d second HRTFs. This segmentation allows selective modification of high-band impulse responses in specific groups to reduce crosstalk while preserving spatial audio effects.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent modifies only the high-band impulse responses of specific HRTFs (a first HRTFs and b second HRTFs) rather than all HRTFs uniformly. This local modification approach reduces crosstalk in critical frequency regions while maintaining the overall three-dimensional audio rendering quality.

Inventive Principle:
Principle #3Local quality

2Reliability

If HRTF convolution is applied to all virtual speakers, then complete spatial coverage is achieved, but interference between left and right channel signals increases

Engineering Contradiction:
Improvespatial audio coverageVSAvoidsignal interference
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent segments the HRTFs into different groups (a first HRTFs, b second HRTFs, c first HRTFs, d second HRTFs) based on their contribution to crosstalk. By selectively modifying only certain groups, the system maintains comprehensive spatial coverage while minimizing interference from problematic HRTFs.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent modifies the high-band impulse response parameter of specific HRTFs to reduce crosstalk. This parameter change is applied selectively to a first HRTFs and b second HRTFs, altering their frequency characteristics to minimize interference while preserving the spatial rendering function.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3833056B1Audio processing method and apparatus
Publication Date: 2026.01.28 HUAWEI TECH CO LTD
  • EP3833056B1 patent drawingFigure 1~2
  • EP3833056B1 patent drawingFigure 3
  • EP3833056B1 patent drawingFigure 4

AI summary

Embodiments of this application provide an audio processing method and apparatus. The method includes: obtaining M audio signals by processing a to-be-processed audio signal by M virtual speakers; obtaining M first HRTFs and M second HRTFs, where the M first HRTFs are HRTFs to which the M audio signals correspond from the M virtual speakers to a left ear position, and the M second HRTFs are HRTFs to which the M audio signals correspond from the M virtual speakers to a right ear position; modifying high-band impulse responses of a first HRTFs, to obtain a first target HRTFs, and modifying high-band impulse responses of b second HRTFs, to obtain b second target HRTFs; and obtaining, based on the a first target HRTFs, c first HRTFs, and the M audio signals, a first target audio signal corresponding to the left ear position, and obtaining, based on d second HRTFs, the b second target HRTFs, and the M audio signals, a second target audio signal corresponding to the right ear position. a + c = M, and b + d = M. In the embodiments of this application, crosstalk between the first target audio signal and the second target audio signal is reduced.