Binaural HRTF High-Band Tuning to Reduce VR Audio Crosstalk
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtual reality audio systems using binaural playback with multi-channel headsets suffer from crosstalk between left and right channel signals due to the convolution of audio signals with head-related transfer functions (HRTFs), which affects the realism of the auditory experience.
Innovation Solution
Modify the high-band impulse responses of the head-related transfer functions (HRTFs) to reduce interference between left and right channel signals by applying specific modification factors to the HRTFs corresponding to virtual speakers positioned on different sides relative to the listener's ears, ensuring energy magnitude consistency between target audio signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If binaural playback is implemented using multi-channel headsets with HRTF convolution, then three-dimensional audio effect is achieved, but crosstalk occurs between left and right channel signals
Solution Approach 1:
The patent divides the M first HRTFs into a first group of a first HRTFs and a second group of c first HRTFs, and the M second HRTFs into a first group of b second HRTFs and a second group of d second HRTFs. This segmentation allows selective modification of high-band impulse responses in specific groups to reduce crosstalk while preserving spatial audio effects.
Solution Approach 2:
The patent modifies only the high-band impulse responses of specific HRTFs (a first HRTFs and b second HRTFs) rather than all HRTFs uniformly. This local modification approach reduces crosstalk in critical frequency regions while maintaining the overall three-dimensional audio rendering quality.
2Reliability
If HRTF convolution is applied to all virtual speakers, then complete spatial coverage is achieved, but interference between left and right channel signals increases
Solution Approach 1:
The patent segments the HRTFs into different groups (a first HRTFs, b second HRTFs, c first HRTFs, d second HRTFs) based on their contribution to crosstalk. By selectively modifying only certain groups, the system maintains comprehensive spatial coverage while minimizing interference from problematic HRTFs.
Solution Approach 2:
The patent modifies the high-band impulse response parameter of specific HRTFs to reduce crosstalk. This parameter change is applied selectively to a first HRTFs and b second HRTFs, altering their frequency characteristics to minimize interference while preserving the spatial rendering function.
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
Embodiments of this application provide an audio processing method and apparatus. The method includes: obtaining M audio signals by processing a to-be-processed audio signal by M virtual speakers; obtaining M first HRTFs and M second HRTFs, where the M first HRTFs are HRTFs to which the M audio signals correspond from the M virtual speakers to a left ear position, and the M second HRTFs are HRTFs to which the M audio signals correspond from the M virtual speakers to a right ear position; modifying high-band impulse responses of a first HRTFs, to obtain a first target HRTFs, and modifying high-band impulse responses of b second HRTFs, to obtain b second target HRTFs; and obtaining, based on the a first target HRTFs, c first HRTFs, and the M audio signals, a first target audio signal corresponding to the left ear position, and obtaining, based on d second HRTFs, the b second target HRTFs, and the M audio signals, a second target audio signal corresponding to the right ear position. a + c = M, and b + d = M. In the embodiments of this application, crosstalk between the first target audio signal and the second target audio signal is reduced.