Binaural Audio Signal Generation Using Frequency-Domain HRTF Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating binaural audio signals from multi-channel audio are complex, resource-intensive, and struggle with representing long echoic or reverberant environments, often resulting in reduced audio quality and increased computational requirements.
Innovation Solution
A system that converts spatial parameters into binaural parameters using binaural perceptual transfer functions, allowing for the generation of binaural audio signals through frequency and time processing, with low complexity and reduced resource demands, capable of handling long impulse responses and echoic environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional binaural synthesis algorithms using HRTFs are used to generate binaural audio signals from multi-channel audio, then spatial sound information can be provided to the user, but the computational complexity and resource requirements increase significantly
Solution Approach 1:
The patent transforms the binaural synthesis process from time-domain convolution with full HRTF impulse responses to frequency-domain processing using HRTF frequency responses at specific frequencies. This parameter transformation reduces computational complexity while maintaining spatial sound information quality, as the invention uses HRTF frequency responses rather than complete time-domain impulse responses.
2Reliability
If conventional binaural synthesis methods are used, then binaural audio signals can be generated, but the processing of long echoic or reverberant environments becomes difficult and audio quality deteriorates
Solution Approach 1:
The patent moves the processing from time-domain convolution to frequency-domain multiplication. By working in the frequency domain, the invention can effectively handle long impulse responses and echoic environments without the computational burden and quality degradation associated with time-domain processing of extended reverberation.
3Reliability
If spatial parameters are converted into binaural parameters using binaural perceptual transfer functions with frequency and time processing, then high-quality binaural audio signals can be generated, but the processing complexity increases
Solution Approach 1:
The patent converts spatial parameters to binaural parameters by multiplying HRTF frequency responses with the power spectral density of the downmix signal at specific frequencies, then transforming back to time domain. This parameter transformation approach maintains high audio quality while reducing overall processing complexity compared to full time-domain convolution.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An apparatus for generating a binaural audio signal comprises a demultiplexer (401) and decoder (403) which receives audio data comprising an audio M-channel audio signal which is a downmix of an N-channel audio signal and spatial parameter data for upmixing the M-channel audio signal to the N-channel audio signal. A conversion processor (411) converts spatial parameters of the spatial parameter data into first binaural parameters in response to at least one binaural perceptual transfer function. A matrix processor (409) converts the M-channel audio signal into a first stereo signal in response to the first binaural parameters. A stereo filter (415, 417) generates the binaural audio signal by filtering the first stereo signal. The filter coefficients for the stereo filter are determined in response to the at least one binaural perceptual transfer function by a coefficient processor (419). The combination of parameter conversion/ processing and filtering allows a high quality binaural signal to be generated with low complexity.