Parametric Audio Renderer for Spatialization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio systems require significant computational resources, power, and memory to perform time-domain convolution of head-related transfer functions (HRTFs) for spatialized audio, making them unsuitable for resource-constrained devices like headsets.
Innovation Solution
A parametric audio system that uses a set of infinite impulse response (IIR) filters and dynamic filters to approximate HRTFs, allowing for efficient generation of spatialized audio content by selecting and configuring an audio time and level difference renderer (TLDR) based on input parameters, reducing computational complexity and memory footprint.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If time-domain convolution of HRTFs is used for spatialized audio generation, then audio quality and spatial accuracy are improved, but computational resources, power consumption, and memory requirements increase significantly
Solution Approach 1:
The patent transforms the HRTF processing from time-domain convolution to frequency-domain multiplication by changing the domain parameter. This parameter change reduces computational complexity from O(N*M) for convolution to O(N) for multiplication, significantly lowering power consumption while maintaining spatial audio accuracy through the use of parametric models that preserve critical spatial cues.
Solution Approach 2:
The patent replaces the computationally intensive mechanical process of time-domain convolution with a more efficient frequency-domain multiplication approach. This substitution uses the Fourier transform domain to achieve the same spatialization effect with fraction of the computational resources and power consumption.
2Measurement precision
If time-domain convolution of HRTFs is used for spatialized audio generation, then audio quality and spatial accuracy are improved, but device complexity and memory footprint increase
Solution Approach 1:
The patent changes the mathematical domain from time to frequency, transforming the convolution operation into multiplication. This parameter change simplifies the computational complexity from requiring multiple multiply-accumulate operations per sample to a single multiplication per frequency bin, making the system suitable for resource-constrained devices.
Solution Approach 2:
The patent segments the HRTF processing into frequency-band specific operations, where different parametric models are applied to different frequency ranges. This segmentation allows for optimized computation in each band while maintaining overall spatial accuracy, reducing total computational complexity.
3Measurement precision
If time-domain convolution of HRTFs is used for spatialized audio generation, then audio quality and spatial accuracy are improved, but memory requirements increase
Solution Approach 1:
The patent changes from storing complete impulse response functions in time domain to storing compact parametric models in frequency domain. This parameter change reduces memory footprint by representing HRTFs with a small set of parameters (gains, frequencies, Q-factors) rather than storing entire convolution kernels, while maintaining spatial accuracy through parametric synthesis.
Solution Approach 2:
Instead of storing and processing complete HRTF impulse responses, the patent inverts the approach by storing compact parametric representations that can be synthesized on-demand. This inversion reduces memory requirements from storing large convolution kernels to storing small parameter sets that generate the necessary spatial cues.
Data Source
AI summary
A system is disclosed for using an audio time and level difference renderer (TLDR) to generate spatialized audio content for multiple channels from an audio signal received at a single channel. The system selects an audio TLDR from a set of audio TLDRs based on received input parameters. The system configures the selected audio TLDR based on received input parameters using a filter parameter model to generate a configured audio TLDR that comprises a set of configured binaural dynamic filters, and a configured delay between the multiple channel. The system applies the configured audio TLDR to an audio signal received at the single channel to generate spatialized multiple channel audio content for each channel of the multiple audio channel and presents the generated spatialized audio content at multiple channels to a user via a headset.


