Binaural HRTF Filter Selection for Dynamic Audio Image Rendering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Binaural decoders with limited sets of Head-Related Transfer Function (HRTF) filters struggle to accurately render dynamic audio images, leading to inconsistencies between intended and perceived audio effects due to incompatible filter sets and increased bitrate from abrupt sound source movements.
Innovation Solution
A method that inputs parametrically encoded audio signals with channel configuration information to select the closest HRTF filter pairs in stepwise motion, maintaining constant angular velocity and handling singular positions, allowing for dynamic binaural control even with limited filters and minimizing bitrate.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a limited set of HRTF filters is used in the binaural decoder, then the device complexity is reduced, but the audio image rendering accuracy deteriorates
Solution Approach 1:
The patent applies dynamics by enabling the binaural decoder to dynamically select and switch between different HRTF filter pairs based on the audio source location. Instead of using a static limited set, the system adapts the filter selection in real-time according to the spatial position of sound sources, thereby maintaining rendering accuracy while keeping the overall filter database size manageable.
Solution Approach 2:
The patent changes the parameter of HRTF filter selection by introducing a systematic method to choose appropriate filters based on audio source location parameters. The decoder determines the spatial position of sound sources and selects HRTF filter pairs that match these parameters, effectively adapting the filter characteristics to the spatial context without requiring an exhaustive filter set for every possible position.
2Speed
If abrupt sound source movements are handled without stepwise motion, then the processing speed is increased, but the bitrate increases due to incompatible filter set transitions
Solution Approach 1:
The patent applies preliminary action by pre-organizing HRTF filters into structured sets corresponding to different spatial regions or azimuth angles. Before processing audio source movements, the system has the filters ready in organized groups, allowing it to predictably transition between adjacent filter sets when sound sources move, rather than making random or abrupt selections that would require higher bitrate to encode the transitions.
Solution Approach 2:
The patent segments the HRTF filter set into multiple subsets, each associated with specific spatial regions or direction ranges. This segmentation allows the decoder to handle sound source movements by transitioning between adjacent segments in a systematic manner, reducing the information required to encode filter transitions compared to using a monolithic filter set.
3Adaptability or versatility
If a sufficient number of HRTF filter pairs are provided for each loudspeaker position, then the audio image rendering flexibility is improved, but the device complexity increases
Solution Approach 1:
The patent applies universality by designing a HRTF filter selection mechanism that can handle multiple loudspeaker configurations and audio source positions using a unified approach. The same filter selection logic and spatial analysis methods work across different scenarios (different numbers of loudspeakers, different layouts, binaural rendering), eliminating the need for separate specialized filter sets for each configuration while maintaining flexibility.
Solution Approach 2:
The patent uses copying by creating virtual loudspeaker positions through HRTF filtering rather than requiring physical loudspeakers at every possible position. The system copies and adapts HRTF filter characteristics to simulate sound sources at various locations, providing flexible audio image rendering without the hardware complexity of having actual loudspeakers at each position.
Data Source
AI summary
Inputting of a parametrically encoded audio signal comprising at least one combined signal of a plurality of audio channels and one or more corresponding sets of side information describing a multi-channel sound image and including channel configuration information is shown along with deriving, from the channel configuration information, audio source location data describing at least one of horizontal and vertical positions of audio sources in the binaural audio signal; selecting, from a predetermined set of head-related transfer function filters, a left-right pair of head-related transfer function filters matching closest to the audio source location data, wherein the left-right pair of head-related transfer function filters is searched in a stepwise motion in a horizontal plane; and synthesizing a binaural audio signal from the at least one processed signal according to side information and the channel configuration information.


