Audio Signal Processing Apparatus for 3D Binaural Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio signal processing methods for generating binaural audio signals from virtual target positions are inefficient due to the complexity and memory requirements of head-related transfer functions (HRTFs), particularly in accurately simulating elevation perception and requiring large datasets for all azimuth angles.
Innovation Solution
An audio signal processing apparatus that extends predefined two-dimensional HRTFs to three dimensions by using infinite impulse response filters to approximate main spectral features, adjusting delays based on azimuth and elevation angles, and employing cascaded biquad filters to compensate for sound travel time differences, thereby reducing computational complexity and memory needs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If HRTFs are measured for all azimuth angles to achieve accurate binaural sound synthesis, then sound localization accuracy is improved, but measurement complexity and time increase
Solution Approach 1:
The patent pre-calculates and stores transfer functions for a limited set of discrete azimuth angles before actual use. This preliminary computation allows the system to achieve accurate sound localization without performing time-consuming measurements during operation, resolving the contradiction between measurement precision and time loss.
Solution Approach 2:
The patent introduces interpolation in the azimuth dimension to extend transfer functions from discrete measured angles to continuous angle ranges. By adding this dimensional extension through mathematical interpolation, the system achieves comprehensive angular coverage without measuring every possible angle, thus improving localization accuracy while minimizing measurement time.
2Measurement precision
If HRTFs for all directions are stored in a database to achieve complete spatial coverage, then sound synthesis quality is improved, but memory requirements increase
Solution Approach 1:
The patent segments the continuous spatial domain into discrete azimuth angles and elevation planes. By storing transfer functions only for these segmented discrete points rather than continuous space, the system achieves complete spatial coverage through interpolation while significantly reducing the quantity of data that must be stored in memory.
Solution Approach 2:
The patent separates the three-dimensional spatial problem into two independent dimensions: azimuth angles and elevation angles. Transfer functions are stored for discrete azimuth angles in the horizontal plane, and elevation effects are handled separately through interpolation and filtering. This dimensional separation reduces memory requirements while maintaining complete spatial synthesis capability.
3Measurement precision
If personalized HRTFs are used to improve sound experience, then sound quality is improved, but acquisition complexity increases
Solution Approach 1:
The patent performs personalized HRTF measurement or calibration once in advance, storing the results for repeated use. This preliminary acquisition of personalized data eliminates the need for complex real-time measurements, achieving high sound quality through pre-acquired personalized transfer functions while minimizing acquisition complexity during actual operation.
4Measurement precision
If HRTFs are measured at multiple nearby positions to enable interpolation, then estimated HRTF accuracy is improved, but measurement complexity increases
Solution Approach 1:
The patent segments the measurement task into a minimal essential set of discrete azimuth angles that form a tetrahedral or spherical configuration. By carefully selecting only these critical segmentation points rather than measuring at all possible positions, the system achieves sufficient accuracy for interpolation while minimizing measurement complexity to the essential minimum.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to an audio signal processing apparatus (100) for processing an input audio (101) signal to be transmitted to a listener in such a way that the listener perceives the input audio signal (101) to come from a virtual target position defined by an azimuth angle and an elevation angle relative to the listener, the audio signal processing apparatus (100) comprising a memory (103) configured to store a set of pairs of predefined left ear and right ear transfer functions, which are predefined for a plurality of reference positions relative to the listener, wherein the plurality of reference positions lie in a a two-dimensional plane, a determiner(105) configured to determine a pair of left ear and right ear transfer functions on the basis of the set of predefined pairs of predefined left ear and right ear transfer functions for the azimuth angle and the elevation angle of the virtual target position and an adjustment filter (107) configured to filter the input audio signal (101) on the basis of the determined pair of left ear and right ear transfer functions and an adjustment function (109) configured to adjust a delay between the left ear transfer function and the right ear transfer function of the determined pair of left ear and right ear transfer functions and a frequency dependence of the left ear transfer function and the right ear transfer function of the determined pair of left ear and right ear transfer functions as a function of the azimuth angle and/or the elevation angle of the virtual target position in order to obtain a left ear output audio signal (111a) and a right ear output audio signal (111b).