Binaural Audio Rendering via Early-Reverberation Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio encoding technologies, such as MPEG Surround and SAOC, face challenges in efficiently representing and communicating head related binaural transfer functions, leading to high data rates and increased complexity due to the long duration of binaural transfer functions, which affects audio quality and flexibility in rendering spatial audio experiences.
Innovation Solution
The approach divides the head related binaural transfer function into an early part and a reverberation part, optimizing processing and representation for each separately, allowing for efficient synchronization and reduced data rates by using FIR filters for the early part and synthetic reverberation models for the reverberation part, enabling accurate emulation of the original transfer function with lower computational resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the head related binaural transfer function is represented in full detail, then audio quality is improved, but data rate increases
Solution Approach 1:
The head related binaural transfer function is divided into two distinct parts: an early part representing direct sound paths and early reflections, and a reverberation part representing late reverberation. This segmentation allows each part to be processed and encoded separately, enabling more efficient data representation while preserving audio quality.
Solution Approach 2:
The patent transforms the transfer function representation by converting it to the frequency domain and applying parametric modeling. Instead of transmitting the complete impulse response, the system encodes spectral parameters and synthesis information that can reconstruct the transfer function, significantly reducing data rate while maintaining audio fidelity.
2Measurement precision
If the head related binaural transfer function is represented in full detail, then audio quality is improved, but processing complexity increases
Solution Approach 1:
By separating the transfer function into early and reverberation parts, the patent reduces processing complexity for each individual part. The early part can be processed with simpler algorithms since it contains fewer reflections, while the reverberation part can use efficient parametric models rather than full impulse response processing.
Solution Approach 2:
Transforming the transfer function to the frequency domain and using parametric representation simplifies processing operations. Frequency domain convolution replaces time domain convolution, and parametric models reduce computational load compared to full impulse response processing, while preserving audio quality.
3Measurement precision
If synchronous processing of early part and reverberation part is performed, then audio quality is improved, but processing complexity increases
Solution Approach 1:
The patent calculates and stores synchronization information (time offsets) during the encoding phase, before the actual audio processing occurs. This preliminary preparation allows the decoder to easily align the early part and reverberation part without complex real-time synchronization algorithms, reducing processing complexity while maintaining audio quality.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An audio renderer comprises a receiver (801) receiving input data comprising early part data indicative of an early part of a head related binaural transfer function; reverberation data indicative of a reverberation part of the transfer function; and a synchronization indication indicative of a time offset between the early part and the reverberation part. An early part circuit (803) generates an audio component by applying a binaural processing to an audio signal where the processing depends on the early part data. A reverberator (807) generates a second audio component by applying a reverberation processing to the audio signal where the reverberation processing depends on the reverberation data. A combiner (809) generates a signal of a binaural stereo signal by combining the two audio components. The relative timing of the audio components is adjusted based on the synchronization indication by a synchronizer (805) which specifically may be a delay.