Audio Decoder With Frequency-Band Convolution for Lower Complexity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing methods for headphone playback, particularly for object-based content, face high computational complexity and memory requirements due to HRIR/BRIR convolution, and require impractical bit rates for delivery, especially on battery-powered devices.
Innovation Solution
A method using multi-tap convolution matrices for low frequencies and high frequency resolution in the encoder, combined with stateless matrices for higher frequencies, reduces computational complexity and memory usage by compensating for reduced decoder frequency resolution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If HRIR/BRIR convolution is applied for headphone playback, then audio spatialization and sound source localization are improved, but computational complexity and memory requirements increase significantly
Solution Approach 1:
The frequency spectrum is divided into multiple bands (low, mid, high frequencies), with different processing methods applied to each band. Multi-tap convolution matrices are used for low frequencies while stateless matrices are used for higher frequencies, reducing overall computational complexity while maintaining spatialization accuracy.
Solution Approach 2:
Different processing techniques are applied to different frequency regions based on their specific characteristics. The patent applies multi-tap convolution for low frequencies where spatial cues are critical, while using simpler stateless matrix processing for higher frequencies, optimizing the balance between quality and complexity.
2Measurement precision
If high frequency resolution is used in the encoder, then audio quality is improved, but bit rate requirements increase to impractical levels
Solution Approach 1:
The patent changes the frequency resolution parameter dynamically based on the frequency band. High frequency resolution is maintained in the encoder for low frequencies, while the decoder uses reduced frequency resolution for higher frequencies, compensated by the stateless matrix processing, thereby reducing bit rate requirements while maintaining perceived audio quality.
3Device complexity
If multi-tap convolution matrices are used for low frequencies, then computational complexity is reduced, but decoder frequency resolution is lowered
Solution Approach 1:
The frequency spectrum is segmented into low and high frequency bands, with multi-tap convolution applied to low frequencies and stateless matrices to high frequencies. This segmentation allows the system to accept reduced frequency resolution at the decoder for low frequencies while maintaining it for high frequencies, balancing complexity and quality.
Data Source
AI summary
A method for representing a second presentation of audio channels or objects as a data stream, the method comprising the steps of: (a) providing a set of base signals, the base signals representing a first presentation of the audio channels or objects; (b) providing a set of transformation parameters, the transformation parameters intended to transform the first presentation into the second presentation; the transformation parameters further being specified for at least two frequency bands and including a set of multi-tap convolution matrix parameters for at least one of the frequency bands.


