Audio Object Decoder Frequency Band Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing object-based audio systems face challenges in efficiently encoding and decoding audio objects while maintaining quality, particularly due to the complexity of parametric coding methods like MPEG SAOC, which require high bit rates and rely on assumptions about audio object properties.
Innovation Solution
The proposed method involves a decoder and encoder system that uses indicators to select and weight downmix signals and decorrelated signals across frequency bands, reducing bit rate requirements and complexity by employing decoding modes that simplify parameter transmission and reconstruction, such as using binary vectors and entropy coding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If parametric coding methods like MPEG SAOC are used to encode audio objects, then the bit rate can be reduced compared to conventional channel-based coding, but the decoding complexity increases and reconstruction quality deteriorates due to mathematical complexity and reliance on assumptions about audio object properties
Solution Approach 1:
The patent segments the frequency spectrum into multiple bands and processes each band independently. The upmix matrix is divided into frequency-dependent components, allowing selective reconstruction based on the frequency band. This segmentation reduces the complexity of processing the entire audio object by handling smaller, more manageable frequency portions separately.
Solution Approach 2:
The patent applies different reconstruction strategies to different frequency bands. For each frequency band, the system determines whether to use a full upmix matrix or a simplified version based on local characteristics. This allows the system to maintain high quality where needed while reducing complexity where the audio content is simpler, resolving the contradiction between quality and complexity.
2Quantity of substance
If parametric coding methods like MPEG SAOC are used to encode audio objects, then the bit rate can be reduced, but the reconstruction quality deteriorates due to mathematical complexity and reliance on assumptions about audio object properties
Solution Approach 1:
The patent dynamically adapts the reconstruction process based on the frequency band and audio content characteristics. The system determines in real-time whether to use full parametric reconstruction or simplified approaches for each frequency band, allowing quality to be optimized where needed while maintaining lower bit rates overall. This dynamic adaptation resolves the contradiction by making quality consistent with the available bit rate allocation.
Solution Approach 2:
The patent changes the parameters used for reconstruction based on frequency band characteristics. Different upmix matrices and weighting parameters are applied to different frequency bands, allowing the system to optimize reconstruction quality for each band while maintaining overall efficiency. This parameter adaptation enables better quality at lower bit rates by tailoring the reconstruction parameters to the specific frequency content.
3Manufacturing precision
If full upmix matrices are used for reconstructing audio objects across all frequency bands, then reconstruction quality is maintained, but the bit rate for transmitting parameters increases
Solution Approach 1:
The patent segments the frequency spectrum and applies different upmix matrix representations to different bands. Instead of transmitting a full upmix matrix for all frequencies, the system transmits simplified or selective matrices for each frequency band, reducing the total parameter bit rate while maintaining quality through frequency-specific optimization.
Solution Approach 2:
The patent applies partial upmix matrices only where necessary. For frequency bands where the audio content is simple or where full reconstruction is not needed, the system uses simplified matrices, reducing the overall parameter transmission requirement. This partial application maintains sufficient quality while significantly reducing the bit rate compared to full matrix transmission across all bands.
Data Source
AI summary
This disclosure falls into the field of audio coding, in particular it is related to the field of spatial audio coding, where the audio information is represented by multiple signals, where the signals may comprise audio channels or/and audio objects. In particular the disclosure provides a method and apparatus for reconstructing audio objects in an audio decoding system. Furthermore, this disclosure provides a method and apparatus for encoding such audio objects.


