Audio Scene Decoding via Selective Matrix Element Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio coding systems for parametric spatial audio face challenges in complex mathematical reconstruction of audio objects, requiring assumptions about audio content properties and increasing computational complexity with the number of channels.
Innovation Solution
The proposed method generates a bit stream containing downmix signals and reconstruction matrix elements, allowing for simpler and more flexible reconstruction of audio objects on the decoder side, with reduced computational complexity and backward compatibility with legacy decoders.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional parametric spatial audio coding systems are used, then audio objects can be reconstructed, but the mathematical complexity increases dramatically as the number of channels increases
Solution Approach 1:
The patent extracts only the essential information needed for reconstruction by transmitting a subset of matrix elements rather than the complete reconstruction matrix. This selective transmission approach reduces the amount of data and computational complexity required while maintaining the ability to reconstruct audio objects accurately.
Solution Approach 2:
Instead of transmitting the complete reconstruction matrix and performing complex mathematical operations at the decoder, the patent inverts the approach by transmitting only essential matrix elements and using simpler inversion principles to achieve the same reconstruction result with reduced complexity.
2Reliability
If traditional audio coding systems are used, then audio reconstruction can be achieved, but additional assumptions about audio content properties are required
Solution Approach 1:
The patent incorporates feedback mechanisms where the decoder adapts to actual audio content characteristics during the decoding process. This allows the system to maintain high reconstruction accuracy without requiring predefined assumptions about audio content properties, as the system learns and adjusts to the actual content characteristics.
3Manufacturing precision
If complex algorithms are deployed in consumer devices, then audio reconstruction quality improves, but the devices become difficult or impossible to upgrade
Solution Approach 1:
The patent segments the audio decoding process into modular components with standardized interfaces. This modular architecture allows different versions of the decoding algorithm to be implemented as separate modules that can be independently upgraded without affecting the entire system, thus maintaining both high reconstruction quality and ease of upgradability.
4Loss of information
If more channels are used in the downmix, then audio spatial information is improved, but the computational complexity and intelligence required on the decoder side increase dramatically
Solution Approach 1:
The patent extracts only the essential spatial information needed for accurate audio reconstruction by transmitting a reduced set of matrix elements. This selective extraction approach preserves spatial information quality while significantly reducing the computational complexity required for processing and reconstruction.
Data Source
AI summary
Exemplary embodiments provide encoding and decoding methods, and associated encoders and decoders, for encoding and decoding of an audio scene which is represented by one or more audio signals. The encoder generates a bit stream which comprises downmix signals and side information which includes individual matrix elements of a reconstruction matrix which enables reconstruction of the one or more audio signals in the decoder.


