Audio Object Encoding for Multichannel Compatibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The Spatial Audio Object Coding (SAOC) standard primarily supports mono- and stereo downmixes, limiting its compatibility with multi-channel applications such as DVD and Blu-Ray, which requires substantial modifications to the standard specification, reducing backwards compatibility and increasing complexity.
Innovation Solution
An audio object encoding and decoding system that mixes N audio objects into M audio channels, deriving K audio channels (where K=1 or 2) to generate an output data stream with audio object upmix parameters, allowing multichannel support without modifying existing standards, enabling efficient reuse of existing functionality and improved backwards compatibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the SAOC standard is modified to support multichannel downmixes, then multichannel application compatibility is improved, but standard complexity and modification difficulty increase
Solution Approach 1:
The patent segments the multichannel audio processing into two independent parts: (1) encoding of M audio channels using existing multichannel codecs, and (2) encoding of audio object parameters relative to K channels (K=1 or 2). This segmentation allows the system to support multichannel applications without modifying the core SAOC standard, as each component can be processed independently using existing technology.
Solution Approach 2:
The patent creates a universal encoding framework that can handle both traditional multichannel audio and audio object-based coding. By deriving audio object parameters from K channels that are themselves derived from M audio channels, the system achieves multi-functionality: it can reproduce the original M-channel signal for legacy devices while simultaneously providing object-level control for advanced applications.
2Adaptability or versatility
If existing standards are modified to support new functionality, then functionality is improved, but backwards compatibility decreases
Solution Approach 1:
The patent introduces K audio channels as an intermediary between the M audio channels and the audio object parameters. The K channels serve as a bridge that allows audio object parameters to be defined without directly modifying the M-channel signal structure. This intermediary approach enables new functionality while preserving the original signal representation for backwards compatibility.
Solution Approach 2:
The patent performs preliminary downmixing of M audio channels to K channels before deriving audio object parameters. This preliminary action creates a stable intermediate representation that can be processed using existing SAOC technology, ensuring that the core encoding/decoding pipeline remains unchanged and backwards compatible while enabling extended functionality.
3Measurement precision
If audio object parameters are defined relative to M audio channels, then multichannel precision is improved, but data rate and complexity increase
Solution Approach 1:
The patent extracts the essential spatial information from M audio channels by deriving audio object parameters relative to K channels (K=1 or 2). This extraction process separates the critical spatial relationships needed for audio object coding from the full M-channel signal, reducing the amount of data that needs to be encoded and transmitted while preserving the necessary spatial precision for audio object manipulation.
Data Source
AI summary
An audio object encoder comprises a receiver (701) which receives N audio objects. A downmixer (703) downmixes the N audio objects to M audio channels, and a channel circuit (707) derives K audio channels from the M audio channels, K=1, 2 and K<M. A parameter circuit (709) generates audio object upmix parameters for at least part of each of the N audio objects relative to the K audio channels and an output circuit (705, 711) generates an output data stream comprising the audio object upmix parameters and the M audio channels. An audio object decoder receives the data stream and includes a channel circuit (805) deriving K audio channels from the M channel downmix; and an object decoder (807) for generating at least part of each of the N audio objects by upmixing the K audio channels based on the audio object upmix parameters. The invention may allow improved object encoding while maintaining backwards compatibility.


