Audio Stream Dependency Metadata for Spatial Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio codecs face limitations in processing multiple audio streams due to varying coding complexity, which affects the immersive spatial audio experience by restricting the number of supported streams.
Innovation Solution
A method and apparatus that utilize metadata to determine dependencies between individual audio streams, encoding related streams as combined multichannel signals and independent streams as mono channels, thereby reducing computational complexity and enhancing processing efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the audio codec processes more individual audio streams, then the immersive spatial audio experience is improved, but the coding complexity increases
Solution Approach 1:
The patent combines multiple individual audio streams into a single multichannel audio signal when dependency analysis shows they are related. This merging reduces the number of separate encoding operations needed, thereby lowering coding complexity while preserving spatial audio quality through the dependency metadata that maintains stream relationships.
Solution Approach 2:
The patent segments audio streams into independent and dependent groups based on dependency field analysis. Independent streams are encoded separately as mono channels, while dependent streams are combined. This segmentation allows the codec to optimize processing for each group, reducing overall complexity while supporting more streams.
2Adaptability or versatility
If the audio codec processes more individual audio streams, then the capacity for spatial audio channels is increased, but the processing time increases
Solution Approach 1:
The patent performs preliminary dependency analysis on audio streams before encoding by examining metadata dependency fields. This pre-processing step identifies which streams can be combined and which must be encoded separately, allowing the codec to optimize the encoding strategy in advance and reduce actual processing time during the encoding phase.
Solution Approach 2:
By combining dependent audio streams into a single multichannel signal before encoding, the patent reduces the total number of encoding operations required. This merging significantly decreases processing time while maintaining the capacity to handle multiple spatial audio channels through the combined signal.
3Productivity
If the audio codec reduces coding complexity, then the processing efficiency is improved, but the number of supported audio streams is limited
Solution Approach 1:
The patent dynamically adjusts the encoding strategy based on dependency analysis of audio streams. The codec automatically determines whether to combine or separately encode streams based on their interdependencies, optimizing processing efficiency for each specific input configuration while maintaining flexibility to support varying numbers of audio streams.
Solution Approach 2:
The patent creates a universal encoding framework that handles both independent and dependent audio streams through a single codec architecture. The dependency field metadata enables the same codec to adaptively process different stream configurations, maintaining high processing efficiency across various scenarios without requiring separate specialized encoders.
Data Source
AI summary
There is disclosed inter alia an apparatus comprising means for receiving an audio format comprising a plurality of individual audio signal streams and metadata, wherein the metadata comprises a dependency field associated with each of the plurality of individual audio signal streams; means for determining that a dependency field associated with a first individual audio signal stream of the plurality of individual audio signal streams indicates that the first individual audio signal stream is related to a second individual audio signal stream of the plurality of individual audio signal streams; and means for encoding the first and second individual audio signal streams as a combined multichannel audio signal by an audio encoder.


