Audio Object Metadata Aggregates Signal Processing Operations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing systems face inefficiencies and errors due to repeated, irreversible, and computationally intensive operations across multiple processors in the audio processing chain, leading to complexity, artifacts, and precision loss.
Innovation Solution
The use of object audio metadata (OAMD) to aggregate signal processing operations, where upstream processors generate audio object gain components for deferred operations and include them in OAMD, allowing downstream rendering stages to perform these operations, reducing computational complexity and improving acoustic output quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multiple audio processors perform repeated signal processing operations in the audio processing chain, then audio content can be processed and delivered to end user devices, but system complexity increases and processing efficiency decreases
Solution Approach 1:
The patent merges multiple repeated signal processing operations into a single aggregation point at the rendering stage. Instead of having multiple audio processors independently performing similar operations (peak limiting, crossfading, program mixing), the system aggregates these operations and executes them once at the rendering stage, thereby reducing overall system complexity while maintaining processing functionality.
Solution Approach 2:
The patent applies preliminary action by having upstream audio processors prepare and package their processing parameters into metadata before the actual signal processing occurs. The rendering stage then receives this pre-prepared metadata and executes the aggregated operations efficiently, avoiding the need for multiple processors to independently perform the same operations.
2Reliability
If signal processing operations are performed by multiple audio processors, then audio processing functionality is distributed, but computational intensity increases and errors accumulate
Solution Approach 1:
The patent combines multiple computational operations into a single rendering stage, reducing the total computational intensity distributed across multiple processors. By aggregating operations such as peak limiting, crossfading, and program mixing into one location, the system performs fewer total calculations while maintaining processing accuracy and reducing error accumulation.
3Manufacturing precision
If repeated peak limiting operations are performed by multiple processors, then audio dynamics are controlled, but irreversible processing and artifacts are introduced
Solution Approach 1:
The patent merges peak limiting operations into a single execution at the rendering stage, eliminating the harmful effect of repeated irreversible processing. Instead of multiple processors independently applying peak limiting which introduces artifacts and precision loss, the aggregated operation performs the limiting function once with higher precision, preserving audio quality.
4Speed
If audio processing operations are distributed across multiple stages, then processing flexibility is maintained, but delays and processing time increase
Solution Approach 1:
The patent applies preliminary action by having upstream processors prepare processing parameters in advance and package them into metadata. This allows the rendering stage to execute aggregated operations more efficiently without requiring complex coordination between multiple processors, thereby reducing processing delays while maintaining flexibility.
Data Source
Figure 1A
Figure 1B
Figure 2A
AI summary
A method comprising: receiving and decoding a coded bitstream encoded with audio content including first audio objects corresponding to a first media content type of two consecutive media content types and second audio objects corresponding to a second media content type of the two consecutive media content types, and audio metadata corresponding to the audio content, the audio metadata including first and second audio object gains, respectively for the first and second audio objects, generated at least in part based on a first fading curve of the first media content type and a second fading curve of the second media content type, respectively; applying the first and second audio object gains to the first and second audio objects, respectively; rendering a sound field represented by the first audio object with the applied first audio object gain and the second audio object with the applied second audio object gain.