Audio Bitstream Boundary Metadata for Adaptive Loudness Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing technologies lack efficient mechanisms to track and correct loudness processing states and program boundaries in audio bitstreams, leading to unnecessary processing and quality degradation across diverse networks and media processing chains.
Innovation Solution
Incorporating loudness processing state metadata (LPSM) and program boundary metadata into audio bitstreams, specifically in AC-3, E-AC-3, and Dolby E formats, to enable adaptive processing and verification, using a buffer memory, audio decoder, and parser to manage and verify metadata within the bitstream.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If audio processing units operate in a blind fashion without metadata, then device complexity is reduced, but processing accuracy and adaptability deteriorate causing unnecessary processing and quality degradation
Solution Approach 1:
The patent embeds loudness processing state metadata and program boundary metadata into the audio bitstream during encoding, so that processing units receive pre-prepared information about previous processing operations. This allows downstream units to make informed decisions without performing unnecessary processing, resolving the contradiction by providing accuracy-enhancing information in advance while keeping individual unit complexity manageable.
Solution Approach 2:
The metadata acts as feedback information that travels with the audio data through the processing chain, informing each processing unit about the processing history and current state. This feedback mechanism enables adaptive processing decisions that improve accuracy while avoiding redundant operations, effectively managing the trade-off between precision and complexity.
2Adaptability or versatility
If multiple audio processing units are scattered across a diverse network, then system versatility is improved, but processing coordination and quality consistency deteriorate
Solution Approach 1:
The patent creates a universal metadata structure that can be embedded in audio bitstreams regardless of which processing unit handles the data. The loudness processing state metadata and program boundary metadata provide universal information that any processing unit in the distributed network can interpret and act upon, ensuring consistent quality across diverse networked units while maintaining distribution flexibility.
3Productivity
If volume leveling is performed on every audio clip without checking processing history, then processing completeness is improved, but energy efficiency and quality deteriorate due to redundant processing
Solution Approach 1:
The loudness processing state metadata is prepared in advance and embedded in the bitstream, allowing processing units to quickly determine whether volume leveling has already been applied. This preliminary information enables units to skip redundant processing operations while maintaining complete processing coverage where needed, improving energy efficiency without sacrificing processing thoroughness.
4Measurement precision
If DIALNORM parameter is manually set by content creator, then processing precision can be improved, but operational complexity and time consumption increase
Solution Approach 1:
The system performs self-service by automatically measuring dialog loudness and setting the DIALNORM parameter based on embedded loudness metadata, eliminating the need for manual content creator intervention. This automation maintains measurement precision through algorithmic analysis while dramatically simplifying operation and reducing time consumption compared to manual parameter setting.
Data Source
Figure 1~2
Figure 3
Figure 4~7
AI summary
Apparatus and methods for generating an encoded audio bitstream, including by including program loudness metadata and audio data in the bitstream, and optionally also program boundary metadata in at least one segment (e.g., frame) of the bitstream. Other aspects are apparatus and methods for decoding such a bitstream, e.g., including by performing adaptive loudness processing of the audio data of an audio program indicated by the bitstream, or authentication and/or validation of metadata and/or audio data of such an audio program. Another aspect is an audio processing unit (e.g., an encoder, decoder, or post-processor) configured (e.g., programmed) to perform any embodiment of the method or which includes a buffer memory which stores at least one frame of an audio bitstream generated in accordance with any embodiment of the method.