Audio Bitstream Metadata Container for Adaptive Loudness Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio data processing systems fail to account for the processing history of audio data, leading to unnecessary and potentially degrading processing, especially in networks with multiple audio units, and lack metadata for loudness processing state and program boundaries, affecting quality and compliance with regulations.
Innovation Solution
Incorporating a metadata container in audio bitstreams with loudness processing state metadata (LPSM) and program boundary metadata, allowing for verification and adaptive processing, and reducing redundant processing by including this metadata in reserved data spaces.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If audio data is processed through multiple audio processing units in a network, then the audio data can be adapted to different devices and formats, but redundant processing occurs and audio quality degrades
Solution Approach 1:
The patent applies preliminary action by embedding processing state metadata (DRC_APPLIED, LAAPplied, LRA_applied flags) into the audio bitstream during encoding. This allows downstream audio processing units to预先 know what processing has already been applied, enabling them to skip redundant processing steps and avoid quality degradation while maintaining adaptability to different devices.
2Stability of the object's composition
If volume leveling processing is applied to audio data, then loudness consistency is improved, but processing time increases and may be redundant
Solution Approach 1:
The patent implements feedback by including processing state metadata (DRC_APPLIED, LAAPplied, LRA_applied flags) in the audio bitstream that provides information about previously applied loudness processing. This feedback mechanism allows audio processing units to determine whether volume leveling has already been applied, enabling them to skip redundant processing and reduce processing time while maintaining loudness consistency when needed.
3Stability of the object's composition
If DIALNORM parameter is used for loudness processing, then playback level consistency is improved, but the parameter may be incorrect and requires manual setting
Solution Approach 1:
The patent applies self-service by automatically generating and embedding processing state metadata (DRC_APPLIED, LAAPplied, LRA_applied flags) into the audio bitstream during encoding. This eliminates the need for manual DIALNORM setting and provides automated information about processing state, allowing the system to self-manage loudness processing without human intervention while maintaining playback level consistency.
4Loss of information
If metadata container is added to audio bitstream, then processing state information is preserved, but data structure complexity increases
Solution Approach 1:
The patent applies segmentation by dividing the metadata container into distinct, well-defined sections: DRC_metadata_section with DRC_APPLIED flag, LAA_metadata_section with LAAPplied flag, and LRA_metadata_section with LRA_applied flag. Each section handles a specific type of processing state information, making the overall structure more manageable and easier to process despite the increased information capacity.
Data Source
Figure 1~2
Figure 3
Figure 4~7
AI summary
Apparatus and methods for generating an encoded audio bitstream, including by including program loudness metadata and audio data in the bitstream, and optionally also program boundary metadata in at least one segment (e.g., frame) of the bitstream. Other aspects are apparatus and methods for decoding such a bitstream, e.g., including by performing adaptive loudness processing of the audio data of an audio program indicated by the bitstream, or authentication and/or validation of metadata and/or audio data of such an audio program. Another aspect is an audio processing unit (e.g., an encoder, decoder, or post-processor) configured (e.g., programmed) to perform any embodiment of the method or which includes a buffer memory which stores at least one frame of an audio bitstream generated in accordance with any embodiment of the method.