Audio Bitstream Metadata for Adaptive Loudness Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing technologies lack efficient methods to accurately determine and maintain consistent loudness levels across audio bitstreams, leading to unnecessary processing and potential degradation of audio quality due to the absence of metadata indicating loudness processing state and program boundaries.
Innovation Solution
Incorporating loudness processing state metadata (LPSM) and program boundary metadata into audio bitstreams, allowing for the verification and adaptive processing of audio data, ensuring compliance with regulations like the CALM Act and improving audio quality by avoiding redundant processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If audio processing units operate in a blind fashion without metadata, then device complexity is reduced, but processing efficiency deteriorates due to redundant operations
Solution Approach 1:
Loudness processing state metadata is embedded in the audio bitstream during encoding, providing advance information about processing history before the audio reaches subsequent processing units. This allows processing units to make informed decisions about whether to apply loudness processing, avoiding redundant operations while maintaining simple device architecture.
2Stability of the object's composition
If loudness processing is applied to all audio segments, then consistent playback loudness is achieved, but audio quality deteriorates due to unnecessary processing
Solution Approach 1:
The loudness processing state metadata provides feedback information about the processing history of audio segments. Processing units can read this metadata to determine whether loudness processing has already been applied, and adjust their processing accordingly. This feedback mechanism ensures consistent playback loudness while avoiding redundant processing that would degrade audio quality.
3Loss of information
If DIALNORM parameter is manually set by content creators, then loudness information can be provided, but measurement accuracy deteriorates due to human error and lack of standardization
Solution Approach 1:
The encoding system automatically measures loudness and generates loudness processing state metadata without requiring manual intervention from content creators. This self-service approach eliminates human error in setting DIALNORM parameters, ensures consistent application of loudness measurement standards, and provides accurate loudness information throughout the audio processing chain.
4Adaptability or versatility
If multiple audio processing units are scattered across a network, then system versatility is improved, but processing coordination deteriorates due to lack of processing history information
Solution Approach 1:
Loudness processing state metadata acts as an intermediary information carrier between distributed audio processing units. As audio segments traverse the network from one processing unit to another, this metadata travels with the audio data, providing each processing unit with knowledge of previous processing operations. This enables reliable coordination across the distributed system while maintaining configuration flexibility.
Data Source
Figure 1~2
Figure 3
Figure 4~7
AI summary
Apparatus and methods for generating an encoded audio bitstream, including by including program loudness metadata and audio data in the bitstream, and optionally also program boundary metadata in at least one segment (e.g., frame) of the bitstream. Other aspects are apparatus and methods for decoding such a bitstream, e.g., including by performing adaptive loudness processing of the audio data of an audio program indicated by the bitstream, or authentication and/or validation of metadata and/or audio data of such an audio program. Another aspect is an audio processing unit (e.g., an encoder, decoder, or post-processor) configured (e.g., programmed) to perform any embodiment of the method or which includes a buffer memory which stores at least one frame of an audio bitstream generated in accordance with any embodiment of the method.