Reserved-Space Audio Metadata for Adaptive Loudness Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Audio data processing systems face issues with unnecessary processing and degradation due to lack of metadata indicating loudness processing state and audio program boundaries, leading to inconsistent audio quality across diverse networks and media rendering devices.
Innovation Solution
Incorporating loudness processing state metadata and program boundary metadata into audio bitstreams, allowing audio processing units to adapt processing based on the current state of the audio data and avoid redundant operations, thereby ensuring consistent quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If audio processing units operate in a blind fashion without metadata, then device complexity is reduced, but audio quality consistency deteriorates due to unnecessary processing and degradation
Solution Approach 1:
The patent introduces metadata as an intermediary element that carries information about loudness processing state and program boundaries between audio processing units. This metadata acts as a mediator that enables processing units to make informed decisions without increasing their inherent complexity, thereby maintaining audio quality consistency across the processing chain.
Solution Approach 2:
The patent applies preliminary action by embedding loudness processing state metadata and program boundary metadata into the audio bitstream before processing. This allows subsequent processing units to know in advance what processing has already been applied and where program boundaries occur, preventing redundant operations and maintaining quality without requiring complex real-time analysis.
2Manufacturing precision
If loudness processing is performed on all audio data, then audio quality is improved, but processing efficiency deteriorates due to redundant operations
Solution Approach 1:
The patent implements feedback by including loudness processing state metadata in the audio bitstream that indicates what processing has already been applied. Processing units use this feedback information to determine whether additional loudness processing is necessary, thereby avoiding redundant operations and maintaining processing efficiency while ensuring audio quality is improved only when needed.
3Measurement precision
If DIALNORM parameter is manually set by content creator, then loudness accuracy is improved, but ease of operation deteriorates due to additional manual steps
Solution Approach 1:
The patent applies self-service by enabling processing units to automatically determine the appropriate DIALNORM value and loudness processing parameters based on the metadata embedded in the bitstream. This eliminates the need for manual content creator intervention while maintaining accurate loudness processing, as the system serves itself by interpreting the processing state metadata and program boundary information.
4Loss of information
If metadata is embedded in reserved data space, then information completeness is improved, but device complexity increases due to additional data handling
Solution Approach 1:
The patent applies universality by using existing reserved data spaces in the audio bitstream format to carry multiple types of metadata information (loudness processing state, program boundaries, DIALNORM values). This multi-functional use of reserved spaces allows comprehensive information to be embedded without requiring additional dedicated data structures, thereby minimizing the increase in device complexity while maximizing information completeness.
Data Source
Figure 1~2
Figure 3
Figure 4~7
AI summary
Apparatus and methods for generating an encoded audio bitstream, including by including program loudness metadata and audio data in the bitstream, and optionally also program boundary metadata in at least one segment (e.g., frame) of the bitstream. Other aspects are apparatus and methods for decoding such a bitstream, e.g., including by performing adaptive loudness processing of the audio data of an audio program indicated by the bitstream, or authentication and/or validation of metadata and/or audio data of such an audio program. Another aspect is an audio processing unit (e.g., an encoder, decoder, or post-processor) configured (e.g., programmed) to perform any embodiment of the method or which includes a buffer memory which stores at least one frame of an audio bitstream generated in accordance with any embodiment of the method.