Encoded Audio Bitstream Metadata for Loudness State Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing technologies lack efficient mechanisms for managing loudness processing state and program boundaries in audio bitstreams, leading to unnecessary processing, degradation, and compliance issues with loudness regulations, particularly due to incorrect DIALNORM values and lack of metadata indicating loudness processing state.
Innovation Solution
Incorporating loudness processing state metadata (LPSM) and program boundary metadata into audio bitstreams, allowing for authentication, validation, and adaptive loudness processing, and accurate determination of program boundaries, thereby enabling efficient and compliant audio processing across diverse networks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If audio processing units operate in a blind fashion without metadata about processing history, then device complexity is reduced, but unnecessary processing occurs and audio quality degrades
Solution Approach 1:
The patent applies preliminary action by embedding processing state metadata (DIALNORM values, loudness processing flags) into the audio bitstream before transmission. This allows downstream processing units to know in advance what processing has already been applied, preventing redundant operations and maintaining audio quality without increasing processing unit complexity.
2Ease of operation
If DIALNORM parameter is manually set by content creators, then loudness processing can be performed, but incorrect DIALNORM values occur due to manual errors or non-compliant measurement methods
Solution Approach 1:
The patent applies self-service by enabling encoders to automatically measure loudness and generate accurate DIALNORM values using standardized measurement methods (such as ITU-R BS.1770). This eliminates reliance on manual content creator input, ensuring measurement precision and compliance while maintaining ease of operation through automated processing.
3Adaptability or versatility
If DIALNORM value is changed during transmission and storage, then bitstream adaptability is improved, but DIALNORM accuracy is lost and loudness processing becomes incorrect
Solution Approach 1:
The patent applies feedback by embedding processing state metadata that tracks the current DIALNORM value and processing history throughout the transmission and storage chain. Each processing unit can read this metadata, verify its accuracy, and provide feedback to maintain correct DIALNORM values, preventing unauthorized or erroneous changes while preserving bitstream adaptability for legitimate processing operations.
4Adaptability or versatility
If multiple audio processing units are scattered across a diverse network, then system versatility is improved, but processing coordination becomes difficult and redundant processing occurs
Solution Approach 1:
The patent applies universality by creating a standardized processing state metadata format that can be read and interpreted by any audio processing unit in the network, regardless of its specific function or location. This universal metadata structure enables efficient coordination across diverse networked processing units, allowing each unit to make informed decisions about whether to perform processing based on the embedded processing history, thereby improving both versatility and productivity.
Data Source
AI summary
Apparatus and methods for generating an encoded audio bitstream, including by including program loudness metadata and audio data in the bitstream, and optionally also program boundary metadata in at least one segment (e.g., frame) of the bitstream. Other aspects are apparatus and methods for decoding such a bitstream, e.g., including by performing adaptive loudness processing of the audio data of an audio program indicated by the bitstream, or authentication and/or validation of metadata and/or audio data of such an audio program. Another aspect is an audio processing unit (e.g., an encoder, decoder, or post-processor) configured (e.g., programmed) to perform any embodiment of the method or which includes a buffer memory which stores at least one frame of an audio bitstream generated in accordance with any embodiment of the method.


