Encoded Audio Bitstream Metadata for Loudness State Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio data processing systems lack efficient methods to accurately determine and maintain consistent loudness levels across audio segments, leading to unnecessary processing and potential degradation of audio quality due to incorrect or missing metadata about loudness processing states and program boundaries.
Innovation Solution
Incorporating loudness processing state metadata (LPSM) and program boundary metadata into audio bitstreams, allowing for the verification and adaptive adjustment of loudness levels and accurate determination of audio program boundaries, thereby enabling efficient and robust loudness regulation compliance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If audio processing units operate in a blind fashion without metadata about processing history, then device complexity is reduced, but manufacturing precision deteriorates due to unnecessary processing and audio quality degradation
Solution Approach 1:
The patent applies preliminary action by embedding loudness processing state metadata (LPSM) and program boundary metadata into the audio bitstream during encoding. This allows downstream processing units to know in advance what processing has already been applied to the audio data, enabling them to skip redundant operations and maintain audio quality without requiring complex decision-making logic.
2Adaptability or versatility
If loudness processing is performed on audio data that has already been processed, then adaptability is improved, but loss of substance increases due to unnecessary processing operations
Solution Approach 1:
The patent implements feedback by including LPSM in the audio bitstream that provides information about the processing state of the audio data. Processing units can read this feedback information to determine whether loudness processing has already been applied, and adjust their behavior accordingly to avoid redundant processing that would degrade audio quality.
3Ease of operation
If DIALNORM parameter is used for loudness processing, then ease of operation is improved, but measurement precision deteriorates due to reliance on manually set values rather than automated measurements
Solution Approach 1:
The patent applies self-service by enabling processing units to automatically read and use LPSM and program boundary metadata from the audio bitstream to perform loudness processing without requiring manual DIALNORM parameter setting. The system serves itself by extracting necessary processing information directly from the encoded audio data, improving both automation and measurement accuracy.
Data Source
AI summary
Apparatus and methods for generating an encoded audio bitstream, including by including program loudness metadata and audio data in the bitstream, and optionally also program boundary metadata in at least one segment (e.g., frame) of the bitstream. Other aspects are apparatus and methods for decoding such a bitstream, e.g., including by performing adaptive loudness processing of the audio data of an audio program indicated by the bitstream, or authentication and/or validation of metadata and/or audio data of such an audio program. Another aspect is an audio processing unit (e.g., an encoder, decoder, or post-processor) configured (e.g., programmed) to perform any embodiment of the method or which includes a buffer memory which stores at least one frame of an audio bitstream generated in accordance with any embodiment of the method.


