Audio Decoder Loudness Leveling With Metadata-Based DRC
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio playback technologies struggle with real-time loudness management and dynamic range control, especially in dynamic encoding scenarios, as they often alter the creative intent of audio content and fail to adapt to device constraints and user preferences.
Innovation Solution
A metadata-based approach for dynamic processing of audio data that includes encoding metadata for loudness leveling and dynamic range compression, allowing decoders to apply personalized and device-specific processing parameters in real-time, using MPEG-D DRC and MPEG-H 3D audio syntax.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the entire audio program is analyzed prior to encoding to determine loudness and DRC parameters, then loudness compliance and dynamic range control are achieved, but the audio content loses its original dynamic range and creative intent
Solution Approach 1:
The audio processing is segmented into two independent parts: the original audio data is encoded without processing, while separate metadata containing loudness and DRC parameters is generated from analysis of the original program. This allows the audio to retain its original dynamic range while the metadata provides the necessary control information for compliance.
Solution Approach 2:
Metadata acts as an intermediary between the original audio content and the playback device. The metadata carries loudness and DRC parameters that enable the playback device to perform loudness management and dynamic range control without directly processing the original audio, thus preserving creative intent while achieving compliance.
2Productivity
If loudness processing is applied in real-time encoding, then loudness compliance is achieved, but the audio is optimized for a single playback environment and cannot adapt to device constraints
Solution Approach 1:
The loudness and DRC parameters are determined in advance from analysis of the original audio program and encoded into metadata. This preliminary action allows real-time encoding to proceed efficiently while the stored metadata enables adaptive processing at playback to suit different device constraints and environments.
Solution Approach 2:
The system transitions from static, single-environment optimization to dynamic adaptability. The metadata contains parameters that allow the playback device to dynamically adjust loudness and dynamic range based on device capabilities, acoustic environment, and user preferences, making the system versatile across different playback scenarios.
3Ease of operation
If dynamic range control is applied to match device capabilities, then playback quality is improved, but the original dynamic range of the audio content is altered
Solution Approach 1:
Metadata serves as an intermediary that carries DRC parameters without directly modifying the original audio. The playback device uses these parameters to control dynamic range adaptation, allowing high-quality playback on devices with limited dynamic range capabilities while preserving the original audio composition.
Solution Approach 2:
The system changes parameters (loudness and DRC values) in the metadata rather than in the audio data itself. This allows the playback device to apply appropriate parameter adjustments based on device capabilities, improving playback quality without permanently altering the original audio content's dynamic range.
Data Source
AI summary
Decoder apparatus, computer program and methods of processing audio data for playback are described. They include receiving a bitstream including encoded audio data and metadata that includes DRC set(s), and for each DRC set, an indication of whether the DRC set is configured for providing a loudness leveling effect. The metadata further includes personalization experience information. The method further includes identifying DRC sets that are configured for providing the dynamic range compensation effect; decoding the encoded audio data to obtain decoded audio data; selecting one of the identified DRC sets configured for providing the loudness leveling effect; extracting from the bitstream one or more DRC gains corresponding to the selected DRC set; applying to the decoded audio data the one or more DRC gains corresponding to the selected DRC set to obtain dynamic loudness compensated audio data; and outputting the dynamic loudness compensated audio data for playback.


