Downmix Loudness Offsets for Playback-Specific Audio Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Consumer devices struggle to consistently reproduce high-quality, wide dynamic range audio content across various media formats and playback environments due to limitations in dynamic range control and audio processing, leading to inconsistent loudness and intelligibility.
Innovation Solution
An audio processing system that includes an encoder and decoder capable of dynamic range control, where the encoder transmits dynamic range compression curves and auditory scene analysis parameters with the audio signal, allowing the decoder to customize audio processing based on the specific playback environment, ensuring consistent loudness levels and maintaining perceptual quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If dynamic range control is applied to maintain consistent loudness across different playback environments, then loudness consistency is improved, but device complexity increases due to the need for encoder and decoder with advanced audio processing capabilities
Solution Approach 1:
The encoder performs preliminary dynamic range control processing on the audio signal before encoding, calculating downmix loudness offsets and embedding them in the bitstream. This preliminary action allows the decoder to simply retrieve and apply the pre-calculated offsets without needing complex real-time analysis capabilities, thus improving loudness consistency while managing device complexity.
Solution Approach 2:
The patent introduces downmix loudness offset parameters as an intermediary element that mediates between the encoder's dynamic range control processing and the decoder's audio reproduction. These offsets serve as a compact representation of complex loudness adjustments, allowing the decoder to achieve consistent loudness across different playback environments without implementing full dynamic range control algorithms.
2Measurement precision
If downmix loudness offset parameters are embedded in the audio bitstream to enable loudness adjustment, then loudness control precision is improved, but loss of information increases due to the additional data required in the bitstream
Solution Approach 1:
The patent applies local quality by embedding downmix loudness offset parameters selectively for different audio channels and downmix configurations. Rather than uniform high-precision loudness data for all channels, the system embeds parameters only where needed and at the appropriate precision level for each specific downmix scenario, optimizing the balance between loudness control precision and bitstream efficiency.
Solution Approach 2:
The patent represents complex loudness adjustments through simplified parameter changes - specifically, downmix loudness offset values that indicate the amount of gain adjustment needed. This parameter-based approach converts complex spectral and temporal loudness analysis results into compact numerical offsets that can be efficiently stored in the bitstream while preserving the essential loudness control information.
3Adaptability or versatility
If real-time loudness analysis is performed during decoding, then adaptability to different playback environments is improved, but processing speed decreases due to the computational requirements of loudness calculation
Solution Approach 1:
The encoder performs real-time loudness analysis and calculates downmix loudness offsets during the encoding phase, storing the results in the bitstream. This shifts the computationally intensive real-time processing to the encoding stage, allowing the decoder to operate at full speed by simply retrieving and applying the pre-calculated offsets, thus maintaining both adaptability and processing speed.
Solution Approach 2:
The system implements dynamic adaptability through the downmix loudness offset parameters that are calculated based on the specific audio content and playback configuration. The offsets dynamically adjust the downmix loudness to match the reference loudness, providing environment-specific optimization without requiring real-time processing during playback, thereby maintaining high processing speed.
Data Source
AI summary
Disclosed is a non-transitory computer readable storage medium which receives, by an audio decoder (operating in a specific playback environment different from a reference channel configuration), an audio signal for the reference channel configuration. The audio signal includes audio sample data and encoder-generated loudness metadata which includes a plurality of portions of loudness metadata for a plurality of playback environments. The plurality of portions of loudness metadata includes one or more respective portions of loudness metadata for each playback environment in the plurality of playback environments. The medium also selects one or more portions of specific loudness metadata (based on the specific playback environment), from among the plurality of portions of loudness metadata for the plurality of playback environments. The one or more portions of specific loudness metadata relating to the specific playback environment determine loudness adjustment gains from the one or more portions of specific loudness metadata for the specific playback environment, apply the loudness adjustment gains as a part of overall gains applied to the audio sample data to generate output audio data.


