Downmix Loudness Offsets for Consistent Speaker Playback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Consumer devices struggle to consistently reproduce high-quality, wide bandwidth audio content across varying media formats and playback environments due to limitations in dynamic range control and loudness consistency.
Innovation Solution
An audio processing system that includes an encoder and decoder capable of dynamic range control, using dynamic range compression curves, auditory scene analysis, and gain smoothing to adjust loudness levels based on specific playback environments, ensuring consistent loudness and spatial balance across different speaker configurations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If dynamic range control is applied to audio content, then loudness consistency across different playback environments is improved, but device complexity increases due to the need for encoder and decoder processing
Solution Approach 1:
The encoder pre-calculates and embeds dynamic range control metadata including downmix loudness offsets before audio playback. This preliminary action allows the decoder to apply loudness adjustments without complex real-time analysis, resolving the contradiction by shifting processing burden to the encoding stage while maintaining reliability.
Solution Approach 2:
The patent introduces downmix loudness offset metadata as an intermediary parameter that bridges the gap between multi-channel source audio and stereo downmixed output. This intermediary carries loudness information that enables consistent reproduction across different playback environments without requiring complex processing at the decoder stage.
2Measurement precision
If downmix loudness offset data is embedded in the audio signal, then loudness accuracy in downmixed content is improved, but information processing requirements increase
Solution Approach 1:
The patent extracts downmix loudness offset data as a separate metadata parameter from the main audio signal. This extraction allows the loudness information to be processed independently and applied only when needed for downmixing operations, improving loudness accuracy while minimizing unnecessary data processing for other audio functions.
Solution Approach 2:
The patent represents loudness information as a compact offset parameter in dBFS units that can be easily stored and transmitted. This parameter change from complex loudness measurements to simple offset values maintains measurement precision while significantly reducing information processing requirements.
3Reliability
If dynamic range compression is applied to prevent clipping, then audio quality is improved, but the natural dynamic range of the audio is reduced
Solution Approach 1:
The patent applies dynamic range compression selectively based on the downmix loudness offset values. Instead of uniform compression, the system dynamically adjusts compression levels to prevent clipping only where necessary, thereby maintaining audio quality while preserving as much of the natural dynamic range as possible in content that doesn't require compression.
Solution Approach 2:
The patent applies different dynamic range compression characteristics to different portions of the audio signal based on local loudness requirements. By using the downmix loudness offset data to identify specific time regions or frequency bands that need clipping prevention, the system maintains high audio quality while preserving dynamic range in regions where it is not needed.
Data Source
AI summary
Audio content coded for a reference speaker configuration is downmixed to downmix audio content coded for a specific speaker configuration. One or more gain adjustments are performed on individual portions of the downmix audio content coded for the specific speaker configuration. Loudness measurements are then performed on the individual portions of the downmix audio content. An audio signal that comprises the audio content coded for the reference speaker configuration and downmix loudness metadata is generated. The downmix loudness metadata is created based at least in part on the loudness measurements on the individual portions of the downmix audio content.


