Downmix Loudness Adjustment for Clipping-Safe Audio Playback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Consumer devices struggle to consistently reproduce high-quality, wide bandwidth and dynamic range audio content across various media formats and playback environments due to limitations in dynamic range control and audio processing.
Innovation Solution
An audio encoder and decoder system that transmits dynamic range compression curves and auditory scene analysis parameters, allowing for customizable audio processing to maintain consistent loudness levels and spatial balance across different playback environments, using techniques like Huffman coding and differential coding for efficient gain management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If dynamic range compression is applied to audio content, then loudness consistency across playback environments is improved, but audio processing complexity increases
Solution Approach 1:
The encoder pre-calculates and embeds DRC gain values and dynamic range compression curves into the audio bitstream before transmission. This preliminary action allows the decoder to apply pre-determined gain adjustments without performing complex real-time analysis, thereby improving loudness consistency while managing processing complexity through advance preparation.
Solution Approach 2:
The patent introduces an intermediary data structure (DRC gain values and compression curves) that mediates between the encoder and decoder. These intermediary elements carry the necessary information for loudness normalization, enabling consistent reproduction across different playback environments without requiring complex processing at either endpoint.
2Adaptability or versatility
If downmixing is performed to adapt multi-channel audio to stereo playback, then compatibility with limited playback devices is improved, but spatial balance and loudness accuracy deteriorate
Solution Approach 1:
The patent applies different processing strategies to different audio channels during downmixing. Specifically, it calculates separate DRC gain values for each channel based on their individual loudness characteristics, then combines them with appropriate spatial weighting. This local quality approach preserves spatial balance by treating each channel's loudness independently while maintaining their relative spatial relationships in the downmixed stereo output.
Solution Approach 2:
The downmixing process uses dynamic gain adjustment based on the program material's loudness characteristics. The system continuously analyzes the audio content and adjusts the downmix gains in real-time to maintain accurate spatial balance and loudness relationships, rather than using fixed static gains. This dynamic adaptation ensures compatibility across devices while preserving spatial accuracy.
3Reliability
If gain adjustments are applied to prevent clipping in downmixed audio, then audio quality is protected, but loudness level accuracy may be compromised
Solution Approach 1:
The system implements a feedback mechanism where the encoder monitors the downmixed audio signal for potential clipping conditions and adjusts the DRC gain values accordingly. By continuously analyzing the combined signal levels and providing feedback to the gain calculation process, the system can apply protective gain reduction only when necessary to prevent clipping, while maintaining accurate loudness levels for the majority of the audio content.
Solution Approach 2:
The patent dynamically changes the DRC gain parameter based on the program material's characteristics and the specific playback configuration. By adjusting the gain parameter in response to measured loudness levels and clipping risk assessment, the system optimizes both protection against clipping and accuracy of loudness reproduction, adapting the parameter values to each specific audio segment and playback scenario.
Data Source
AI summary
Audio content coded for a reference speaker configuration is downmixed to downmix audio content coded for a specific speaker configuration. One or more gain adjustments are performed on individual portions of the downmix audio content coded for the specific speaker configuration. Loudness measurements are then performed on the individual portions of the downmix audio content. An audio signal that comprises the audio content coded for the reference speaker configuration and downmix loudness metadata is generated. The downmix loudness metadata is created based at least in part on the loudness measurements on the individual portions of the downmix audio content.


