Downmixed Audio Loudness Adjustment Using Segment Gain Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing media processing devices struggle to consistently reproduce high-quality, wide bandwidth, and wide dynamic range audio content across varying media formats and content types, often resulting in inconsistent loudness and intelligibility.
Innovation Solution
The implementation of dynamic range control techniques, including the transmission of dynamic range compression curves and reference loudness levels, allows for customized audio processing operations in various playback environments, ensuring consistent loudness and intelligibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If dynamic range control techniques are implemented with transmission of dynamic range compression curves and reference loudness levels, then consistent loudness and intelligibility are achieved across varying media formats and content types, but device complexity increases
Solution Approach 1:
The encoder pre-calculates and transmits dynamic range compression curves and reference loudness levels in the metadata before playback. This preliminary preparation allows the decoder to apply consistent loudness adjustment without performing complex real-time calculations, resolving the contradiction between reliability and device complexity
Solution Approach 2:
Dynamic range compression curves and reference loudness levels serve as intermediary parameters that bridge the encoder and decoder. These transmitted metadata enable the decoder to achieve consistent loudness reproduction across different devices without requiring each device to implement identical complex processing algorithms
2Reliability
If loudness levels are dynamically adjusted based on playback environments, then listening experience is improved and clipping is prevented, but processing time increases
Solution Approach 1:
The system pre-determines dynamic range compression parameters and reference loudness levels during encoding. During playback, the decoder simply applies these pre-calculated parameters to adjust loudness dynamically, preventing clipping without requiring time-consuming real-time analysis
Solution Approach 2:
The system implements dynamic loudness adjustment by applying pre-determined compression curves that adapt to different playback environments. The decoder uses the transmitted reference loudness levels to dynamically scale audio output, achieving clipping prevention with minimal processing time
3Adaptability or versatility
If downmix operations are performed from multi-channel to two-channel configuration, then spatial balance may be compromised, but device compatibility is improved
Solution Approach 1:
Downmix loudness offset parameters serve as intermediary metadata that guide the downmixing process from multi-channel to two-channel configurations. These parameters enable the decoder to maintain spatial balance during downmixing by providing reference loudness levels for each channel, achieving both device compatibility and spatial balance
Data Source
AI summary
Audio content coded for a reference speaker configuration is downmixed to downmix audio content coded for a specific speaker configuration. One or more gain adjustments are performed on individual portions of the downmix audio content coded for the specific speaker configuration. Loudness measurements are then performed on the individual portions of the downmix audio content. An audio signal that comprises the audio content coded for the reference speaker configuration and downmix loudness metadata is generated. The downmix loudness metadata is created based at least in part on the loudness measurements on the individual portions of the downmix audio content.


