Downmix Loudness Metadata for Consistent Multi-Speaker Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Consumer devices struggle to consistently reproduce high-quality, wide dynamic range audio content across varying media formats and playback environments due to inconsistencies in loudness and intelligibility, as they often cannot distinguish between different audio processing parameters like dynamic range control, gain smoothing, and gain limiting.

Innovation Solution

An audio processing system that encodes dynamic range compression curves and other metadata with audio signals, allowing decoders to customize audio processing based on specific playback environments, differentiate loudness levels, and maintain spatial balance across channels, using techniques like auditory scene analysis and differential coding to adjust gains dynamically.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If dynamic range control, gain smoothing, and gain limiting are applied to audio signals, then audio quality and loudness consistency are improved, but device complexity increases due to the need to process and distinguish multiple audio processing parameters

Engineering Contradiction:
Improveaudio qualityVSAvoiddevice complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments audio processing parameters by separating downmix loudness parameters from other audio parameters. The downmix loudness parameters are extracted and processed independently to determine loudness compensation gains, while other parameters are handled separately. This segmentation reduces the complexity of processing all parameters simultaneously while maintaining audio quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts downmix loudness parameters from the encoded audio signal metadata and processes them separately to determine loudness compensation gains. This extraction allows the decoder to focus on specific loudness adjustment without being overwhelmed by the full set of audio processing parameters, reducing computational complexity while improving loudness consistency.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If loudness compensation is applied to downmixed audio content, then loudness consistency across playback environments is improved, but processing time increases due to additional analysis and adjustment steps

Engineering Contradiction:
Improveloudness consistencyVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by encoding downmix loudness parameters and loudness metadata with the audio signal during the encoding stage. This pre-computation of loudness information allows the decoder to quickly retrieve and apply loudness compensation gains without performing complex real-time analysis, thereby reducing processing time while maintaining loudness consistency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements self-service by embedding all necessary loudness information (downmix loudness parameters, reference loudness levels, and loudness metadata) directly in the encoded audio signal. The decoder can independently determine and apply loudness compensation without requiring external analysis or additional processing steps, reducing processing time while ensuring reliable loudness consistency.

Inventive Principle:
Principle #25Self-service

3Productivity

If downmix loudness parameters are extracted and processed separately, then processing efficiency is improved, but information loss may occur if critical audio parameters are not properly preserved

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidinformation loss
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent uses feedback by utilizing the extracted downmix loudness parameters to determine loudness compensation gains, which are then applied to the downmixed audio content. The process includes measuring the loudness of the downmixed content and comparing it to reference loudness levels, using the difference to adjust gains. This feedback mechanism ensures that critical audio information is preserved while improving processing efficiency through targeted parameter extraction.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3044786B1Loudness adjustment for downmixed audio content
Publication Date: 2024.04.24 DOLBY LABORATORIES LICENSING CORP
  • EP3044786B1 patent drawingFigure 1A
  • EP3044786B1 patent drawingFigure 1B
  • EP3044786B1 patent drawingFigure 2A

AI summary

Audio content coded for a reference speaker configuration is downmixed to downmix audio content coded for a specific speaker configuration. One or more gain adjustments are performed on individual portions of the downmix audio content coded for the specific speaker configuration. Loudness measurements are then performed on the individual portions of the downmix audio content. An audio signal that comprises the audio content coded for the reference speaker configuration and downmix loudness metadata is generated. The downmix loudness metadata is created based at least in part on the loudness measurements on the individual portions of the downmix audio content.