Audio Mix Gain Optimization for Multi-Target Loudness Compliance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing technologies struggle to simultaneously meet multiple audio targets, such as dialogue and program loudness, without disrupting the balance and artistic intent of the original soundtrack, especially in the context of media distribution where compliance with various standards is required.
Innovation Solution
A system that optimizes audio content by computing audio feature measurement values and iteratively adjusting gain values to meet multiple energy-based targets, preserving the relative energy balance between audio components, using a loss function optimization in a down-sampled audio feature domain.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional audio processing methods are used to meet multiple audio targets, then compliance with audio standards is achieved, but processing time is excessively long and the balance of the original soundtrack is disrupted
Solution Approach 1:
The audio signal is divided into multiple audio frames, and gain values are optimized for each frame independently. This segmentation allows parallel processing of different time segments, significantly reducing overall processing time while maintaining compliance with audio standards for each segment.
Solution Approach 2:
The patent transforms the audio processing problem from the time domain to the frequency domain by computing audio feature measurement values (such as energy, spectral centroid, and zero-crossing rate). This parameter transformation enables more efficient optimization and reduces processing time while ensuring standard compliance.
2Reliability
If gain values are adjusted to meet energy-based targets, then compliance with loudness standards is achieved, but the relative energy balance between audio components is disrupted
Solution Approach 1:
The patent computes and optimizes gain values separately for different audio components (dialogue, music, sound effects) based on their specific energy measurements. This local optimization ensures that each component meets the required loudness standards while preserving the relative energy balance between components by considering their individual characteristics.
Solution Approach 2:
The system iteratively computes audio feature measurements, evaluates compliance with energy-based targets, and adjusts gain values accordingly. This feedback loop continues until both loudness standard compliance and energy balance preservation are achieved, ensuring that adjustments to one component do not adversely affect others.
3Reliability
If multiple audio targets are optimized simultaneously, then comprehensive compliance is achieved, but computational complexity increases
Solution Approach 1:
The optimization problem is segmented by dividing the audio into frames and processing each frame independently with pre-computed target gain values. This segmentation transforms a complex simultaneous optimization problem into multiple simpler sequential optimizations, reducing computational complexity while achieving comprehensive compliance.
Solution Approach 2:
The patent pre-computes target gain values based on audio feature measurements before applying them to the audio signal. This preliminary action simplifies the actual optimization process by having the difficult computational work done in advance, reducing real-time computational complexity while ensuring comprehensive compliance with multiple standards.
Data Source
AI summary
Some implementations of the disclosure relate to a non-transitory computer-readable medium having executable instructions stored thereon that, when executed by a processor, cause a system to perform operations comprising: obtaining a first energy-based target for audio; obtaining a first version of a sound mix including one or more audio components; computing, for each audio frame of multiple audio frames of each of the one or more audio components, a first audio feature measurement value; optimizing, based at least on the first energy-based target and the first audio feature measurement values, gain values of the audio frames; and after optimizing the gain values, applying the gain values to the first version of sound mix to obtain a second version of the sound mix.


