Object-Based Audio Balancing for Dialog Loudness Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio technologies fail to effectively adjust the balance between dialog and non-dialog signals in audio programs, leading to user dissatisfaction due to inconsistent listening comfort and the need for frequent volume adjustments, especially for individuals with hearing impairments or those listening in adverse conditions.
Innovation Solution
An object-based digital audio processing system that uses separate dialog and non-dialog object signals, applying gain or attenuation to achieve a user-defined dialog-to-non-dialog balance, with long-term and short-term correction mechanisms to maintain optimal listening comfort across different audio genres and environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If dialog signal level is increased to improve listening comfort for users with hearing impairments or in adverse listening conditions, then dialog prominence is improved, but overall audio balance and listening comfort for other users deteriorates
Solution Approach 1:
The audio signal is segmented into dialog components and non-dialog components using dialog detection techniques. This allows independent processing of dialog signals to enhance prominence for users with hearing impairments or in adverse listening conditions, while maintaining the original balance for other users through separate gain control of each component.
Solution Approach 2:
The system dynamically adjusts dialog signal levels based on detected dialog presence and user preferences. The gain control is not static but adapts in real-time to maintain optimal dialog prominence while preserving overall audio balance, resolving the contradiction between enhanced dialog and overall balance.
2Stability of the object's composition
If a fixed VRA ratio is locked-in and maintained for the rest of the program, then production mix consistency is improved, but adaptability to different user preferences and listening conditions deteriorates
Solution Approach 1:
The system replaces fixed VRA ratio with dynamic adjustment mechanisms that respond to user preferences and detected dialog conditions. The dialog gain control adapts throughout the program while maintaining production mix consistency through controlled adjustments based on user-specific parameters and real-time dialog detection.
Solution Approach 2:
The system changes the VRA ratio parameter dynamically based on user preferences and listening conditions rather than maintaining a fixed value. This allows adaptation to different user needs (e.g., users with hearing impairments) while preserving the core production mix integrity through controlled parameter modification.
3Ease of operation
If dialog detection techniques are used to selectively process dialog components, then dialog prominence control is improved, but device complexity increases
Solution Approach 1:
The system uses universal dialog detection techniques that can be integrated into existing audio processing pipelines. The same detection mechanism serves multiple functions: identifying dialog for enhancement, maintaining audio balance, and adapting to user preferences, thereby reducing overall system complexity despite the added capability.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Systems, devices, and methods are described herein for adjusting a relationship between dialog and non-dialog signals in an audio program. In an example, information about a long-term dialog balance for an audio program can be received. The long-term loudness dialog balance can indicate a dialog-to-non-dialog loudness relationship of the audio program. A dialog loudness preference can be received, such as from a user, from a database, or from another source. A desired long-term gain or attenuation can be determined according to a difference between the received long-term dialog balance for the audio program and the received dialog balance preference. The long-term gain or attenuation can be applied to at least one of the dialog signal and the non-dialog signal of the audio program to render an audio program that is enhanced according to the loudness preference.