Adaptive Loudness Levelling with Temporal Masking Gain Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing technologies fail to maintain consistent perceived loudness across varying audio sources, leading to frequent manual volume adjustments when switching between different channels or content types, resulting in an unpleasant listening experience due to signal distortion and unrealistic sound production.
Innovation Solution
A method and apparatus that analyze audio signals by segmenting them into frames, classifying signal states, detecting important events, and applying appropriate gains with smoothing to maintain consistent volume levels, utilizing a frequency weighting block and loudness engine that models human listening sensitivity and temporal masking effects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Stability of the object's composition
If a Dynamic Range Compressor (DRC) is used to suppress loud signals and boost soft ones, then volume consistency is improved, but signal distortion occurs and realistic sound quality deteriorates
Solution Approach 1:
The audio signal is divided into multiple frequency sub-bands, and a separate DRC is applied to each sub-band independently. This segmentation allows different compression parameters to be applied to different frequency ranges, maintaining volume consistency while preserving the natural characteristics of each frequency band, thereby reducing overall signal distortion
Solution Approach 2:
Different compression characteristics and parameters are applied to different sub-bands based on their specific properties. Each sub-band can have customized attack and release times, threshold levels, and ratio settings, allowing the system to maintain realistic sound quality while achieving volume consistency across the entire audio spectrum
2Object-generated harmful factors
If DRC is applied to each sub-band independently with customized parameters, then signal quality is improved, but device complexity increases
Solution Approach 1:
The system uses a universal processing framework that handles multiple sub-bands through a single integrated architecture. The same basic DRC algorithm and processing steps are applied across all sub-bands, with parameters dynamically adjusted based on input characteristics. This multi-functional approach achieves sub-band independent processing while maintaining relatively simple device structure
Solution Approach 2:
The system dynamically adjusts compression parameters based on the actual audio content and perceived loudness in each sub-band. Rather than using fixed complex configurations, the parameters adapt in real-time to signal conditions, simplifying the device structure while maintaining high signal quality through context-aware processing
3Device complexity
If traditional DRC responds to signal level according to predefined curve, then processing simplicity is maintained, but attack portion is attenuated and release portion is amplified causing unrealistic sounds
Solution Approach 1:
The system performs preliminary analysis of the audio signal to identify attack and release portions before applying compression. By detecting transient events and signal envelope characteristics in advance, the system can apply appropriate compression only where needed, preserving the natural dynamics of attack portions while still controlling overall loudness, thereby avoiding unrealistic sound artifacts
Data Source
AI summary
A time-domain method of adaptively levelling the loudness of a digital audio signal is proposed. It selects a proper frequency weighting curve to relate the volume level to the human auditory system. The audio signal is segmented into frames of a suitable duration for content analysis. Each frame is classified to one of several predefined states and events of perceptual interest is detected. Four quantities are updated each frame according to the classified state and detected event to keep track of the signal. One quantity measures the long-term loudness and is the main criterion for state classification of a frame. The second quantity is the short-term loudness that is mainly used for deriving the target gain. The third quantity measures the low-level loudness when the signal is deemed to not contain important content, giving a reasonable estimate of noise floor. A fourth quantity measures the peak loudness level that is used to simulate the temporal masking effect. The target gain to maintain the audio signal to the desired loudness level is calculated by a volume leveller, regulated by a gain controller that simulates the temporal masking effect to get rid of unnecessary gain fluctuations, ensuring a pleasant sound.


