Loudness Control with Noise Classification and Drop Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing loudness control systems fail to distinguish between desired content and unwanted noise, leading to undesirable amplification of background noise and fluctuations in signal-to-noise ratio, and struggle to maintain uniform average loudness levels without adversely affecting short-term signal dynamics.
Innovation Solution
A loudness control system that includes modules for content-versus-noise classification, temporal smoothing, and gain correction, using frequency or time domain noise detection to adjust smoothing factors and apply time-varying gains, while also detecting loudness drops to minimize artifacts during transitions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the loudness control system amplifies all low-level audio content above a predetermined threshold, then the average loudness level is normalized, but the background noise is also amplified reducing the signal-to-noise ratio
Solution Approach 1:
The audio signal is segmented into speech segments and non-speech segments using a speech detector. The loudness control system applies different processing to these segments: amplifying speech content while preventing amplification of background noise in non-speech segments, thereby normalizing loudness without degrading signal-to-noise ratio
Solution Approach 2:
The system dynamically adjusts the gain application based on real-time speech detection. When speech is detected, normal loudness control amplification is applied; when no speech is detected, amplification is prevented or reduced. This dynamic switching allows the system to maintain loudness normalization for desired content while avoiding noise amplification
2Measurement precision
If the loudness control system reacts quickly to loudness changes, then the target loudness level is consistently achieved, but the short-term signal dynamics are reduced
Solution Approach 1:
The system employs dynamic smoothing factors that adapt to the temporal characteristics of the input signal. For transient signals with rapid loudness changes, the smoothing factor is reduced to preserve short-term dynamics. For steady-state signals, the smoothing factor is increased to achieve consistent target loudness level. This dynamic adjustment resolves the contradiction between response speed and dynamics preservation
Solution Approach 2:
The smoothing factor parameter is changed based on the temporal derivative of the loudness signal. When the rate of change exceeds a threshold indicating transient content, the smoothing factor is automatically reduced. This parameter adaptation allows the system to maintain precision for steady signals while preserving dynamics for transient signals
3Stability of the object's composition
If the loudness control system uses a slow response to loudness changes, then the short-term signal dynamics are preserved, but the loudness level control becomes ineffective and artifacts appear
Solution Approach 1:
The system dynamically adjusts the smoothing factor based on the temporal characteristics of the input signal. For transient signals where dynamics preservation is critical, the smoothing factor is reduced to maintain responsiveness. For steady-state signals where dynamics are less important, the smoothing factor is increased to achieve precise loudness control. This resolves the contradiction by adapting the response speed to the signal content
4Device complexity
If the loudness control system applies uniform processing to all audio content, then the implementation is simple, but the system cannot distinguish between desired content and unwanted noise
Solution Approach 1:
The audio signal is segmented into speech and non-speech portions using a speech detector. Different processing rules are applied to each segment type: speech segments receive normal loudness control processing while non-speech segments have amplification prevented or reduced. This segmentation enables the system to distinguish desired content from noise without requiring complex analysis
Solution Approach 2:
A speech detector acts as an intermediary component that classifies audio segments before they reach the loudness control processor. This intermediary provides the discrimination capability needed to differentiate desired content from noise, while the main loudness control system can remain relatively simple in its processing logic
Data Source
AI summary
Loudness control systems or methods may normalize audio signals to a predetermined loudness level. If the audio signal includes moderate background noise, then the background noise may also be normalized to the target loudness level. Noise signals may be detected using content-versus-noise classification, and a loudness control system or method may be adjusted based on the detection of noise. Noise signals may be detected by signal analysis in the frequency domain or in the time domain. Loudness control systems may also produce undesirable audio effects when content shifts from a high overall loudness level to a lower overall loudness level. Such loudness drops may be detected, and the loudness control system may be adjusted to minimize the undesirable effects during the transition between loudness levels.


