Content-Aware Audio Leveling for Noise Pumping Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional dynamic processing techniques for audio, such as AGC and DRC, struggle to effectively manage audio levels in diverse content, often boosting low-level unwanted noise and introducing noise pumping, especially in user-generated content recorded in poor environments.
Innovation Solution
The method involves performing content-aware audio processing by source separating the audio signal into voice-related and residual components, determining a dynamic audio gain based on these components, and applying this gain for audio level adjustment, thereby improving audio quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional AGC and DRC techniques multiply the whole audio mixture by a time-varying gain, then the overall audio level is adjusted, but low-level unwanted noise is boosted and noise pumping is introduced
Solution Approach 1:
The patent segments the audio signal into multiple audio objects (speech, music, sound effects, noise) and applies different processing to each object. This allows the system to adjust the gain for speech and music separately from noise, preventing noise boosting while maintaining proper audio level control.
Solution Approach 2:
The patent applies different quality characteristics to different audio objects by adjusting their respective gain parameters independently. Speech objects receive different gain treatment compared to music objects, and noise objects can be suppressed separately, creating locally optimized audio quality for each component.
2Stability of the object's composition
If conventional AGC adjusts the long-term level to target level with slow gain changes, then the average loudness is consistent, but short-term level fluctuates significantly
Solution Approach 1:
The patent implements dynamic gain adjustment for each audio object based on its instantaneous level and importance. The system can rapidly adjust the gain of prominent speech or music objects in real-time, providing both long-term stability through overall level control and short-term responsiveness through object-specific dynamic adjustment.
3Speed
If conventional DRC adjusts the short-term level and limits fluctuations, then the dynamic range is compressed, but soft sounds are mapped to higher levels and loud sounds to lower values
Solution Approach 1:
The patent applies different compression characteristics to different audio objects. Music objects can undergo DRC processing to limit fluctuations, while speech objects maintain their dynamic range. The system preserves the natural dynamic range of each object type while still providing short-term level control where appropriate.
4Ease of operation
If volume leveler processes both PGC and UGC with the same technique, then the system is simple to operate, but it cannot effectively handle the diverse characteristics of user-generated content
Solution Approach 1:
The patent dynamically adapts its processing strategy based on the detected content type. The system automatically identifies whether the input is PGC or UGC and adjusts the audio object processing parameters accordingly, providing specialized handling for each content type while maintaining a unified user interface and operation flow.
Data Source
AI summary
Described herein is a method of performing content-aware audio processing for an audio signal comprising a plurality of audio components of different types. The method includes source separating the audio signal into at least a voice-related audio component and a residual audio component. The method further includes determining a dynamic audio gain based on the voice-related audio component and the residual audio component. The method also includes performing audio level adjustment for the audio signal based on the determined audio gain. Further described are corresponding apparatus, programs, and computer-readable storage media.

