Speech Loudness Leveling Using Frame-Wise Gain Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for leveling speech content are either time-consuming and labor-intensive, such as manual gain adjustments, or result in quality degradation when using dynamic compressors, especially when significant gain reduction is required, failing to ensure consistent loudness across varying input ranges.
Innovation Solution
A method and system for automatic loudness leveling using a time-varying gain, incorporating a Voice Activity Detector, optional signal-to-noise ratio estimator, Speaker Diarization, and a denoising module, which analyzes loudness on both short-term and long-term scales to adjust gains and maintain a target loudness range without degrading quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual gain adjustments are used for leveling speech content, then the quality of leveling is improved, but the time consumption and labor intensity increase significantly
Solution Approach 1:
The system automatically detects speech presence, estimates loudness, and applies gain adjustments without human intervention. The automated leveling system serves itself by using algorithms to analyze audio characteristics and make real-time gain decisions, eliminating the need for manual engineer intervention while maintaining professional quality standards
Solution Approach 2:
The patent replaces manual mechanical gain adjustment with an automated digital signal processing system. The mechanical/manual fader operation is substituted by electronic algorithms including Voice Activity Detection, loudness estimation, and automatic gain calculation, achieving both speed and quality
2Productivity
If dynamic compressors are used for automatic leveling, then the time consumption is reduced, but the quality of speech content degrades when significant gain reduction is required
Solution Approach 1:
The system dynamically changes gain parameters based on real-time speech detection and loudness estimation. Instead of fixed compressor thresholds, the gain adjustment parameters are continuously adapted to match the speech characteristics and desired loudness range, preventing quality degradation while maintaining processing speed
Solution Approach 2:
The automated system incorporates feedback loops where the estimated loudness and speech presence detection continuously inform the gain adjustment decisions. This closed-loop control ensures that gain reduction is applied intelligently only when necessary, preserving speech quality while achieving automatic leveling
3Manufacturing precision
If manual gain adjustments are used, then the leveling quality is improved, but the adaptability to different loudness ranges is not ensured
Solution Approach 1:
The system dynamically adapts to different input loudness ranges by continuously estimating the actual loudness level and adjusting the target gain accordingly. The automated nature allows real-time adaptation to varying speech characteristics, speaker volumes, and recording conditions, ensuring consistent output quality across diverse input scenarios
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
Embodiments are disclosed for automatic leveling of speech content. In an embodiment, a method comprises: receiving, using one or more processors, frames of an audio recording including speech and non-speech content; for each frame: determining, using the one or more processors, a speech probability; analyzing, using the one or more processors, a perceptual loudness of the frame; obtaining, using the one or more processors, a target loudness range for the frame; computing, using the one or more processors, gains to apply to the frame based on the target loudness range and the perceptual loudness analysis, where the gains include dynamic gains that change frame-by-frame and that are scaled based on the speech probability; and applying the gains to the frame so that a resulting loudness range of the speech content in the audio recording fits within the target loudness range.