Speech Loudness Leveling Using Frame-by-Frame VAD Gain Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech content editing and processing methods, such as manual gain adjustments and dynamic compressors, are time-consuming and can degrade quality, failing to ensure consistent loudness and professional quality in speech recordings.
Innovation Solution
An automated method using a Voice Activity Detector (VAD), optional signal-to-noise ratio (SNR) estimator, Speaker Diarization, and loudness analysis to apply time-varying gains, boosting soft sections and attenuating loud sections, ensuring the loudness range fits within a target range.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual gain adjustments are used for leveling speech content, then the best quality results are achieved, but the process is time consuming
Solution Approach 1:
The system automatically detects speech frames, analyzes their loudness, and applies appropriate gain adjustments without requiring manual intervention. The automated leveling system serves itself by making intelligent decisions about which frames need adjustment and by how much, based on speech probability detection and loudness analysis.
Solution Approach 2:
The system dynamically changes the gain parameter for each frame based on its speech probability and loudness characteristics. By adjusting the gain parameter automatically according to detected speech content and loudness levels, the system achieves manual-quality results without the time cost of manual adjustment.
2Productivity
If dynamic compressors are used for automatic leveling, then processing time is reduced, but quality degradation occurs when lots of gain reduction is required
Solution Approach 1:
Instead of applying uniform compression across all audio content, the system detects speech frames and applies gain adjustments specifically to those frames. By treating speech and non-speech content differently and applying localized gain control only where needed, the system maintains high quality while achieving automatic processing speed.
Solution Approach 2:
The system uses dynamic gain adjustment based on real-time speech probability and loudness analysis of each frame. Rather than static compression thresholds, the gain applied to each frame is dynamically determined by its speech characteristics and loudness level, preserving quality while maintaining processing efficiency.
3Stability of the object's composition
If manual leveling is applied, then loudness consistency is achieved, but the output audio does not fit into the desired loudness range for varying input audio
Solution Approach 1:
The system analyzes the loudness of each detected speech frame and uses this feedback to determine the appropriate gain adjustment. By continuously monitoring loudness and adjusting gain based on this feedback, the system achieves both loudness consistency and adaptability to fit the desired loudness range for varying input audio.
Solution Approach 2:
The system performs preliminary detection of speech frames and their loudness characteristics before applying gain adjustments. This preliminary analysis allows the system to plan and apply appropriate gain changes that will ensure the final output fits within the desired loudness range while maintaining consistency.
Data Source
AI summary
Embodiments are disclosed for automatic leveling of speech content. In an embodiment, a method comprises: receiving, using one or more processors, frames of an audio recording including speech and non-speech content; for each frame: determining, using the one or more processors, a speech probability; analyzing, using the one or more processors, a perceptual loudness of the frame; obtaining, using the one or more processors, a target loudness range for the frame; computing, using the one or more processors, gains to apply to the frame based on the target loudness range and the perceptual loudness analysis, where the gains include dynamic gains that change frame-by-frame and that are scaled based on the speech probability; and applying the gains to the frame so that a resulting loudness range of the speech content in the audio recording fits within the target loudness range.


