Speech Loudness Leveling Using Frame-by-Frame VAD Gain Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech content editing and processing methods, such as manual gain adjustments and dynamic compressors, are time-consuming and can degrade quality, failing to ensure consistent loudness and professional quality in speech recordings.

Innovation Solution

An automated method using a Voice Activity Detector (VAD), optional signal-to-noise ratio (SNR) estimator, Speaker Diarization, and loudness analysis to apply time-varying gains, boosting soft sections and attenuating loud sections, ensuring the loudness range fits within a target range.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual gain adjustments are used for leveling speech content, then the best quality results are achieved, but the process is time consuming

Engineering Contradiction:
Improveleveling qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system automatically detects speech frames, analyzes their loudness, and applies appropriate gain adjustments without requiring manual intervention. The automated leveling system serves itself by making intelligent decisions about which frames need adjustment and by how much, based on speech probability detection and loudness analysis.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system dynamically changes the gain parameter for each frame based on its speech probability and loudness characteristics. By adjusting the gain parameter automatically according to detected speech content and loudness levels, the system achieves manual-quality results without the time cost of manual adjustment.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If dynamic compressors are used for automatic leveling, then processing time is reduced, but quality degradation occurs when lots of gain reduction is required

Engineering Contradiction:
Improveprocessing speedVSAvoidoutput quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

Instead of applying uniform compression across all audio content, the system detects speech frames and applies gain adjustments specifically to those frames. By treating speech and non-speech content differently and applying localized gain control only where needed, the system maintains high quality while achieving automatic processing speed.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system uses dynamic gain adjustment based on real-time speech probability and loudness analysis of each frame. Rather than static compression thresholds, the gain applied to each frame is dynamically determined by its speech characteristics and loudness level, preserving quality while maintaining processing efficiency.

Inventive Principle:
Principle #15Dynamics

3Stability of the object's composition

If manual leveling is applied, then loudness consistency is achieved, but the output audio does not fit into the desired loudness range for varying input audio

Engineering Contradiction:
Improveloudness consistencyVSAvoidloudness range adaptation
Core Design Contradiction:
Stability of the object's compositionVSAdaptability or versatility

Solution Approach 1:

The system analyzes the loudness of each detected speech frame and uses this feedback to determine the appropriate gain adjustment. By continuously monitoring loudness and adjusting gain based on this feedback, the system achieves both loudness consistency and adaptability to fit the desired loudness range for varying input audio.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary detection of speech frames and their loudness characteristics before applying gain adjustments. This preliminary analysis allows the system to plan and apply appropriate gain changes that will ensure the final output fits within the desired loudness range while maintaining consistency.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12412595B2Automatic leveling of speech content
Publication Date: 2025.09.09 DOLBY LABORATORIES LICENSING CORP
  • US12412595B2 patent drawing
  • US12412595B2 patent drawing
  • US12412595B2 patent drawing

AI summary

Embodiments are disclosed for automatic leveling of speech content. In an embodiment, a method comprises: receiving, using one or more processors, frames of an audio recording including speech and non-speech content; for each frame: determining, using the one or more processors, a speech probability; analyzing, using the one or more processors, a perceptual loudness of the frame; obtaining, using the one or more processors, a target loudness range for the frame; computing, using the one or more processors, gains to apply to the frame based on the target loudness range and the perceptual loudness analysis, where the gains include dynamic gains that change frame-by-frame and that are scaled based on the speech probability; and applying the gains to the frame so that a resulting loudness range of the speech content in the audio recording fits within the target loudness range.