Adaptive Audio Gain Smoothing for Pumping Effect Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio ducking technologies struggle to achieve aesthetically pleasing background audio gain smoothing in real-time scenarios due to difficulties in balancing foreground and background signal attenuation, leading to unpleasant fluctuations and reduced intelligibility of speech.

Innovation Solution

An apparatus and method for adaptive background audio gain smoothing, which determines a sequence of output gains based on signal characteristics, including foreground and background signal levels, using a signal characteristics provider and a gain sequence generator to gradually change current gain values to target values, ensuring smooth transitions and appropriate attenuation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If fast time-constants are used for real-time ducking processing, then the system can react quickly to signal changes, but it causes rapid fluctuation of background level (pumping) which is unpleasant

Engineering Contradiction:
Improvereaction speed to signal changesVSAvoidpumping effect
Core Design Contradiction:
SpeedVSObject-affected harmful factors

Solution Approach 1:

The patent applies dynamics by making the time-constants adaptive rather than fixed. The system dynamically adjusts attack and release time-constants based on the detected signal state and characteristics, allowing fast reaction when needed while preventing pumping effects during stable periods. This is achieved through state machine logic that modifies processing parameters in real-time based on input signal analysis.

Inventive Principle:
Principle #15Dynamics

2Object-affected harmful factors

If slow time-constants are used for stationary signals, then the behavior is smooth and pleasant, but the system cannot react quickly to significant signal changes

Engineering Contradiction:
Improvesmooth behaviorVSAvoidreaction speed to signal changes
Core Design Contradiction:
Object-affected harmful factorsVSSpeed

Solution Approach 1:

The system dynamically switches between slow and fast time-constants based on the current signal state. During stationary periods, slow time-constants provide smooth behavior. When significant signal changes are detected (e.g., new speech onset, loud background events), the system transitions to fast time-constants for rapid response. This state-dependent adaptation resolves the contradiction between smoothness and responsiveness.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter values of time-constants based on signal characteristics and system state. By modifying attack and release time-constant parameters dynamically, the system achieves both smooth behavior for stationary signals and fast reaction to changes, depending on the operational context.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If static attenuation is applied to background signal, then foreground speech intelligibility is ensured, but unnecessary attenuation damages the esthetic quality of background signal

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidesthetic quality damage
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The system applies local quality by differentiating between different segments and characteristics of the background signal. Instead of uniform static attenuation, the system applies selective attenuation based on local signal properties such as speech presence detection, signal type classification (music, noise, speech), and temporal context. This ensures speech intelligibility is protected while preserving esthetic quality of non-speech background portions.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The attenuation applied to the background signal is dynamic rather than static. The system continuously analyzes the background signal characteristics and adjusts the attenuation level accordingly. When speech is detected or signal characteristics indicate potential quality degradation, attenuation is applied or increased. When no speech is present and signal quality is good, attenuation is reduced or removed, preserving esthetic quality.

Inventive Principle:
Principle #15Dynamics

4Object-affected harmful factors

If release is delayed for subsequent ducking events, then listening experience is improved, but foreground speech intelligibility may be compromised

Engineering Contradiction:
Improvelistening experienceVSAvoidspeech intelligibility
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The release timing is made dynamic and adaptive rather than fixed. The system monitors for subsequent ducking events and adjusts release behavior accordingly. When a new ducking event is detected within a certain time window, the release is delayed to maintain smooth transitions and improve listening experience. However, this adaptive behavior continuously monitors speech intelligibility conditions and can override or adjust the delayed release if speech clarity is compromised, balancing both requirements dynamically.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP4305623B1Apparatus and method for adaptive background audio gain smoothing
Publication Date: 2024.12.11 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP4305623B1 patent drawingFigure 1
  • EP4305623B1 patent drawingFigure 2
  • EP4305623B1 patent drawingFigure 3

AI summary

An apparatus (100) for providing a sequence of output gains, wherein the sequence of output gains is suitable for attenuating a background signal of an audio signal is provided. The apparatus (100) comprises a signal characteristics provider (110) configured to receive or to determine signal characteristics information on one or more characteristics of the audio signal, wherein the signal characteristics information depends on the background signal, wherein the signal characteristics information comprises a sequence of input gains which depends on the background signal and on a foreground signal of the audio signal. Moreover, the apparatus (100) comprises a gain sequence generator (120) configured to determine the sequence of output gains depending on the sequence of input gains. To determine the sequence of outputs gains, to modify a current gain value of a current gain of the sequence of output gains to a target gain value, the gain sequence generator (120) is configured to determine a plurality of succeeding gains, which succeed the current gain in the sequence of output gains, by gradually changing the current gain value according to a modification rule during a transition period to the target gain value. The modification rule depends on the signal characteristics information; and/or the gain sequence generator (120) is configured to determine the target gain value depending on a further one of the one or more signal characteristics in addition to the sequence of input gains.