Audio Ducking via Real-Time Signal Metrics and Key Frame Automation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multimedia editing applications face challenges in accurately and efficiently performing audio ducking, as manual techniques are tedious and prone to inaccuracy, while automated methods can be inaccurate and computationally inefficient, especially when handling changes to audio signals.

Innovation Solution

A computer system generates metrics for audio slices of a foreground audio signal, using a summed-area table to compute total metrics efficiently, allowing for accurate and fast addition and updating of key frames to control a background audio signal, thereby automating audio ducking with improved precision and reduced computational overhead.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual key frame techniques are used, then the user can control audio ducking, but the process is tedious and time consuming

Engineering Contradiction:
Improveease of audio ducking controlVSAvoidtime for placing key frames
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system automatically analyzes the foreground audio signal and places key frames without user intervention. The computer system computes audio metrics, detects dialogue regions, and generates key frames autonomously, eliminating the manual operation burden while maintaining accurate audio ducking control.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The manual mechanical process of visually inspecting waveforms and manually placing key frames is replaced with an automated computational system. The system uses audio metric computation and algorithms to automatically determine key frame locations, substituting human manual operations with automated signal processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated key frame techniques are used, then the workflow is more efficient, but the accuracy of audio ducking decreases

Engineering Contradiction:
Improveworkflow efficiencyVSAvoidaccuracy of audio ducking
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system computes audio metrics continuously and uses this feedback to dynamically adjust key frame placement. By monitoring audio properties such as RMS levels and detecting dialogue regions through metric analysis, the system accurately identifies when background audio should be ducked, ensuring both efficiency and precision.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system analyzes multiple audio parameters including RMS levels, peak detection, and temporal patterns to determine key frame locations. By changing and comparing these parameters across different time windows, the system accurately identifies dialogue regions and places key frames at precise locations for optimal audio ducking.

Inventive Principle:
Principle #35Parameter changes

3Extent of automation

If side chaining is used for automated ducking, then the process is automated, but the configuration is complex and requires delay introduction

Engineering Contradiction:
Improveautomation of audio duckingVSAvoidcomplexity of system configuration
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The system extracts the essential function of side chaining (using foreground audio to control background audio levels) but removes the complex routing configuration and delay requirements. By directly computing audio metrics from the foreground signal and using these metrics to control background audio, the system achieves automation without the complexity of traditional side chaining architecture.

Inventive Principle:
Principle #2Taking out (Extraction)

4Ease of manufacture

If existing automated techniques add key frames at audio clip boundaries, then the process is simple, but the audio ducking is inaccurate

Engineering Contradiction:
Improvesimplicity of key frame generationVSAvoidaccuracy of key frame placement
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The system segments the audio signal into analysis windows and computes metrics for each segment. By dividing the foreground audio into overlapping windows and analyzing audio properties within each segment, the system can precisely locate dialogue regions and place key frames at accurate positions rather than relying on arbitrary clip boundaries.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11327710B2Automatic audio ducking with real time feedback based on fast integration of signal levels
Publication Date: 2022.05.10 ADOBE INC
  • US11327710B2 patent drawing
  • US11327710B2 patent drawing
  • US11327710B2 patent drawing

AI summary

A computer-implemented method for audio signal processing includes analyzing a foreground audio signal to determine metrics corresponding to audio slices of the foreground audio signal. Each such metric indicates a value for an audio property of a respective audio slice. The method further includes computing a total metric for an audio slice as a function of a set of the metrics corresponding to a set of the audio slices including the audio slice. The method further includes adding a key frame to a track based on the total metric. The track includes the foreground audio signal and a background audio signal, and a location of the key frame corresponds to a location of the audio slice on the track. The key frame indicates a change to the audio property of the background audio signal at the location on the track, and the key frame is utilizable for audio ducking.