Automatic Audio Gain Calculation for Narration Mixing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face challenges in adding commentary to audio and video recordings captured on mobile devices, particularly in noisy environments, where ambient noise drowns out their voice, requiring separate devices or post-processing delays for mixing narration with pre-recorded content.

Innovation Solution

A method and apparatus for a mobile device to automatically mix first and second audio data associated with video, by calculating a loudness measure of the first audio data, attenuating it based on the loudness measure, and synchronizing the attenuated audio with the second audio data to form a new content item, allowing contemporaneous commentary recording.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If users record commentary contemporaneously with pre-recorded content, then the workflow efficiency is improved, but ambient noise drowns out the voiceover

Engineering Contradiction:
Improveworkflow efficiencyVSAvoidambient noise interference
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent converts the harmful ambient noise into a useful signal by using it as a reference for noise cancellation. The system captures ambient noise during pre-recording, then uses this captured noise profile to subtract ambient noise from the contemporaneous commentary recording, transforming the harmful interference into a benefit for improving voiceover clarity.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The patent performs preliminary noise profiling during the pre-recording phase. Before the actual commentary is recorded, the system captures and analyzes the ambient noise characteristics, storing this noise profile for later use. This preliminary action enables the system to preemptively prepare for noise cancellation during the commentary recording phase.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If users use separate devices for audio and video capture, then the audio quality is improved, but the device complexity increases

Engineering Contradiction:
Improveaudio qualityVSAvoidnumber of devices
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges audio and video capture capabilities into a single mobile device. By integrating the voiceover recording, pre-recorded content playback, and noise cancellation processing within one device, the system eliminates the need for separate audio recording equipment while maintaining acceptable audio quality through software-based noise management.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If users perform post-processing for audio mixing, then the audio mixing quality is improved, but the time delay increases

Engineering Contradiction:
Improveaudio mixing qualityVSAvoidpost-processing delay
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements automated audio mixing and noise cancellation processing that operates in real-time without requiring manual post-processing intervention. The system automatically analyzes the audio signals, identifies ambient noise patterns, performs noise subtraction, and mixes the commentary with pre-recorded content autonomously, eliminating the time delay associated with manual post-processing while maintaining professional-quality results.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10297269B2Automatic calculation of gains for mixing narration into pre-recorded content
Publication Date: 2019.05.21 DOLBY LABORATORIES LICENSING CORP
  • US10297269B2 patent drawing
  • US10297269B2 patent drawing
  • US10297269B2 patent drawing

AI summary

A system and method of mixing narration into content. The system automatically reduces the volume of the content according to a threshold value and a knee value. In this manner, the audio of the content does not overwhelm the narration.