Speaker-Specific Volume Equalization for Consistent Dialogue Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Audio content often features varying audio levels between speakers and background noise, leading to discomfort for listeners who need to repeatedly adjust volume settings, which can affect the creative intent of the audio source.

Innovation Solution

Implementing a system for speaker-specific volume level equalization using machine learning to identify and adjust the volume of individual speakers and background noise, allowing users to set preferred volume levels through a remote control or voice commands, while maintaining the emotional context of the audio content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If volume level is flattened among different speakers, then listener comfort is improved, but creative intent of the audio source is adversely affected

Engineering Contradiction:
Improvelistener comfortVSAvoidcreative intent
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The system applies different volume adjustment strategies to different audio components (dialogue, music, sound effects) based on their specific characteristics and importance to creative intent. Dialogue receives aggressive normalization for comfort, while music and SFX preserve their original dynamic range to maintain artistic expression.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts volume parameters based on detected speaker characteristics and contextual information. It modifies gain levels, compression ratios, and limiting thresholds adaptively to balance listener comfort with preservation of creative intent across different audio segments.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If manual volume adjustments are allowed, then creative intent is preserved, but listener convenience deteriorates due to repeated adjustments

Engineering Contradiction:
Improvecreative intentVSAvoidtime for volume adjustments
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system pre-processes audio content by analyzing speaker characteristics, voice patterns, and contextual information before playback. It prepares normalized volume levels and pre-configured adjustment profiles, so that minimal real-time intervention is needed during actual listening.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically detects and compensates for volume inconsistencies between speakers without requiring manual user input. It self-adjusts gain levels, applies appropriate compression, and maintains creative intent preservation algorithms autonomously throughout playback.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If automatic volume normalization is applied, then listener comfort is improved, but system complexity increases

Engineering Contradiction:
Improvelistener comfortVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system divides audio processing into distinct modules: speaker detection, voice activity detection, volume analysis, normalization processing, and creative intent preservation. Each module handles a specific aspect independently, making the overall complex system manageable and maintainable through clear separation of concerns.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11848655B1Multi-channel volume level equalization based on user preferences
Publication Date: 2023.12.19 AMAZON TECH INC
  • US11848655B1 patent drawing
  • US11848655B1 patent drawing
  • US11848655B1 patent drawing

AI summary

Systems, devices, and methods are provided for multi-stem volume equalization, wherein the volume levels of each stem may be adjusted non-uniformly. Audio may be diarized into a plurality of stems, including background noise separate. Mean and variance of the volume levels of the stems may be computed. Each audio stem may be automatically adjusted based on a stem-specific preference that a user may specify. View may adjust actor volume relative to the mean/variance that maintains a relative difference in volume levels between stems.