Audio Classification Gain Blending for Stable Volume Levels

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio media systems face issues with varying volume levels across different sources and media types, leading to irritating changes in audio perception due to continuous dynamic range compression, which degrades the audio quality.

Innovation Solution

The implementation of audio classification techniques to determine the category of an audio signal and apply a targeted gain value, minimizing the need for dynamic range compression by adjusting the volume based on classification groups and real-time input measurements, thereby maintaining a consistent volume range.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Stability of the object's composition

If dynamic range compression is continuously applied to normalize volume levels across different audio sources, then volume consistency is improved, but audio quality deteriorates due to perceptible changes and artifacts

Engineering Contradiction:
Improvevolume consistencyVSAvoidaudio quality degradation
Core Design Contradiction:
Stability of the object's compositionVSObject-generated harmful factors

Solution Approach 1:

The system performs preliminary classification of the audio signal into categories (e.g., music, speech, nature sounds) before applying volume adjustment. Based on the classification, appropriate gain values are pre-determined and applied proactively, preventing the need for continuous dynamic range compression and thereby preserving audio quality while achieving volume consistency

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes the gain parameter based on the classified audio category rather than continuously adjusting compression parameters. By selecting from pre-determined gain values associated with different audio categories, the system achieves volume normalization with minimal compression, thus maintaining audio fidelity

Inventive Principle:
Principle #35Parameter changes

2Object-generated harmful factors

If classification-based gain adjustment is applied, then audio quality is preserved by minimizing compression, but system complexity increases due to audio classification requirements

Engineering Contradiction:
Improveaudio quality degradationVSAvoidsystem complexity
Core Design Contradiction:
Object-generated harmful factorsVSDevice complexity

Solution Approach 1:

The audio signal space is segmented into distinct categories (music, speech, nature sounds, etc.), each with pre-determined gain values. This segmentation allows the system to handle diverse audio types through a manageable set of classification groups rather than requiring continuous analysis and adjustment, thus balancing accuracy with computational efficiency

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Gain values for different audio categories are pre-determined and stored in association with each classification group. This preliminary preparation eliminates the need for real-time gain calculation, reducing computational complexity while maintaining audio quality through classification-based selection

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11775250B2Methods and apparatus for dynamic volume adjustment via audio classification
Publication Date: 2023.10.03 GRACENOTE INC
  • US11775250B2 patent drawing
  • US11775250B2 patent drawing
  • US11775250B2 patent drawing

AI summary

Methods, apparatus, systems and articles of manufacture are disclosed for dynamic volume adjustment via audio classification. Example apparatus include at least one memory; instructions; and at least one processor to execute the instructions to: analyze, with a neural network, a parameter of an audio signal associated with a first volume level to determine a classification group associated with the audio signal; determine an input volume of the audio signal; determine a classification gain value based on the classification group; determine an intermediate gain value as an intermediate between the input volume and the classification gain value by applying a first weight to the input volume and a second weight to the classification gain value; apply the intermediate gain value to the audio signal, the intermediate gain value to modify the first volume level to a second volume level; and apply a compression value to the audio signal, the compression value to modify the second volume level to a third volume level that satisfies a target volume threshold.