Audio Classification Gain Blending for Low-Artifact Volume Normalization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio media systems face issues with varying volume levels from different sources, leading to irritating differences in sound quality due to continuous dynamic range compression, which degrades the perceived audio quality by introducing artifacts.

Innovation Solution

The implementation of audio classification to determine the category of an audio signal and apply a targeted gain value, minimizing the need for dynamic range compression, by using a combination of classification and real-time input measurements to adjust the volume to a consistent target range, while also considering initial volume adjustments and source changes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Stability of the object's composition

If dynamic range compression is continuously applied to normalize volume levels from different audio sources, then volume consistency is improved, but audio quality deteriorates due to introduced artifacts

Engineering Contradiction:
Improvevolume consistencyVSAvoidaudio quality degradation
Core Design Contradiction:
Stability of the object's compositionVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary classification of the audio signal to determine its category (e.g., music genre, speech, type of content) before applying volume adjustment. This preliminary action allows the system to select an appropriate targeted gain value specific to the classified category, thereby minimizing the need for aggressive dynamic range compression and reducing audio quality degradation while still achieving volume normalization.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes the parameter of gain adjustment based on the classified audio category. Instead of applying a fixed compression ratio, the system applies different targeted gain values depending on the audio signal classification (e.g., different gain values for music, speech, podcasts). This parameter change allows volume normalization with minimal compression, preserving audio quality while achieving consistency across different sources.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If a fixed compression ratio is applied to all audio sources, then volume normalization is simplified, but audio quality across different genres deteriorates

Engineering Contradiction:
Improvevolume normalization processVSAvoidaudio quality across genres
Core Design Contradiction:
Device complexityVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary classification of the audio signal to determine its category (e.g., music genre, speech, type of content) before applying volume adjustment. This preliminary action allows the system to select an appropriate targeted gain value specific to the classified category, thereby minimizing the need for aggressive dynamic range compression and reducing audio quality degradation while still achieving volume normalization.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes the parameter of gain adjustment based on the classified audio category. Instead of applying a fixed compression ratio, the system applies different targeted gain values depending on the audio signal classification (e.g., different gain values for music, speech, podcasts). This parameter change allows volume normalization with minimal compression, preserving audio quality while achieving consistency across different sources.

Inventive Principle:
Principle #35Parameter changes

3Stability of the object's composition

If aggressive compression is used to achieve target volume threshold, then volume consistency is improved, but perceptible artifacts are introduced

Engineering Contradiction:
Improvevolume consistencyVSAvoidperceptible artifacts
Core Design Contradiction:
Stability of the object's compositionVSObject-generated harmful factors

Solution Approach 1:

The system performs preliminary classification of the audio signal to determine its category (e.g., music genre, speech, type of content) before applying volume adjustment. This preliminary action allows the system to select an appropriate targeted gain value specific to the classified category, thereby minimizing the need for aggressive dynamic range compression and reducing audio quality degradation while still achieving volume normalization.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes the parameter of gain adjustment based on the classified audio category. Instead of applying a fixed compression ratio, the system applies different targeted gain values depending on the audio signal classification (e.g., different gain values for music, speech, podcasts). This parameter change allows volume normalization with minimal compression, preserving audio quality while achieving consistency across different sources.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12061840B2Methods and apparatus for dynamic volume adjustment via audio classification
Publication Date: 2024.08.13 GRACENOTE INC
  • US12061840B2 patent drawing
  • US12061840B2 patent drawing
  • US12061840B2 patent drawing

AI summary

Methods, apparatus, systems and articles of manufacture are disclosed for dynamic volume adjustment via audio classification. Example apparatus include at least one memory; instructions; and at least one processor to execute the instructions to: analyze, with a neural network, a parameter of an audio signal associated with a first volume level to determine a classification group associated with the audio signal; determine an input volume of the audio signal; determine a classification gain value based on the classification group; determine an intermediate gain value as an intermediate between the input volume and the classification gain value by applying a first weight to the input volume and a second weight to the classification gain value; apply the intermediate gain value to the audio signal, the intermediate gain value to modify the first volume level to a second volume level; and apply a compression value to the audio signal, the compression value to modify the second volume level to a third volume level that satisfies a target volume threshold.