Audio Classification Gain Control for Consistent Volume

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio media systems face challenges in maintaining consistent volume levels across different audio sources, leading to noticeable and irritating volume fluctuations, as dynamic range compression often degrades audio quality by altering the original dynamics of the audio signal.

Innovation Solution

The system employs audio classification to determine the category of the audio signal and applies a targeted gain value, minimizing the need for dynamic range compression by using a combination of classification and real-time input measurements to adjust the volume, thereby reducing the amount of compression required to bring the audio signal within a target volume range.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Stability of the object's composition

If dynamic range compression is used to adjust volume levels, then volume consistency across different audio sources is improved, but audio quality deteriorates due to alteration of original dynamics

Engineering Contradiction:
Improvevolume consistencyVSAvoidaudio quality
Core Design Contradiction:
Stability of the object's compositionVSManufacturing precision

Solution Approach 1:

The system performs preliminary audio classification before volume adjustment to identify the type of audio content (music, speech, effects). This preliminary action enables the selection of appropriate processing parameters that preserve original dynamics while achieving volume consistency, thereby resolving the contradiction between volume stability and audio quality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies different processing strategies based on the classified audio type. For example, music audio receives different treatment compared to speech or sound effects. This localized approach ensures that each audio type is processed in a way that maintains its characteristic dynamics while achieving consistent volume levels across different sources

Inventive Principle:
Principle #3Local quality

2Manufacturing precision

If classification-based gain adjustment is used, then audio quality is preserved by minimizing compression, but system complexity increases due to additional classification and measurement components

Engineering Contradiction:
Improveaudio qualityVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The audio classification engine serves multiple functions: it identifies audio types for gain adjustment selection and provides information for optimizing compression parameters. This multi-functionality reduces the need for separate processing paths, thereby managing system complexity while maintaining audio quality through classification-based gain adjustment

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses real-time input measurements taken from the audio signal itself to determine appropriate gain values, rather than relying entirely on pre-defined classification rules. This self-service approach allows the system to adapt to actual signal characteristics, preserving audio quality while keeping the processing logic relatively simple

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240354053A1Methods and Apparatus for Dynamic Volume Adjustment Via Audio Classification
Publication Date: 2024.10.24 GRACENOTE INC
  • US20240354053A1 patent drawing
  • US20240354053A1 patent drawing
  • US20240354053A1 patent drawing

AI summary

Methods, apparatus, systems and articles of manufacture are disclosed for dynamic volume adjustment via audio classification. Example apparatus include at least one memory; instructions; and at least one processor to execute the instructions to: analyze, with a neural network, a parameter of an audio signal associated with a first volume level to determine a classification group associated with the audio signal; determine an input volume of the audio signal; determine a classification gain value based on the classification group; determine an intermediate gain value as an intermediate between the input volume and the classification gain value by applying a first weight to the input volume and a second weight to the classification gain value; apply the intermediate gain value to the audio signal, the intermediate gain value to modify the first volume level to a second volume level; and apply a compression value to the audio signal, the compression value to modify the second volume level to a third volume level that satisfies a target volume threshold.