Audio Classification Gain Blending for Low-Artifact Volume Normalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio media systems face issues with varying volume levels from different sources, leading to irritating differences in sound quality due to continuous dynamic range compression, which degrades the perceived audio quality by introducing artifacts.
Innovation Solution
The implementation of audio classification to determine the category of an audio signal and apply a targeted gain value, minimizing the need for dynamic range compression, by using a combination of classification and real-time input measurements to adjust the volume to a consistent target range, while also considering initial volume adjustments and source changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Stability of the object's composition
If dynamic range compression is continuously applied to normalize volume levels from different audio sources, then volume consistency is improved, but audio quality deteriorates due to introduced artifacts
Solution Approach 1:
The system performs preliminary classification of the audio signal to determine its category (e.g., music genre, speech, type of content) before applying volume adjustment. This preliminary action allows the system to select an appropriate targeted gain value specific to the classified category, thereby minimizing the need for aggressive dynamic range compression and reducing audio quality degradation while still achieving volume normalization.
Solution Approach 2:
The system changes the parameter of gain adjustment based on the classified audio category. Instead of applying a fixed compression ratio, the system applies different targeted gain values depending on the audio signal classification (e.g., different gain values for music, speech, podcasts). This parameter change allows volume normalization with minimal compression, preserving audio quality while achieving consistency across different sources.
2Device complexity
If a fixed compression ratio is applied to all audio sources, then volume normalization is simplified, but audio quality across different genres deteriorates
Solution Approach 1:
The system performs preliminary classification of the audio signal to determine its category (e.g., music genre, speech, type of content) before applying volume adjustment. This preliminary action allows the system to select an appropriate targeted gain value specific to the classified category, thereby minimizing the need for aggressive dynamic range compression and reducing audio quality degradation while still achieving volume normalization.
Solution Approach 2:
The system changes the parameter of gain adjustment based on the classified audio category. Instead of applying a fixed compression ratio, the system applies different targeted gain values depending on the audio signal classification (e.g., different gain values for music, speech, podcasts). This parameter change allows volume normalization with minimal compression, preserving audio quality while achieving consistency across different sources.
3Stability of the object's composition
If aggressive compression is used to achieve target volume threshold, then volume consistency is improved, but perceptible artifacts are introduced
Solution Approach 1:
The system performs preliminary classification of the audio signal to determine its category (e.g., music genre, speech, type of content) before applying volume adjustment. This preliminary action allows the system to select an appropriate targeted gain value specific to the classified category, thereby minimizing the need for aggressive dynamic range compression and reducing audio quality degradation while still achieving volume normalization.
Solution Approach 2:
The system changes the parameter of gain adjustment based on the classified audio category. Instead of applying a fixed compression ratio, the system applies different targeted gain values depending on the audio signal classification (e.g., different gain values for music, speech, podcasts). This parameter change allows volume normalization with minimal compression, preserving audio quality while achieving consistency across different sources.
Data Source
AI summary
Methods, apparatus, systems and articles of manufacture are disclosed for dynamic volume adjustment via audio classification. Example apparatus include at least one memory; instructions; and at least one processor to execute the instructions to: analyze, with a neural network, a parameter of an audio signal associated with a first volume level to determine a classification group associated with the audio signal; determine an input volume of the audio signal; determine a classification gain value based on the classification group; determine an intermediate gain value as an intermediate between the input volume and the classification gain value by applying a first weight to the input volume and a second weight to the classification gain value; apply the intermediate gain value to the audio signal, the intermediate gain value to modify the first volume level to a second volume level; and apply a compression value to the audio signal, the compression value to modify the second volume level to a third volume level that satisfies a target volume threshold.


