Audio Classification Gain Blending for Stable Volume Levels
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio media systems face issues with varying volume levels across different sources and media types, leading to irritating changes in audio perception due to continuous dynamic range compression, which degrades the audio quality.
Innovation Solution
The implementation of audio classification techniques to determine the category of an audio signal and apply a targeted gain value, minimizing the need for dynamic range compression by adjusting the volume based on classification groups and real-time input measurements, thereby maintaining a consistent volume range.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Stability of the object's composition
If dynamic range compression is continuously applied to normalize volume levels across different audio sources, then volume consistency is improved, but audio quality deteriorates due to perceptible changes and artifacts
Solution Approach 1:
The system performs preliminary classification of the audio signal into categories (e.g., music, speech, nature sounds) before applying volume adjustment. Based on the classification, appropriate gain values are pre-determined and applied proactively, preventing the need for continuous dynamic range compression and thereby preserving audio quality while achieving volume consistency
Solution Approach 2:
The system changes the gain parameter based on the classified audio category rather than continuously adjusting compression parameters. By selecting from pre-determined gain values associated with different audio categories, the system achieves volume normalization with minimal compression, thus maintaining audio fidelity
2Object-generated harmful factors
If classification-based gain adjustment is applied, then audio quality is preserved by minimizing compression, but system complexity increases due to audio classification requirements
Solution Approach 1:
The audio signal space is segmented into distinct categories (music, speech, nature sounds, etc.), each with pre-determined gain values. This segmentation allows the system to handle diverse audio types through a manageable set of classification groups rather than requiring continuous analysis and adjustment, thus balancing accuracy with computational efficiency
Solution Approach 2:
Gain values for different audio categories are pre-determined and stored in association with each classification group. This preliminary preparation eliminates the need for real-time gain calculation, reducing computational complexity while maintaining audio quality through classification-based selection
Data Source
AI summary
Methods, apparatus, systems and articles of manufacture are disclosed for dynamic volume adjustment via audio classification. Example apparatus include at least one memory; instructions; and at least one processor to execute the instructions to: analyze, with a neural network, a parameter of an audio signal associated with a first volume level to determine a classification group associated with the audio signal; determine an input volume of the audio signal; determine a classification gain value based on the classification group; determine an intermediate gain value as an intermediate between the input volume and the classification gain value by applying a first weight to the input volume and a second weight to the classification gain value; apply the intermediate gain value to the audio signal, the intermediate gain value to modify the first volume level to a second volume level; and apply a compression value to the audio signal, the compression value to modify the second volume level to a third volume level that satisfies a target volume threshold.


