Compressed Audio Classification Using Side-Information Descriptors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for classifying compressed audio files are computationally complex and not directly applicable to generic internet applications, often requiring decoding back to the time domain and failing to utilize side-information descriptors, making them inefficient for rapid identification and comparison of compressed audio files.
Innovation Solution
The method involves dividing audio files into frames, compressing them using a psycho-acoustic algorithm, and generating parameters from frequency sub-bands, including average spectral power and side-information related to rhythm, which are used to classify and compare audio files by calculating differences with weighting factors to account for human auditory sensitivity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing classification methods for compressed audio files are used, then classification accuracy can be maintained, but computational complexity increases and processing time extends
Solution Approach 1:
The patent extracts and utilizes side-information descriptors that are already present in the compressed audio file format, rather than performing full decoding and re-analysis. This extraction approach maintains classification accuracy by using relevant features while significantly reducing computational complexity by working directly with the compressed representation.
Solution Approach 2:
The classification system leverages information that is already embedded in the compressed audio files during their encoding process. By using descriptors that were generated as part of the compression process itself, the system eliminates the need for separate, computationally intensive analysis stages, thereby reducing overall computational complexity while maintaining accuracy.
2Measurement precision
If existing classification methods for compressed audio files are used, then classification accuracy can be maintained, but processing speed decreases
Solution Approach 1:
The side-information descriptors are generated during the audio compression process itself, before the classification task begins. This preliminary generation of relevant features means that when classification is needed, the system can immediately use pre-computed descriptors without performing additional heavy computations, thereby maintaining accuracy while improving processing speed.
3Adaptability or versatility
If decoding back to time domain is performed for classification, then classification can be applied to compressed files, but computational efficiency is lost
Solution Approach 1:
The patent uses side-information descriptors as an intermediary representation that bridges the gap between compressed audio files and classification algorithms. These descriptors serve as a mediator that allows classification to be performed directly on compressed data without full decoding, maintaining versatility while preserving computational efficiency by avoiding the time-domain conversion.
4Productivity
If side-information descriptors are utilized for classification, then computational efficiency improves, but existing methods fail to use this information
Solution Approach 1:
The classification system is designed to utilize side-information descriptors that are automatically generated during the audio compression process. By making the system adaptive to use this already-available information, the patent improves computational efficiency without requiring additional processing or external data sources, thereby enhancing both efficiency and information utilization.
Data Source
AI summary
An audio file is divided into frames in the time domain and each frame is compressed, according to a psycho-acoustic algorithm, into file in the frequency domain. Each frame is divided into sub-bands and each sub-band is further divided into split sub-bands. The spectral energy over each split sub-band is averaged for all frames. The resulting quantity for each split sub-band provides a parameter. The set of parameters can be compared to a corresponding set of parameters generated from a different audio file to determine whether the audio files are similar. In order to provide for the higher sensitivity of the auditory response, the comparison of individual split sub-bands of the lower order sub-bands can be performed. Selected constants can be used in the comparison process to improve further the sensitivity of the comparison. In the side-information generated by the psycho-acoustic compression, data related to the rhythm, i.e., related percussive effects, is present. The data known as attack flags can also be used as part of the audio frame comparison.


