Audio Classification by Integrated Feature Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods fail to accurately classify moving images based on audio content, specifically when the audio consists of a mixture of various sounds, as they cannot determine the type of event or situation represented by the audio.
Innovation Solution
An audio classification device that extracts section features from audio signals, calculates section similarity, and integrates these features to classify the audio signal by comparing them to reference features, allowing for the identification of the event or situation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional audio classification methods are used, then individual sound segments can be classified, but the overall event or situation cannot be determined
Solution Approach 1:
The patent merges section-level features into an integrated feature that represents the overall audio content. The integrated feature extraction unit combines multiple section features (extracted from individual sound segments) to create a comprehensive representation of the entire audio signal, enabling classification of events or situations rather than just individual sounds.
Solution Approach 2:
The patent segments the audio signal into multiple sections with predetermined lengths, extracts features from each segment, and then integrates these segmented features to classify the overall audio. This segmentation allows detailed analysis of individual sound segments while the integration step recovers the broader event context.
2Reliability
If manual classification methods are used, then moving images can be categorized, but user effort and time increase
Solution Approach 1:
The system performs automatic audio classification without requiring manual user input. The classification unit automatically compares integrated features against reference features and assigns categories to moving images autonomously, eliminating the need for users to manually tag or categorize their media files.
Solution Approach 2:
The system extracts and integrates audio features in advance during the classification process, preparing the data structure needed for comparison before the actual classification decision is made. This preliminary feature extraction and integration enables rapid automated classification without requiring real-time manual analysis.
Data Source
AI summary
To classify moving images using audio signals. An audio signal is acquired, a section feature relating to an audio frequency distribution is extracted with respect to each of a plurality of sections each having a predetermined length contained in the acquired audio signal, each extracted section feature is compared with each of reference section features to calculate a section similarity indicating a degree of correlation between each section feature and each reference section feature. An integrated feature relating to the plurality of sections and being calculated based on the section similarity calculated with respect to each of the plurality of sections is extracted from the acquired audio signal. The extracted integrated feature is compared with each of one or more reference integrated features, and the audio signal is classified based on comparison result. Then, classification result is used for moving image classification.


