Audio Scene Classification Using Foreground-Background Feature Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current acoustic scene classification (ASC) methods face challenges in real-life environments due to the unbounded and hard-to-generalize nature of acoustic events, requiring manual definition and selection of sound events, which is unrealistic and inefficient.
Innovation Solution
The proposed solution involves merging frame-level features with binary features characterizing affinity to foreground or background in an 'event-informed' deep neural network (DNN), where the binary layer feature is used as a target in pre-training and a control parameter in training and classification stages to improve feature effectiveness and adapt to the acoustic scene's nature.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual definition and selection of sound events is used, then classification accuracy can be improved for constrained scenarios, but the method becomes unrealistic and inefficient for real-life environments with unbounded acoustic events
Solution Approach 1:
The system automatically discovers and selects acoustic events through unsupervised clustering algorithms without requiring manual definition. The algorithm processes audio features, clusters them into event types, and uses these discovered events for scene classification, enabling the system to serve itself in real-life environments with unbounded events
Solution Approach 2:
The invention changes the approach from fixed manual event definitions to dynamic event discovery by modifying key parameters: using unsupervised clustering instead of manual labeling, and adapting the number and types of events based on the specific acoustic environment being analyzed
2Reliability
If all sound events in real-life environments are manually defined, then comprehensive scene characterization can be achieved, but the complexity and time required becomes unrealistic
Solution Approach 1:
The system performs preliminary unsupervised clustering of acoustic events before classification, automatically organizing sound events into meaningful categories without manual intervention. This preliminary organization enables comprehensive scene characterization to be achieved rapidly without manual definition of all possible events
Solution Approach 2:
Instead of requiring complete manual definition of all possible events, the system uses unsupervised clustering to discover the most relevant events present in the specific acoustic environment, achieving sufficient completeness without the excessive time cost of defining all potential events
3Adaptability or versatility
If current ASC schemes are applied to softly constrained problems, then the approach fails to generalize because the set of acoustic events is unbounded and extremely hard to generalise
Solution Approach 1:
The unsupervised event discovery algorithm serves multiple functions: it automatically adapts to different acoustic environments, discovers environment-specific events, and provides a universal framework that works across softly constrained problems without requiring environment-specific manual definitions, thereby achieving both versatility and reliability
Data Source
Figure 1a~1b
Figure 2
Figure 3
AI summary
The invention relates to an audio processing apparatus (200) configured to classify an audio signal into one or more audio scene classes, the audio signal comprising a component signal. The apparatus (200) comprises: processing circuitry configured to classify the component signal of the audio signal as a foreground layer component signal or a background layer component signal; obtain an audio signal feature on the basis of the audio signal; select, depending on the classification of the component signal, a first set of weights or a second set of weights; and to classify the audio signal on the basis of the audio signal features, the foreground layer component signal or the background layer component signal and the selected set of weights.