Audio Scene Classification Using Foreground-Background Feature Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current acoustic scene classification (ASC) methods face challenges in real-life environments due to the unbounded and hard-to-generalize nature of acoustic events, requiring manual definition and selection of sound events, which is unrealistic and inefficient.

Innovation Solution

The proposed solution involves merging frame-level features with binary features characterizing affinity to foreground or background in an 'event-informed' deep neural network (DNN), where the binary layer feature is used as a target in pre-training and a control parameter in training and classification stages to improve feature effectiveness and adapt to the acoustic scene's nature.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual definition and selection of sound events is used, then classification accuracy can be improved for constrained scenarios, but the method becomes unrealistic and inefficient for real-life environments with unbounded acoustic events

Engineering Contradiction:
Improveclassification accuracyVSAvoidoperational efficiency
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system automatically discovers and selects acoustic events through unsupervised clustering algorithms without requiring manual definition. The algorithm processes audio features, clusters them into event types, and uses these discovered events for scene classification, enabling the system to serve itself in real-life environments with unbounded events

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The invention changes the approach from fixed manual event definitions to dynamic event discovery by modifying key parameters: using unsupervised clustering instead of manual labeling, and adapting the number and types of events based on the specific acoustic environment being analyzed

Inventive Principle:
Principle #35Parameter changes

2Reliability

If all sound events in real-life environments are manually defined, then comprehensive scene characterization can be achieved, but the complexity and time required becomes unrealistic

Engineering Contradiction:
Improvescene characterization completenessVSAvoidevent definition time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary unsupervised clustering of acoustic events before classification, automatically organizing sound events into meaningful categories without manual intervention. This preliminary organization enables comprehensive scene characterization to be achieved rapidly without manual definition of all possible events

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of requiring complete manual definition of all possible events, the system uses unsupervised clustering to discover the most relevant events present in the specific acoustic environment, achieving sufficient completeness without the excessive time cost of defining all potential events

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If current ASC schemes are applied to softly constrained problems, then the approach fails to generalize because the set of acoustic events is unbounded and extremely hard to generalise

Engineering Contradiction:
Improvegeneralization capabilityVSAvoidclassification reliability
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The unsupervised event discovery algorithm serves multiple functions: it automatically adapts to different acoustic environments, discovers environment-specific events, and provides a universal framework that works across softly constrained problems without requiring environment-specific manual definitions, thereby achieving both versatility and reliability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3847646B1An audio processing apparatus and method for audio scene classification
Publication Date: 2023.10.04 HUAWEI TECH CO LTD
  • EP3847646B1 patent drawingFigure 1a~1b
  • EP3847646B1 patent drawingFigure 2
  • EP3847646B1 patent drawingFigure 3

AI summary

The invention relates to an audio processing apparatus (200) configured to classify an audio signal into one or more audio scene classes, the audio signal comprising a component signal. The apparatus (200) comprises: processing circuitry configured to classify the component signal of the audio signal as a foreground layer component signal or a background layer component signal; obtain an audio signal feature on the basis of the audio signal; select, depending on the classification of the component signal, a first set of weights or a second set of weights; and to classify the audio signal on the basis of the audio signal features, the foreground layer component signal or the background layer component signal and the selected set of weights.