Audio Classification Using Joint Likelihood Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio classification methods are not reliable for general types of audio signals with mixed sounds and severe background noises, as they are typically designed for specific applications like music genre classification or individual audio event detection and struggle in environments with multiple sound sources and background interference.

Innovation Solution

A method that uses a joint likelihood model in conjunction with semantic concept detectors to analyze audio signals, where the detectors are trained together to determine the co-occurrence of semantic concepts, enhancing the classification accuracy in noisy and complex audio environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If independent audio event detection methods are used, then individual audio events can be detected, but performance deteriorates when multiple audio events occur together

Engineering Contradiction:
Improveaudio event detection accuracyVSAvoidperformance with multiple simultaneous events
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent merges multiple independent audio event detection models into a unified framework that processes multiple events simultaneously. The joint likelihood model combines predictions from individual event detectors, allowing the system to accurately identify and distinguish multiple audio events occurring at the same time, thereby resolving the contradiction between individual event detection accuracy and performance with multiple simultaneous events.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The joint likelihood model acts as an intermediary that receives outputs from individual audio event detectors and mediates their combined information. This intermediary component integrates the detections and uses pair-wise likelihoods to determine the most probable event combinations, enabling accurate detection of multiple simultaneous events while maintaining the benefits of specialized individual detectors.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If existing audio classification methods are used, then classification can be performed on simple audio signals, but reliability deteriorates in the presence of background noise and mixed sounds

Engineering Contradiction:
Improveclassification reliabilityVSAvoidbackground noise interference
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary action by detecting multiple audio events simultaneously and using the joint likelihood model to predict their co-occurrence before final classification. This preliminary joint analysis allows the system to anticipate and compensate for background noise interference, improving classification reliability in noisy environments by pre-establishing the probabilistic relationships between events.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The joint likelihood model provides feedback by continuously evaluating the consistency of event detections against learned pair-wise likelihoods. When background noise causes false detections, the feedback mechanism identifies inconsistencies and adjusts the classification, thereby improving reliability in the presence of harmful background factors.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If music genre classification methods are used, then music audio can be classified, but adaptability deteriorates for general audio signals with mixed sounds

Engineering Contradiction:
Improveapplicability to general audio signalsVSAvoidperformance on mixed sounds with background noise
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent implements universality by designing a multi-functional system where the same joint likelihood model and semantic concept detectors can handle various audio types including music, speech, and environmental sounds. The framework is not limited to music genre classification but extends to general audio signals with mixed sounds, improving both adaptability and reliability across diverse audio scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system achieves adaptability through parameter changes by adjusting the semantic concepts and pair-wise likelihoods based on the specific audio signal characteristics. The joint likelihood model dynamically adapts its parameters to suit different audio environments, enabling reliable classification of mixed sounds with background noise while maintaining versatility across various audio types.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8880444B2Audio based control of equipment and systems
Publication Date: 2014.11.04 KODAK ALARIS LLC
  • US8880444B2 patent drawing
  • US8880444B2 patent drawing
  • US8880444B2 patent drawing

AI summary

A method for controlling a device responsive to an audio signal captured using an audio sensor. A data processor is used to automatically analyze the audio signal using a plurality of semantic concept detectors to determine corresponding preliminary semantic concept detection values, each semantic concept detector being adapted to detect a particular semantic concept. The preliminary semantic concept detection values are analyzed using a joint likelihood model based on predetermined pair-wise likelihoods that particular pairs of semantic concepts co-occur to determine updated semantic concept detection values. One or more semantic concepts are determined based on the updated semantic concept detection values, and the device is controlled responsive to identified semantic concepts. The semantic concept detectors and the joint likelihood model are trained together with a joint training process using training audio signals, at least some of which are known to be associated with a plurality of semantic concepts.