Audio Classification Using Joint Likelihood Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio classification methods are not reliable for general types of audio signals with mixed sounds and severe background noises, as they are typically designed for specific applications like music genre classification or individual audio event detection and struggle in environments with multiple sound sources and background interference.
Innovation Solution
A method that uses a joint likelihood model in conjunction with semantic concept detectors to analyze audio signals, where the detectors are trained together to determine the co-occurrence of semantic concepts, enhancing the classification accuracy in noisy and complex audio environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If independent audio event detection methods are used, then individual audio events can be detected, but performance deteriorates when multiple audio events occur together
Solution Approach 1:
The patent merges multiple independent audio event detection models into a unified framework that processes multiple events simultaneously. The joint likelihood model combines predictions from individual event detectors, allowing the system to accurately identify and distinguish multiple audio events occurring at the same time, thereby resolving the contradiction between individual event detection accuracy and performance with multiple simultaneous events.
Solution Approach 2:
The joint likelihood model acts as an intermediary that receives outputs from individual audio event detectors and mediates their combined information. This intermediary component integrates the detections and uses pair-wise likelihoods to determine the most probable event combinations, enabling accurate detection of multiple simultaneous events while maintaining the benefits of specialized individual detectors.
2Reliability
If existing audio classification methods are used, then classification can be performed on simple audio signals, but reliability deteriorates in the presence of background noise and mixed sounds
Solution Approach 1:
The system performs preliminary action by detecting multiple audio events simultaneously and using the joint likelihood model to predict their co-occurrence before final classification. This preliminary joint analysis allows the system to anticipate and compensate for background noise interference, improving classification reliability in noisy environments by pre-establishing the probabilistic relationships between events.
Solution Approach 2:
The joint likelihood model provides feedback by continuously evaluating the consistency of event detections against learned pair-wise likelihoods. When background noise causes false detections, the feedback mechanism identifies inconsistencies and adjusts the classification, thereby improving reliability in the presence of harmful background factors.
3Adaptability or versatility
If music genre classification methods are used, then music audio can be classified, but adaptability deteriorates for general audio signals with mixed sounds
Solution Approach 1:
The patent implements universality by designing a multi-functional system where the same joint likelihood model and semantic concept detectors can handle various audio types including music, speech, and environmental sounds. The framework is not limited to music genre classification but extends to general audio signals with mixed sounds, improving both adaptability and reliability across diverse audio scenarios.
Solution Approach 2:
The system achieves adaptability through parameter changes by adjusting the semantic concepts and pair-wise likelihoods based on the specific audio signal characteristics. The joint likelihood model dynamically adapts its parameters to suit different audio environments, enabling reliable classification of mixed sounds with background noise while maintaining versatility across various audio types.
Data Source
AI summary
A method for controlling a device responsive to an audio signal captured using an audio sensor. A data processor is used to automatically analyze the audio signal using a plurality of semantic concept detectors to determine corresponding preliminary semantic concept detection values, each semantic concept detector being adapted to detect a particular semantic concept. The preliminary semantic concept detection values are analyzed using a joint likelihood model based on predetermined pair-wise likelihoods that particular pairs of semantic concepts co-occur to determine updated semantic concept detection values. One or more semantic concepts are determined based on the updated semantic concept detection values, and the device is controlled responsive to identified semantic concepts. The semantic concept detectors and the joint likelihood model are trained together with a joint training process using training audio signals, at least some of which are known to be associated with a plurality of semantic concepts.


