Audio Scene Understanding via Grammar-Based Event Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio monitoring systems are inadequate for detecting complex events involving multiple objects in an environment over prolonged periods, as they focus on narrow classes of actions for single objects and lack the capability to identify interactions between multiple objects.
Innovation Solution
An audio monitoring system that trains classifiers for specific object actions and generates a scene grammar model to identify sound events by integrating object relationships and temporal data, allowing for the detection of complex events in a scene through a processor and sound sensors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional audio monitoring systems are used to detect sound events, then detection of narrow classes of actions for single objects is achieved, but capability to identify complex events involving multiple objects is lost
Solution Approach 1:
The patent combines multiple audio sensors into a distributed sensor network that collectively monitors multiple objects in an environment. The system merges data from multiple sensors and integrates it with scene grammar models to detect complex events involving interactions between multiple objects, thereby achieving both precise single-object detection and comprehensive multi-object event identification.
Solution Approach 2:
The scene grammar model serves as a universal framework that can represent multiple types of sound events, object interactions, and temporal patterns within a single system. This multi-functional approach enables the system to detect various complex events including sequences of actions, concurrent actions, and causal relationships between different objects without requiring separate specialized systems for each event type.
2Loss of information
If video monitoring is deployed to monitor environment and identify complex events, then comprehensive scene understanding is achieved, but intrusiveness and privacy concerns increase
Solution Approach 1:
The patent replaces video-based optical monitoring with audio-based sensing to achieve scene understanding. Audio sensors capture sound events that provide information about object interactions and environmental activities without requiring visual observation. This substitution maintains comprehensive scene monitoring capability while eliminating the intrusiveness and privacy concerns associated with video cameras in sensitive environments.
3Adaptability or versatility
If multiple smart devices are deployed to monitor multiple objects, then monitoring capability for complex events is improved, but system complexity and cost increase
Solution Approach 1:
The patent segments the monitoring system into distinct functional components: audio sensors for data collection, a scene grammar model for event representation, and a detection engine for event identification. This segmentation allows each component to be optimized independently and enables modular deployment where systems can be scaled from simple to complex configurations based on specific needs, reducing overall system complexity while maintaining comprehensive monitoring capability.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method of operating an audio monitoring system includes generating with a sound sensor audio data corresponding to a sound event generated by an object in a scene around the sound sensor, identifying with a processor a type and action of the object in the scene that generated the sound with reference to the audio data, generating with the processor a timestamp corresponding to a time of the detection of the sound event, and updating a scene state model corresponding to sound events generated by a plurality of objects in the scene with reference to the identified type of object, action taken by the object, and the timestamp. The method further includes identifying a sound event in the scene with reference to the scene state model and a predetermined scene grammar stored in a memory, and generating with the processor an output corresponding to the sound event.