Surveillance Aggression Prediction Using Edge Knowledge Graph Reasoning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Surveillance cameras lack sufficient computing power to run complex AI models for accurate real-time detection of aggressive behavior, limiting their ability to process both audio and video data effectively, especially in multi-detection scenarios.
Innovation Solution
Deploy multiple classification models for audio and video information into networked cameras that relay metadata to a central system with an edge server running a knowledge graph and query-based inference engine, utilizing a local computing system for metadata generation and a remote system for knowledge graph-based reasoning to predict aggressive behavior.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If complex AI models are deployed in surveillance cameras for accurate aggressive behavior detection, then detection precision is improved, but device complexity and computing power requirements increase beyond what camera systems can provide
Solution Approach 1:
The system divides the aggressive behavior detection task into two segments: (1) camera-based object classification and event detection for basic identification, and (2) remote knowledge graph-based reasoning for comprehensive aggressive behavior analysis. This segmentation allows cameras to perform simple tasks within their computing constraints while achieving accurate aggressive behavior detection through the combined system.
Solution Approach 2:
A remote computing system with knowledge graph technology acts as an intermediary between the camera system and the final detection output. The camera captures and processes basic video data, then transfers relevant information to the remote system which performs complex reasoning about aggressive behaviors, returning results to the camera system.
2Measurement precision
If comprehensive AI models process both audio and video data in real-time, then detection accuracy is improved, but processing speed decreases due to computational complexity
Solution Approach 1:
The processing pipeline is segmented into fast camera-based classification for immediate response and slower remote knowledge graph reasoning for comprehensive analysis. This allows the system to maintain real-time responsiveness for basic events while performing detailed aggressive behavior analysis at appropriate intervals.
Solution Approach 2:
The camera system performs partial processing of audio and video data locally, focusing on detecting basic objects and events. The complete multi-sensor analysis is performed partially by the remote system, which processes only the necessary metadata and classifications rather than all raw sensor data, achieving good detection accuracy without full real-time processing of all data.
3Loss of information
If large AI models are used for holistic environmental understanding, then contextual comprehension is improved, but the computing power required exceeds camera system capabilities
Solution Approach 1:
The remote computing system serves as an intermediary that performs holistic environmental understanding using knowledge graphs. The camera system captures environmental data and transfers relevant information to the remote system, which comprehensively analyzes the environmental context and returns actionable insights, distributing the computational burden appropriately.
Solution Approach 2:
Instead of running large AI models directly on camera systems, the patent creates simplified copies or representations of environmental data (metadata, classifications, event detections) that are transmitted to the remote system. The remote system then performs comprehensive environmental understanding on these copied representations rather than on the original full-resolution data streams.
Data Source
AI summary
Methods and systems for predicting aggressive behavior associated with a surveillance scene. Images are generated from a camera of a surveillance scene, and audio is also generated. A local computing system executes an object classification model on the images to predict one or more classes of objects in the scene. The local computing system also executes a sound-event detection model on the audio to predict one or more classes of events occurring in the scene. Metadata is generated associated with the image-based classes and audio-based classes. The metadata is transferred to a remote computing system that executes a knowledge graph on the metadata to implement knowledge graph-based reasoning to predict aggressive behavior occurring in the surveillance scene based on the metadata. The metadata associated with the predicted aggressive behavior is labeled as such, and control commands are output accordingly.


