AI Video Threat Detection With Merged Bounding Boxes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video analysis systems for surveillance are computationally intensive, leading to slow processing, high energy consumption, scalability challenges, and vulnerability to attacks, as they render every video frame with annotations, making it impractical for security personnel to continuously view live feeds.
Innovation Solution
A system that reduces and merges bounding boxes in video frames to identify suspicious objects, selectively rendering only relevant frames on a user interface, using a surveillance computer with a processor and memory to define, merge, and annotate bounding boxes, and generate alerts for suspicious objects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If every video frame is rendered with annotations on the user interface, then complete visual information is provided to security personnel, but computational intensity increases and processing speed decreases
Solution Approach 1:
The patent extracts only the most critical information (suspicious objects with high confidence scores) from the complete video stream and presents only those frames to security personnel. This extraction approach maintains essential information while dramatically reducing the volume of data requiring human review, thus improving processing speed without significant loss of security-relevant information.
Solution Approach 2:
The system performs partial rendering by displaying only a subset of video frames (specifically, frames containing detected suspicious objects) rather than all frames. This partial action strategy provides sufficient security monitoring capability while reducing computational load and improving overall processing throughput.
2Productivity
If bounding boxes are merged to reduce the number of objects, then processing efficiency improves, but detection precision may be affected
Solution Approach 1:
The patent merges bounding boxes based on spatial proximity and overlap thresholds to reduce the number of individual object detections. This merging approach consolidates redundant detections while maintaining the ability to identify and track multiple suspicious objects, thus improving processing efficiency without significantly compromising detection precision for security-critical objects.
Solution Approach 2:
The system dynamically adjusts detection parameters including confidence score thresholds and bounding box overlap criteria to balance processing efficiency with detection precision. By optimizing these parameters, the system maintains high precision for identifying suspicious objects while achieving improved processing throughput through bounding box merging.
3Reliability
If all video frames are processed and rendered, then complete surveillance coverage is achieved, but energy consumption increases
Solution Approach 1:
The system extracts and processes only the subset of video frames containing suspicious objects rather than processing every frame. This extraction approach maintains effective surveillance coverage for security monitoring while significantly reducing computational energy consumption, as only critical frames require full processing and rendering.
Solution Approach 2:
The system employs periodic processing by continuously monitoring video streams and processing frames at intervals based on detected events. This periodic action maintains surveillance effectiveness by processing frames when suspicious activity is detected while avoiding continuous processing of all frames, thus reducing overall energy consumption.
4Speed
If high frame rates are processed, then real-time surveillance capability is improved, but system complexity and scalability challenges increase
Solution Approach 1:
The system processes only the essential subset of frames (those containing suspicious objects) rather than all frames at high frame rates. This partial processing approach maintains real-time surveillance capability for security events while reducing the overall computational load and system complexity, making the system more scalable to multiple cameras and higher frame rates.
Data Source
AI summary
A device for efficient object detection and selective display of video frames is disclosed. A plurality of bounding boxes to visually bound one or more subjects are defined in the received video frames; each bounding box includes padding above and laterally with respect to the one or more subjects in the video frame. The device may reduce the plurality of bounding boxes by merging bounding boxes to create reduced plurality of merged bounding boxes. The device performs object detection on the reduced plurality of merged bounding boxes by searching a training database for detecting and identifying suspicious objects. The identified objects are annotated. On identifying that an object is suspicious, based on a trained database of firearms in a plurality of environments across a plurality of industries, the video frame is transmitted to a graphical user interface of a user device.


