AI Video Threat Detection With Merged Bounding Box Rendering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video analysis systems face computational intensity, scalability challenges, and high energy consumption due to the rendering of every video frame with annotations, which impacts reaction time and system vulnerability, especially in surveillance applications where faster processing is critical for identifying suspicious objects.
Innovation Solution
A system that reduces and merges bounding boxes in video frames to create merged bounding boxes while maintaining a pre-defined aspect ratio, detects and annotates objects, and renders frames only when suspicious objects are identified, using a surveillance computer with a processor and memory to manage video frames and generate alerts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If every video frame is rendered with annotations, then complete visual information is provided to security personnel, but computational intensity and energy consumption increase significantly
Solution Approach 1:
The patent extracts only the essential information (suspicious objects with bounding boxes) from the complete video frame data and transmits only this extracted information to security personnel, rather than transmitting the entire annotated frame. This reduces energy consumption while maintaining critical information delivery.
Solution Approach 2:
Instead of rendering every frame completely (excessive action), the system performs partial rendering by displaying only the bounding boxes and annotated suspicious objects on a simplified interface, reducing computational load while providing sufficient visual information for security personnel.
2Loss of information
If every video frame is rendered with annotations, then all detected objects are visible, but processing speed decreases and reaction time increases
Solution Approach 1:
The system extracts only the critical detection results (bounding boxes and suspicious object labels) from the complete frame processing and transmits only this condensed information, enabling faster processing speeds while maintaining complete object detection coverage through the bounding box annotations.
Solution Approach 2:
The patent segments the video frame processing into two stages: complete processing for suspicious object detection, and simplified rendering for display. This segmentation allows the system to maintain high processing speed by avoiding redundant rendering operations while preserving complete detection information through the bounding box annotations.
3Reliability
If all video frames are processed and rendered, then comprehensive surveillance coverage is achieved, but hardware requirements and costs increase
Solution Approach 1:
The system extracts only the essential detection information (bounding boxes and object labels) from full frame processing and transmits only this extracted data, reducing hardware requirements while maintaining comprehensive surveillance coverage through efficient data transmission and processing.
Solution Approach 2:
Instead of performing complete rendering operations on all frames (excessive action), the system performs partial processing by displaying only annotated suspicious objects and bounding boxes, reducing hardware requirements while maintaining adequate surveillance coverage for security personnel.
4Productivity
If bounding boxes are merged to reduce the number of objects, then processing efficiency improves, but the precision of individual object location may be affected
Solution Approach 1:
The patent merges bounding boxes that are spatially adjacent or overlapping to reduce the total number of detection objects, improving processing efficiency by reducing the number of items to display and process, while maintaining location precision through the merged bounding box that still accurately encompasses the original objects.
Data Source
AI summary
A device for efficient object detection and selective display of video frames is disclosed. A plurality of bounding boxes to visually bound one or more subjects are defined in the received video frames; each bounding box includes padding above and laterally with respect to the one or more subjects in the video frame. The device may reduce the plurality of bounding boxes by merging bounding boxes to create reduced plurality of merged bounding boxes. The device performs object detection on the reduced plurality of merged bounding boxes by searching a training database for detecting and identifying suspicious objects. The identified objects are annotated. On identifying that an object is suspicious, based on a trained database of firearms in a plurality of environments across a plurality of industries, the video frame is transmitted to a graphical user interface of a user device.


