Video Analysis System for Target Item Interaction Monitoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current intelligent video analysis systems have low accuracy and are inefficient in monitoring specific scenarios, such as front desk personnel usage, leading to missed video information and high resource wastage.
Innovation Solution
A method that determines the interaction between a target item and a monitored target object by analyzing video frames for the presence of the target item, head, hand, and face orientation, and adjusts thresholds based on verification results to improve detection accuracy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If human operators stare at the screen TV wall for long time to monitor video information, then monitoring coverage is comprehensive, but human resources are wasted and video information may be missed
Solution Approach 1:
The patent replaces the mechanical human monitoring system with an intelligent video analysis system that uses computer vision algorithms to automatically detect and analyze video content. The system processes video frames through multiple detection stages (target item detection, head detection, hand detection, face orientation analysis) to identify interactions between targets and monitored objects, eliminating the need for continuous human observation while maintaining or improving monitoring reliability.
2Extent of automation
If current intelligent video analysis systems are used for special detection scenarios, then automation is achieved, but detection accuracy is low
Solution Approach 1:
The patent divides the video analysis process into multiple segmented detection stages: first detecting target items in video frames, then detecting heads of monitored objects, followed by hand detection, and finally face orientation analysis. This segmentation allows each detection module to focus on specific features, improving overall detection accuracy while maintaining full automation. The multi-stage approach enables the system to process complex interaction scenarios by breaking them down into manageable detection steps.
Solution Approach 2:
The patent applies local quality by performing detections only in relevant regions of video frames. After detecting target items, the system focuses subsequent detections (head, hand, face orientation) specifically on regions containing these targets. This localized approach improves detection accuracy by concentrating computational resources on areas of interest while maintaining automation efficiency.
3Reliability
If comprehensive video analysis is performed on all frames, then detection thoroughness is high, but processing efficiency is low
Solution Approach 1:
The patent implements preliminary action by performing target item detection on video frames before conducting more computationally intensive detections. The system first identifies regions containing target items, then only performs head, hand, and face orientation detections on frames where targets are present. This preliminary filtering reduces the total number of complex detections needed, improving processing efficiency while maintaining detection thoroughness for relevant frames.
Solution Approach 2:
The patent applies partial action by selectively performing comprehensive analysis only on video frames that contain target items and monitored objects. Rather than analyzing all frames equally, the system focuses computational resources on frames where interactions are likely to occur, determined through preliminary target detection. This approach maintains detection thoroughness for critical frames while reducing overall processing load.
Data Source
AI summary
A method, an apparatus, a computing device, and a computer-readable storage medium for monitoring use of target item are disclosed. The method includes obtaining a RoI of each of video frames, for each frame, determining whether the target item and a head of a monitored target object are both present in the RoI, if yes, determining whether a hand of the monitored target object is present in the RoI, if the hand not present, determining a face orientation of the monitored target object to determining whether the each frame is in a first video frame status, if the hand presents, determining a relative position relationship between the hand and the target item to determining whether the each frame is in the first video frame status, and based on a number of video frames continuously in the first video frame status, determining whether the monitored target object uses the target item.


