Video Image Processing with Asymmetric Convolution for Real-Time Gesture Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video image processing technologies face challenges in achieving real-time gesture recognition with poor performance and high false detection rates due to interference from light, and they require large models with low calculation speeds.
Innovation Solution
A video image processing method that determines a target object position region in a current frame and performs a series of convolution processing on the target object tracking image, using a reduced number of convolutions in the first set compared to subsequent sets, to efficiently track and recognize objects in subsequent frames, employing depthwise separable convolution processing and a small neural network model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If large models are used for gesture recognition, then recognition accuracy can be improved, but calculation speed decreases and real-time processing becomes difficult
Solution Approach 1:
The patent segments the video processing into two distinct phases: a coarse detection phase that identifies potential gesture regions, and a fine recognition phase that applies convolution processing only to those specific regions. This segmentation allows the system to use simpler processing for most of the frame while applying more accurate (but computationally intensive) processing only where needed, thus maintaining both speed and accuracy
Solution Approach 2:
The patent applies different processing qualities to different regions of the video frame. The majority of the frame receives minimal processing (coarse detection), while only the identified gesture regions receive full convolution processing (fine recognition). This local quality approach ensures high recognition accuracy is applied only where necessary, optimizing the balance between computational resources and recognition performance
2Measurement precision
If conventional convolution processing is applied to the entire frame, then detection accuracy can be improved, but computational resources and processing time increase significantly
Solution Approach 1:
The patent extracts and processes only the relevant portions of the video frame for gesture detection. By first performing coarse detection to identify potential gesture regions, the system then extracts only those specific regions for detailed convolution processing. This extraction approach eliminates the need to apply computationally intensive processing to the entire frame, significantly reducing energy consumption and computational resources while maintaining detection accuracy for the actual gestures
Solution Approach 2:
The patent applies partial action by performing convolution processing on only a subset of the video frame - specifically, only the regions identified as potential gestures through coarse detection. Rather than applying full processing to the entire frame (excessive action), the system applies processing only where necessary (partial action), optimizing the balance between detection accuracy and computational resource usage
Data Source
AI summary
This present disclosure describes a video image processing method and apparatus, a computer-readable medium and an electronic device, relating to the field of image processing technologies. The method includes determining, by a device, a target-object region in a current frame in a video. The device includes a memory storing instructions and a processor in communication with the memory. The method also includes determining, by the device, a target-object tracking image in a next frame and corresponding to the target-object region; and sequentially performing, by the device, a plurality of sets of convolution processing on the target-object tracking image to determine a target-object region in the next frame. A quantity of convolutions of a first set of convolution processing in the plurality of sets of convolution processing is less than a quantity of convolutions of any other set of convolution processing.


