Video Image Processing with Asymmetric Convolution for Real-Time Gesture Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video image processing technologies face challenges in achieving real-time gesture recognition with poor performance and high false detection rates due to interference from light, and they require large models with low calculation speeds.

Innovation Solution

A video image processing method that determines a target object position region in a current frame and performs a series of convolution processing on the target object tracking image, using a reduced number of convolutions in the first set compared to subsequent sets, to efficiently track and recognize objects in subsequent frames, employing depthwise separable convolution processing and a small neural network model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If large models are used for gesture recognition, then recognition accuracy can be improved, but calculation speed decreases and real-time processing becomes difficult

Engineering Contradiction:
Improvegesture recognition accuracyVSAvoidcalculation speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the video processing into two distinct phases: a coarse detection phase that identifies potential gesture regions, and a fine recognition phase that applies convolution processing only to those specific regions. This segmentation allows the system to use simpler processing for most of the frame while applying more accurate (but computationally intensive) processing only where needed, thus maintaining both speed and accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing qualities to different regions of the video frame. The majority of the frame receives minimal processing (coarse detection), while only the identified gesture regions receive full convolution processing (fine recognition). This local quality approach ensures high recognition accuracy is applied only where necessary, optimizing the balance between computational resources and recognition performance

Inventive Principle:
Principle #3Local quality

2Measurement precision

If conventional convolution processing is applied to the entire frame, then detection accuracy can be improved, but computational resources and processing time increase significantly

Engineering Contradiction:
Improvetarget object detection accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts and processes only the relevant portions of the video frame for gesture detection. By first performing coarse detection to identify potential gesture regions, the system then extracts only those specific regions for detailed convolution processing. This extraction approach eliminates the need to apply computationally intensive processing to the entire frame, significantly reducing energy consumption and computational resources while maintaining detection accuracy for the actual gestures

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by performing convolution processing on only a subset of the video frame - specifically, only the regions identified as potential gestures through coarse detection. Rather than applying full processing to the entire frame (excessive action), the system applies processing only where necessary (partial action), optimizing the balance between detection accuracy and computational resource usage

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11436739B2Method, apparatus, and storage medium for processing video image
Publication Date: 2022.09.06 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US11436739B2 patent drawing
  • US11436739B2 patent drawing
  • US11436739B2 patent drawing

AI summary

This present disclosure describes a video image processing method and apparatus, a computer-readable medium and an electronic device, relating to the field of image processing technologies. The method includes determining, by a device, a target-object region in a current frame in a video. The device includes a memory storing instructions and a processor in communication with the memory. The method also includes determining, by the device, a target-object tracking image in a next frame and corresponding to the target-object region; and sequentially performing, by the device, a plurality of sets of convolution processing on the target-object tracking image to determine a target-object region in the next frame. A quantity of convolutions of a first set of convolution processing in the plurality of sets of convolution processing is less than a quantity of convolutions of any other set of convolution processing.