Hand Gesture Detection via Virtual Space
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for remote control of devices using hand gestures face challenges such as high computational cost, low accuracy, and false positives in complex backgrounds, long distances, and low-light environments, particularly when multiple humans are present.
Innovation Solution
A machine-learning based system that uses an end-to-end trained approach for detecting and recognizing hand gestures by defining a virtual gesture-space around a user's face, allowing for real-time detection and recognition with reduced computational resources and improved accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If gesture segmentation and recognition is performed on a sequence of digital video frames using existing machine vision-based detection, then remote control functionality is achieved, but computational cost increases significantly and accuracy remains relatively low
Solution Approach 1:
The patent segments the video processing task by first detecting human bodies and isolating regions containing human hands, then performing gesture recognition only on these segmented regions. This divides the complex full-frame gesture recognition into manageable parts, reducing computational cost while maintaining accuracy.
Solution Approach 2:
The patent extracts and focuses on specific regions containing hands from the entire video frame sequence. By taking out only the relevant hand regions for processing, the system eliminates unnecessary computational overhead from processing entire frames, thereby reducing computational cost while preserving detection accuracy.
2Reliability
If gesture detection is performed in complex backgrounds with multiple humans, then comprehensive detection coverage is achieved, but false positives and false negatives increase
Solution Approach 1:
The patent segments the scene by detecting individual human bodies and associating hands with specific persons. This segmentation allows the system to distinguish between hands belonging to different people in complex backgrounds, reducing false positives and false negatives caused by multiple humans.
Solution Approach 2:
The patent introduces body detection as an intermediary step between frame analysis and gesture recognition. By using body detection results to guide hand detection and gesture recognition, the system effectively filters out irrelevant hands from other people, reducing false positives in complex backgrounds with multiple humans.
3Reliability
If high resolution frames are used to detect smaller hands at farther distances, then detection coverage is improved, but computational cost increases significantly
Solution Approach 1:
The patent applies partial processing by analyzing only the regions containing hands at high resolution, rather than processing entire high-resolution frames. This partial action approach maintains detection accuracy for distant hands while significantly reducing the computational cost associated with processing full high-resolution frames.
Solution Approach 2:
The patent applies different processing qualities to different regions: high-resolution processing is applied locally only to hand-containing regions identified through body detection, while the rest of the frame receives minimal or no processing. This local quality approach ensures accurate detection of distant hands without the prohibitive computational cost of processing entire high-resolution frames.
4Reliability
If the entire field of view is processed for hand detection, then comprehensive gesture detection is achieved, but processing efficiency decreases
Solution Approach 1:
The patent segments the field of view into body-containing regions and processes only these segments for hand detection. This segmentation maintains comprehensive gesture detection coverage within relevant areas while improving processing efficiency by excluding empty or irrelevant regions from intensive processing.
Solution Approach 2:
The patent extracts and processes only the relevant body-containing regions from the entire field of view. By taking out and processing only these essential regions, the system maintains comprehensive gesture detection coverage while significantly improving processing efficiency compared to processing the entire field of view.
Data Source
AI summary
Methods and systems for gesture-based control of a device are described. An input frame is processed to determine a location of a distinguishing anatomical feature in the input frame. A virtual gesture-space is defined based on the location of the distinguishing anatomical feature, the virtual gesture-space being a defined space for detecting a gesture input. The input frame is processed in only the virtual gesture-space, to detect and track a hand. Using information generated from detecting and tracking the at least one hand, a gesture class is determined for the at least one hand. The device may be a smart television, a smart phone, a tablet, etc.


