3D Gesture Recognition Using Depth Cameras and EM Algorithm
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current gesture recognition systems are limited in capturing 3D space-time gesture variations, restricting the design of gesture vocabularies and requiring users to perform gestures in specific 2D planes, often necessitating handheld devices or visual feedback.
Innovation Solution
A 3D free-form gesture recognition system using a 3D camera to capture hand gestures, processing images through a computer that extracts features, constructs template matrices, and employs the Expectation Maximization algorithm to recognize gestures without the need for handheld devices or visual feedback, allowing gestures to be performed in any plane relative to the camera.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single video camera is used for gesture recognition, then the system is simple and inexpensive, but it can only recognize planar gestures in the x-y plane, greatly restricting gesture vocabulary design
Solution Approach 1:
The patent transitions from 2D planar gesture recognition to 3D gesture recognition by introducing depth information through multiple cameras. The system captures gestures in three-dimensional space, allowing users to perform gestures in any orientation and plane, not just parallel to the camera plane. This dimensional expansion enables a much richer gesture vocabulary while maintaining reasonable system complexity.
2Adaptability or versatility
If 3D vision camera systems are used to overcome 2D system limitations, then gesture recognition capability is improved, but the system complexity and cost increase
Solution Approach 1:
The patent employs multiple video cameras that serve dual purposes: capturing spatial information for 3D gesture recognition and providing redundant data for improved accuracy. The system processes gestures in any 3D plane using the same camera setup, making the system universally applicable to various gesture types without requiring specialized hardware for each gesture category.
3Measurement precision
If existing 3D camera systems require gestures to be performed in a specific embedded 2D plane, then recognition accuracy is improved, but user freedom and ease of operation are reduced
Solution Approach 1:
The patent removes the constraint of requiring gestures to be performed in a specific 2D plane by utilizing full 3D space for gesture recognition. The system captures and processes hand movements in three-dimensional space, allowing users to perform gestures naturally in any orientation while maintaining high recognition accuracy through sophisticated 3D trajectory analysis.
4Measurement precision
If handheld pointing devices are used for gesture tracking, then motion detection accuracy is improved, but the system requires additional equipment and reduces ease of operation
Solution Approach 1:
The patent enables the user's hand to serve as both the gesture tool and the tracking target. By using multiple video cameras to capture the hand's natural movements in 3D space, the system eliminates the need for handheld pointing devices or markers. The hand itself provides the motion information needed for accurate gesture recognition, making the system more intuitive and easier to operate.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A system for 3D free-form gesture recognition, based on character input, which comprises (a) a 3D camera for acquiring images of the 3D gestures when performed in front of the camera; (b) a memory for storing the images; (c) a computer for analyzing images representing trajectories of the gesture and for extracting typical features related to the character during a training stage; (d) a database for storing the typical features after the training stage; and (e) a processing unit for comparing the typical features extracted during the training stage to features extracted online.