3D Gesture Recognition Using Depth Cameras and EM Algorithm

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current gesture recognition systems are limited in capturing 3D space-time gesture variations, restricting the design of gesture vocabularies and requiring users to perform gestures in specific 2D planes, often necessitating handheld devices or visual feedback.

Innovation Solution

A 3D free-form gesture recognition system using a 3D camera to capture hand gestures, processing images through a computer that extracts features, constructs template matrices, and employs the Expectation Maximization algorithm to recognize gestures without the need for handheld devices or visual feedback, allowing gestures to be performed in any plane relative to the camera.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single video camera is used for gesture recognition, then the system is simple and inexpensive, but it can only recognize planar gestures in the x-y plane, greatly restricting gesture vocabulary design

Engineering Contradiction:
Improvesystem complexityVSAvoidgesture vocabulary
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent transitions from 2D planar gesture recognition to 3D gesture recognition by introducing depth information through multiple cameras. The system captures gestures in three-dimensional space, allowing users to perform gestures in any orientation and plane, not just parallel to the camera plane. This dimensional expansion enables a much richer gesture vocabulary while maintaining reasonable system complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If 3D vision camera systems are used to overcome 2D system limitations, then gesture recognition capability is improved, but the system complexity and cost increase

Engineering Contradiction:
Improvegesture recognition capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent employs multiple video cameras that serve dual purposes: capturing spatial information for 3D gesture recognition and providing redundant data for improved accuracy. The system processes gestures in any 3D plane using the same camera setup, making the system universally applicable to various gesture types without requiring specialized hardware for each gesture category.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If existing 3D camera systems require gestures to be performed in a specific embedded 2D plane, then recognition accuracy is improved, but user freedom and ease of operation are reduced

Engineering Contradiction:
Improverecognition accuracyVSAvoiduser freedom
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent removes the constraint of requiring gestures to be performed in a specific 2D plane by utilizing full 3D space for gesture recognition. The system captures and processes hand movements in three-dimensional space, allowing users to perform gestures naturally in any orientation while maintaining high recognition accuracy through sophisticated 3D trajectory analysis.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Measurement precision

If handheld pointing devices are used for gesture tracking, then motion detection accuracy is improved, but the system requires additional equipment and reduces ease of operation

Engineering Contradiction:
Improvemotion detection accuracyVSAvoidequipment requirement
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent enables the user's hand to serve as both the gesture tool and the tracking target. By using multiple video cameras to capture the hand's natural movements in 3D space, the system eliminates the need for handheld pointing devices or markers. The hand itself provides the motion information needed for accurate gesture recognition, making the system more intuitive and easier to operate.

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP2535787B13D free-form gesture recognition system and method for character input
Publication Date: 2019.08.07 DEUTSCHE TELEKOM AG
  • EP2535787B1 patent drawingFigure 1
  • EP2535787B1 patent drawingFigure 2
  • EP2535787B1 patent drawingFigure 3

AI summary

A system for 3D free-form gesture recognition, based on character input, which comprises (a) a 3D camera for acquiring images of the 3D gestures when performed in front of the camera; (b) a memory for storing the images; (c) a computer for analyzing images representing trajectories of the gesture and for extracting typical features related to the character during a training stage; (d) a database for storing the typical features after the training stage; and (e) a processing unit for comparing the typical features extracted during the training stage to features extracted online.