Gesture Recognition Using Depth Maps and Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current gesture recognition systems are insufficient in recognizing gestures and providing subsequent processing, limiting their effectiveness in various operating contexts and applications.

Innovation Solution

The development of gesture recognition systems utilizing machine vision and computer-aided methods, including the use of time-of-flight sensors and stereoscopic camera configurations, to generate depth maps and recognize user gestures, enabling interaction with computers and performing tasks without physical contact.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional gesture recognition systems are used, then basic gesture detection is possible, but recognition accuracy and system effectiveness are insufficient

Engineering Contradiction:
Improvegesture recognition accuracyVSAvoidsystem effectiveness
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent combines multiple imaging modalities (stereoscopic cameras for depth mapping and time-of-flight sensors for 3D gesture capture) into a unified gesture recognition system. This merging of sensing technologies enables more accurate and reliable gesture detection by compensating for the limitations of individual sensors and providing complementary data streams for comprehensive gesture analysis.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system transitions from traditional 2D image-based gesture recognition to 3D depth-based gesture recognition by incorporating stereoscopic camera pairs and time-of-flight sensors. This dimensional enhancement allows the system to capture spatial information, hand orientation, and gesture trajectories in three-dimensional space, significantly improving recognition accuracy and reducing ambiguity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of operation

If contact-based interaction is used, then precise input is possible, but accessibility for individuals with motor impairments is limited

Engineering Contradiction:
ImproveaccessibilityVSAvoidinput precision
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent replaces mechanical contact-based interaction (physical touching of screens or buttons) with contactless optical sensing systems. Stereoscopic cameras and time-of-flight sensors detect hand gestures, finger movements, and body language in 3D space, enabling individuals with motor impairments to interact with computers through intuitive gestures without requiring physical contact, while maintaining precise input capability through advanced image processing and gesture recognition algorithms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables accurate gesture recognition in diverse environments, facilitating human-computer interaction, enhancing accessibility for individuals with motor impairments, and supporting applications like interactive gaming, secure access, and vehicle guidance.

Implementation Method 1

time-of-flight sensors and stereoscopic camera configurations, to generate depth maps

Methodology Applied
Scientific EffectTime of flight: Time of Flight

Implementation Method 2

stereoscopic camera configurations, to generate depth maps

Methodology Applied
Scientific EffectStereoscopy: Parallax

Data Source

PatentUS12105887B1Gesture recognition systems
Publication Date: 2024.10.01 GOLDEN EDGE HOLDING CORP
  • US12105887B1 patent drawing
  • US12105887B1 patent drawing
  • US12105887B1 patent drawing

AI summary

A method and apparatus for performing gesture recognition. In one embodiment of the invention, the method includes the steps of receiving one or more raw frames from one or more cameras, each of the one or more raw frames representing a time sequence of images, determining one or more regions of the one or more received raw frames that comprise highly textured regions, segmenting the one or more determined highly textured regions in accordance textured features thereof to determine one or more segments thereof, determining one or more regions of the one or more received raw frames that comprise other than highly textured regions, and segmenting the one or more determined other than highly textured regions in accordance with color thereof to determine one or more segments thereof. One or more of the segments are then tracked through the one or more raw frames representing the time sequence of images.