Depth Camera Gesture Recognition Without Markers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies are unable to accurately interpret human movements without the use of special reflective tags or markers, limiting the ability of computers to assess and respond to human gestures in a natural and intuitive manner.

Innovation Solution

A depth camera system that models human movements using virtual skeletons, allowing users to control interactive interfaces through physical gestures, such as spell-casting gestures, without the need for conventional controllers or markers, by capturing and interpreting depth images to recognize and translate body movements into machine-readable commands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If reflective tags or markers are used to track human movements, then measurement precision is improved, but device complexity and ease of operation deteriorate due to requiring special equipment and preparation

Engineering Contradiction:
Improvemovement tracking accuracyVSAvoidgesture recognition simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent extracts the tracking markers from the system by using a depth camera to directly capture human body geometry and movements without requiring reflective tags or external markers. The depth camera captures unassisted human movements by analyzing depth information from the human body itself, eliminating the need for special equipment attachment.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces virtual skeletons as an intermediary representation between the depth camera data and the gesture recognition system. The virtual skeleton model serves as a mediator that maps real human movements to interpretable gesture commands, enabling accurate movement tracking without physical markers.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multiple cameras and tracking tags are used to achieve accurate 3D position triangulation, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improve3D position accuracyVSAvoidcamera and tag system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent eliminates the need for multiple cameras and external tracking tags by using a single depth camera to capture three-dimensional depth information directly. The depth camera provides volumetric data that enables 3D position calculation without requiring triangulation from multiple camera viewpoints or external markers.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical triangulation system (multiple cameras and physical tags) with an optical depth sensing system. The depth camera uses light time-of-flight or phase-shift measurement to directly obtain 3D position information, substituting complex mechanical tracking infrastructure with a more compact optical sensing approach.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If conventional controllers or markers are required for interaction, then measurement precision is maintained, but ease of operation and adaptability deteriorate

Engineering Contradiction:
Improvegesture detection accuracyVSAvoidinterface control flexibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal gesture-based interface that can control multiple applications and games without requiring application-specific controllers or markers. The depth camera system recognizes a variety of gestures (spelling gestures, aiming gestures, casting gestures) that can be mapped to different functions across diverse applications, providing a single versatile control mechanism.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system enables users to interact directly with the computer interface using their natural body movements without requiring external controllers. The depth camera captures and interprets gestures performed with the user's own body, allowing the interface to serve itself through natural human-computer interaction rather than requiring mediating control devices.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8657683B2Action selection gesturing
Publication Date: 2014.02.25 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8657683B2 patent drawing
  • US8657683B2 patent drawing
  • US8657683B2 patent drawing

AI summary

Gestures of a computer user are observed with a depth camera. A first gesture of the computer user is identified as one of a plurality of different action selection gestures, each action selection gesture associated with a different action performable within an interactive interface controlled by gestures of the computer user. A second gesture is identified as a triggering gesture that causes performance of the action associated with the action selection gesture.