Depth Image Target Tracking via Gesture Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computing applications face challenges in providing intuitive controls for users, as conventional input methods like controllers and keyboards can be difficult to learn and abstracted from actual actions, creating a barrier between users and applications.

Innovation Solution

A target recognition, analysis, and tracking system processes depth images to track users, allowing their gestures and movements to directly control in-game actions or application functions, using a capture device that captures depth information and processes images to isolate and refine the target, creating a binary mask for precise control.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional input controls (controllers, keyboards, mice) are used, then application control functionality is provided, but user learning difficulty increases and intuitiveness decreases

Engineering Contradiction:
Improveintuitiveness of controlVSAvoidcontrol mechanism complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces conventional mechanical input devices (controllers, keyboards, mice) with an optical capture device that tracks natural body movements. The depth image capture device and processing system substitute mechanical control interfaces with direct gesture recognition, allowing users to interact naturally without learning complex control schemes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system creates a digital copy or representation of the user's physical gestures through depth image capture and processing. The capture device records actual body movements, and the processing system generates corresponding control signals by analyzing the captured depth data, effectively copying natural motion into application control.

Inventive Principle:
Principle #26Copying

2Measurement precision

If depth image processing is performed to isolate and refine targets, then tracking precision is improved, but computational complexity increases

Engineering Contradiction:
Improvetarget tracking precisionVSAvoidimage processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the depth image processing into distinct functional stages: initial capture, downsampling to reduce resolution, shadow/noise/missing portion identification and estimation, planar surface detection, target scanning, and binary mask creation. This segmentation allows complex processing to be broken into manageable steps that progressively refine target identification.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary processing actions before final target identification, including downsampling the depth image to reduce computational load, pre-identifying and estimating shadow and noise regions, and detecting planar surfaces early in the process. These preliminary actions prepare the data for more efficient subsequent target isolation and masking operations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8988432B2Systems and methods for processing an image for target tracking
Publication Date: 2015.03.24 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8988432B2 patent drawing
  • US8988432B2 patent drawing
  • US8988432B2 patent drawing

AI summary

An image such as a depth image of a scene may be received, observed, or captured by a device. The image may then be processed. For example, the image may be downsampled, a shadow, noise, and/or a missing potion in the image may be determined, pixels in the image that may be outside a range defined by a capture device associated with the image may be determined, a portion of the image associated with a floor may be detected. Additionally, a target in the image may be determined and scanned. A refined image may then be rendered based on the processed image. The refined image may then be processed to, for example, track a user.