Depth Image Target Tracking via Gesture Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computing applications face challenges in providing intuitive controls for users, as conventional input methods like controllers and keyboards can be difficult to learn and abstracted from actual actions, creating a barrier between users and applications.
Innovation Solution
A target recognition, analysis, and tracking system processes depth images to track users, allowing their gestures and movements to directly control in-game actions or application functions, using a capture device that captures depth information and processes images to isolate and refine the target, creating a binary mask for precise control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional input controls (controllers, keyboards, mice) are used, then application control functionality is provided, but user learning difficulty increases and intuitiveness decreases
Solution Approach 1:
The patent replaces conventional mechanical input devices (controllers, keyboards, mice) with an optical capture device that tracks natural body movements. The depth image capture device and processing system substitute mechanical control interfaces with direct gesture recognition, allowing users to interact naturally without learning complex control schemes.
Solution Approach 2:
The system creates a digital copy or representation of the user's physical gestures through depth image capture and processing. The capture device records actual body movements, and the processing system generates corresponding control signals by analyzing the captured depth data, effectively copying natural motion into application control.
2Measurement precision
If depth image processing is performed to isolate and refine targets, then tracking precision is improved, but computational complexity increases
Solution Approach 1:
The patent segments the depth image processing into distinct functional stages: initial capture, downsampling to reduce resolution, shadow/noise/missing portion identification and estimation, planar surface detection, target scanning, and binary mask creation. This segmentation allows complex processing to be broken into manageable steps that progressively refine target identification.
Solution Approach 2:
The system performs preliminary processing actions before final target identification, including downsampling the depth image to reduce computational load, pre-identifying and estimating shadow and noise regions, and detecting planar surfaces early in the process. These preliminary actions prepare the data for more efficient subsequent target isolation and masking operations.
Data Source
AI summary
An image such as a depth image of a scene may be received, observed, or captured by a device. The image may then be processed. For example, the image may be downsampled, a shadow, noise, and/or a missing potion in the image may be determined, pixels in the image that may be outside a range defined by a capture device associated with the image may be determined, a portion of the image associated with a floor may be detected. Additionally, a target in the image may be determined and scanned. A refined image may then be rendered based on the processed image. The refined image may then be processed to, for example, track a user.


