Voxel Grid Skeletal Model Tracking for Intuitive Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computing applications face challenges in providing intuitive controls for users, as conventional input methods can be difficult to learn and may not directly correspond to actual actions, creating a barrier between users and applications.

Innovation Solution

A system and method for tracking a user in a scene using depth images to generate a grid of voxels, isolate foreground objects, determine extremities, and adjust a skeletal model to mimic user movements, allowing for more natural and intuitive control inputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional input controls (keyboards, mice, controllers) are used, then applications can be operated, but the controls are difficult to learn and do not correspond to actual actions

Engineering Contradiction:
Improveease of controlVSAvoidcontrol learning difficulty
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent creates a virtual model that copies the user's physical body and movements. The system captures real-world user movements and replicates them in a virtual skeletal model, allowing natural physical actions to directly control application functions without requiring learning complex control schemes

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces traditional mechanical input devices (keyboards, mice, controllers) with a vision-based motion capture system. Depth cameras and image processing algorithms substitute for physical controls, translating natural body movements directly into digital commands

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If depth images are processed at high resolution, then accurate extremity detection is achieved, but processing time and computational resources increase

Engineering Contradiction:
Improveextremity detection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent divides the depth image into multiple blocks or regions and processes each block independently to generate voxels. This segmentation allows parallel processing of different image regions, maintaining detection accuracy while reducing overall processing time through distributed computation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms 2D depth image data into a 3D voxel representation. By adding a depth dimension and organizing data spatially as voxels, the system enables more efficient processing and analysis of extremity positions while maintaining high measurement precision

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS9582717B2Systems and methods for tracking a model
Publication Date: 2017.02.28 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9582717B2 patent drawing
  • US9582717B2 patent drawing
  • US9582717B2 patent drawing

AI summary

An image such as a depth image of a scene may be received, observed, or captured by a device. A grid of voxels may then be generated based on the depth image such that the depth image may be downsampled. A model may be adjusted based on a location or position of one or more extremities estimated or determined for a human target in the grid of voxels. The model may also be adjusted based on a default location or position of the model in a default pose such as a T-pose, a DaVinci pose, and/or a natural pose.