Voxel Grid Skeletal Model Tracking for Intuitive Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computing applications face challenges in providing intuitive controls for users, as conventional input methods can be difficult to learn and may not directly correspond to actual actions, creating a barrier between users and applications.
Innovation Solution
A system and method for tracking a user in a scene using depth images to generate a grid of voxels, isolate foreground objects, determine extremities, and adjust a skeletal model to mimic user movements, allowing for more natural and intuitive control inputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional input controls (keyboards, mice, controllers) are used, then applications can be operated, but the controls are difficult to learn and do not correspond to actual actions
Solution Approach 1:
The patent creates a virtual model that copies the user's physical body and movements. The system captures real-world user movements and replicates them in a virtual skeletal model, allowing natural physical actions to directly control application functions without requiring learning complex control schemes
Solution Approach 2:
The patent replaces traditional mechanical input devices (keyboards, mice, controllers) with a vision-based motion capture system. Depth cameras and image processing algorithms substitute for physical controls, translating natural body movements directly into digital commands
2Measurement precision
If depth images are processed at high resolution, then accurate extremity detection is achieved, but processing time and computational resources increase
Solution Approach 1:
The patent divides the depth image into multiple blocks or regions and processes each block independently to generate voxels. This segmentation allows parallel processing of different image regions, maintaining detection accuracy while reducing overall processing time through distributed computation
Solution Approach 2:
The patent transforms 2D depth image data into a 3D voxel representation. By adding a depth dimension and organizing data spatially as voxels, the system enables more efficient processing and analysis of extremity positions while maintaining high measurement precision
Data Source
AI summary
An image such as a depth image of a scene may be received, observed, or captured by a device. A grid of voxels may then be generated based on the depth image such that the depth image may be downsampled. A model may be adjusted based on a location or position of one or more extremities estimated or determined for a human target in the grid of voxels. The model may also be adjusted based on a default location or position of the model in a default pose such as a T-pose, a DaVinci pose, and/or a natural pose.


