Human Tracking System Using Depth Image Voxel Grids
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computing applications face challenges in providing intuitive controls for users, as conventional input methods like controllers and keyboards can be difficult to learn and may not directly correspond to actual actions, creating a barrier between users and applications.
Innovation Solution
A system that tracks a user in a scene using depth images to generate a grid of voxels, isolates foreground objects, determines extremities, and adjusts a skeletal model to map user movements onto avatars or game characters, allowing natural body motions to control applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional input controls (controllers, keyboards, mice) are used, then application functionality is achieved, but user learning difficulty increases and intuitiveness decreases
Solution Approach 1:
The patent creates a virtual copy of the user's physical body (avatar) that mirrors real-world movements. Instead of learning abstract control schemes, users directly manipulate their digital representation through natural body motions, making the interaction intuitive and eliminating the learning curve associated with conventional controls
Solution Approach 2:
The patent replaces mechanical input devices (controllers, keyboards, mice) with a vision-based motion tracking system. Depth cameras capture user movements and automatically translate them into application commands, substituting physical manipulation of controls with natural gestural interaction
2Ease of operation
If conventional controls are used, then application actions are performed, but correspondence between control and actual action is lost
Solution Approach 1:
The system creates a virtual skeleton model that copies the user's anatomical structure and movement patterns. This skeletal representation maintains direct correspondence between physical body actions and their digital interpretation, allowing users to perform natural movements that directly translate to application actions without abstract mapping
Solution Approach 2:
The motion tracking system serves multiple functions simultaneously: it captures user movements, tracks them over time, maps them to skeletal joints, and translates them into application commands. This multi-functional approach eliminates the need for separate control mechanisms while maintaining direct action correspondence
Data Source
Figure 1A
Figure 1B
Figure 2
AI summary
An image such as a depth image of a scene may be received, observed, or captured by a device. A grid of voxels may then be generated based on the depth image such that the depth image may be downsampled. A background included in the grid of voxels may also be removed to isolate one or more voxels associated with a foreground object such as a human target. A location or position of one or more extremities of the isolated human target may be determined and a model may be adjusted based on the location or position of the one or more extremities.