Depth Camera Human Tracking via Skeletal Model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computing applications that use human gestures for input are often complex and require specialized sensors, making it difficult for users to interact naturally with games and other applications, as the controls do not correspond to actual movements and require post-processing for accurate interpretation.
Innovation Solution
A compact device that recognizes and tracks humans in three-dimensional space without special sensors, using a capture device with depth cameras to generate a real-time multi-point skeletal model, allowing users to interact with applications through natural movements and voice commands, and providing a direct representation for a wide range of applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If traditional controls (keyboards, mice, controllers) are used to manipulate game characters, then users can interact with applications, but the controls are difficult to learn and do not correspond to actual movements
Solution Approach 1:
The patent replaces traditional mechanical input devices (keyboards, mice, controllers) with a vision-based system that uses depth cameras and image processing to detect and interpret human gestures. The capture device captures images of the user's movements, and the processor translates these visual inputs into control commands, eliminating the need for physical controllers and creating a more intuitive, natural interface that directly maps user actions to game character movements.
2Measurement precision
If complex sensor systems are used to track humans in three dimensional space, then accurate tracking is achieved, but the system becomes complex and requires post-processing
Solution Approach 1:
The patent extracts and utilizes the depth information component from multi-camera setups, focusing specifically on the Z-axis depth data to simplify the tracking system. By isolating and processing only the depth information rather than full color images, the system achieves accurate three-dimensional tracking with reduced computational complexity and eliminates the need for complex post-processing of complete image data.
Solution Approach 2:
The patent creates a simplified skeletal model representation that copies only the essential structural information from the captured human form. This skeletal model includes key body points and segment connections that replicate the necessary movement data without requiring processing of the entire complex image data, enabling efficient real-time tracking with reduced computational requirements.
3Adaptability or versatility
If multiple people are to be tracked and identified, then the system must recognize and distinguish individuals, but this increases processing complexity
Solution Approach 1:
The patent performs preliminary identification and assignment of unique identifiers to each detected user before tracking begins. The system detects multiple users in the scene, assigns each a unique ID, and establishes their initial skeletal models in advance. This preliminary action allows the tracking algorithm to focus on maintaining and updating pre-established user identities rather than continuously identifying and re-identifying individuals, significantly reducing processing complexity during ongoing tracking.
Data Source
AI summary
A system recognizes human beings in their natural environment, without special sensing devices attached to the subjects, uniquely identifies them and tracks them in three dimensional space. The resulting representation is presented directly to applications as a multi-point skeletal model delivered in real-time. The device efficiently tracks humans and their natural movements by understanding the natural mechanics and capabilities of the human muscular-skeletal system. The device also uniquely recognizes individuals in order to allow multiple people to interact with the system via natural movements of their limbs and body as well as voice commands/responses.


