Depth Camera Gesture Recognition Without Markers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies are unable to accurately interpret human movements without the use of special reflective tags or markers, limiting the ability of computers to assess and respond to human gestures in a natural and intuitive manner.
Innovation Solution
A depth camera system that models human movements using virtual skeletons, allowing users to control interactive interfaces through physical gestures, such as spell-casting gestures, without the need for conventional controllers or markers, by capturing and interpreting depth images to recognize and translate body movements into machine-readable commands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If reflective tags or markers are used to track human movements, then measurement precision is improved, but device complexity and ease of operation deteriorate due to requiring special equipment and preparation
Solution Approach 1:
The patent extracts the tracking markers from the system by using a depth camera to directly capture human body geometry and movements without requiring reflective tags or external markers. The depth camera captures unassisted human movements by analyzing depth information from the human body itself, eliminating the need for special equipment attachment.
Solution Approach 2:
The patent introduces virtual skeletons as an intermediary representation between the depth camera data and the gesture recognition system. The virtual skeleton model serves as a mediator that maps real human movements to interpretable gesture commands, enabling accurate movement tracking without physical markers.
2Measurement precision
If multiple cameras and tracking tags are used to achieve accurate 3D position triangulation, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent eliminates the need for multiple cameras and external tracking tags by using a single depth camera to capture three-dimensional depth information directly. The depth camera provides volumetric data that enables 3D position calculation without requiring triangulation from multiple camera viewpoints or external markers.
Solution Approach 2:
The patent replaces the mechanical triangulation system (multiple cameras and physical tags) with an optical depth sensing system. The depth camera uses light time-of-flight or phase-shift measurement to directly obtain 3D position information, substituting complex mechanical tracking infrastructure with a more compact optical sensing approach.
3Measurement precision
If conventional controllers or markers are required for interaction, then measurement precision is maintained, but ease of operation and adaptability deteriorate
Solution Approach 1:
The patent creates a universal gesture-based interface that can control multiple applications and games without requiring application-specific controllers or markers. The depth camera system recognizes a variety of gestures (spelling gestures, aiming gestures, casting gestures) that can be mapped to different functions across diverse applications, providing a single versatile control mechanism.
Solution Approach 2:
The system enables users to interact directly with the computer interface using their natural body movements without requiring external controllers. The depth camera captures and interprets gestures performed with the user's own body, allowing the interface to serve itself through natural human-computer interaction rather than requiring mediating control devices.
Data Source
AI summary
Gestures of a computer user are observed with a depth camera. A first gesture of the computer user is identified as one of a plurality of different action selection gestures, each action selection gesture associated with a different action performable within an interactive interface controlled by gestures of the computer user. A second gesture is identified as a triggering gesture that causes performance of the action associated with the action selection gesture.


