Pointing Vector Determination for Depth Camera Gestures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing camera systems struggle to accurately determine pointing vectors for hand gestures, especially when hands are close to or far from the camera, leading to unreliable depth data and limited functionality in virtual reality and augmented reality applications.
Innovation Solution
A method using a 3D camera system that defines a minimum Z distance (minZ) for reliable depth data, employing a 2D convolution filter to isolate the hand and find the fingertip position, and combining depth data with the camera's location to determine the pointing vector, even when hands are close to the camera by utilizing disparity between images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a depth camera is used to track hand gestures, then hand tracking capability is improved, but measurement precision deteriorates when hands are close to or far from the camera
Solution Approach 1:
The patent introduces an intermediary computational process that combines multiple data sources (depth camera data, RGB camera data, and hand model predictions) to mediate the unreliable depth measurements. When the hand is out of the reliable depth range, the system uses the hand model and RGB image data as intermediaries to infer hand position and orientation, thereby maintaining measurement precision across all distances.
Solution Approach 2:
The system dynamically changes parameters based on depth reliability. It monitors the hand's distance from the camera and switches between using raw depth data (when reliable) and using computed estimates from hand models and image processing (when unreliable). This parameter switching resolves the contradiction by adapting the measurement approach to the current depth conditions.
2Ease of operation
If the hand is placed close to the camera for natural interaction, then ease of operation is improved, but measurement precision deteriorates due to unreliable depth data
Solution Approach 1:
The system prepares in advance for potential depth reliability issues by having a fallback hand model and computation pipeline ready. When the hand enters the unreliable near-field zone, the system seamlessly transitions to using the hand model that predicts hand geometry and pose from RGB images, cushioning the user from any degradation in tracking accuracy while maintaining natural close-proximity interaction.
Solution Approach 2:
The hand model acts as an intermediary that translates unreliable depth information and RGB image data into accurate hand pose estimates. This intermediary computation allows users to interact naturally at close distances without suffering from the depth camera's reliability limitations.
3Device complexity
If a simple camera system is used, then device complexity is reduced, but measurement precision and functionality are limited
Solution Approach 1:
The patent segments the hand tracking problem into multiple independent components: depth camera processing, RGB camera processing, hand model prediction, and fusion/computation. Each segment can be processed independently and combined to achieve high precision. This segmentation allows the use of simpler individual camera systems while achieving complex measurement goals through coordinated processing of multiple data streams.
Solution Approach 2:
The hand model serves multiple functions: it predicts hand geometry, determines pose, estimates position, and provides fallback tracking when depth data is unreliable. This multi-functionality allows a relatively simple camera system to achieve sophisticated measurement capabilities that would otherwise require complex specialized hardware.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables accurate and consistent determination of pointing vectors across a wide range of distances, enhancing user interaction comfort and accuracy in hand gesture systems, particularly in head-mounted displays and portable devices.
Implementation Method 1
Other camera systems use a rangefinder or proximity sensor either for particular points in the image or for the whole image such as a time-of-flight camera
Implementation Method 2
A method using a 3D camera system that defines a minimum Z distance (minZ) for reliable depth data, employing a 2D convolution filter to isolate the hand and find the fingertip position
Implementation Method 3
combining depth data with the camera's location to determine the pointing vector, even when hands are close to the camera by utilizing disparity between images
Data Source
AI summary
A pointing vector is determined for a gesture that is performed before a depth camera. One example includes receiving a first and a second image of a pointing gesture in a depth camera, the depth camera having a first and a second image sensor, applying erosion and dilation to the first image using a 2D convolution filter to isolate the gesture from other objects, finding the imaged gesture in the filtered first image of the camera, finding a pointing tip of the imaged gesture, determining a position of the pointing tip of the imaged gesture using the second image, and determining a pointing vector using the determined position of the pointing tip.


