Pointing Vector Determination for Depth Camera Gestures

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing camera systems struggle to accurately determine pointing vectors for hand gestures, especially when hands are close to or far from the camera, leading to unreliable depth data and limited functionality in virtual reality and augmented reality applications.

Innovation Solution

A method using a 3D camera system that defines a minimum Z distance (minZ) for reliable depth data, employing a 2D convolution filter to isolate the hand and find the fingertip position, and combining depth data with the camera's location to determine the pointing vector, even when hands are close to the camera by utilizing disparity between images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a depth camera is used to track hand gestures, then hand tracking capability is improved, but measurement precision deteriorates when hands are close to or far from the camera

Engineering Contradiction:
Improvehand tracking capabilityVSAvoiddepth data accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary computational process that combines multiple data sources (depth camera data, RGB camera data, and hand model predictions) to mediate the unreliable depth measurements. When the hand is out of the reliable depth range, the system uses the hand model and RGB image data as intermediaries to infer hand position and orientation, thereby maintaining measurement precision across all distances.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system dynamically changes parameters based on depth reliability. It monitors the hand's distance from the camera and switches between using raw depth data (when reliable) and using computed estimates from hand models and image processing (when unreliable). This parameter switching resolves the contradiction by adapting the measurement approach to the current depth conditions.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If the hand is placed close to the camera for natural interaction, then ease of operation is improved, but measurement precision deteriorates due to unreliable depth data

Engineering Contradiction:
Improveuser interaction naturalnessVSAvoiddepth data reliability
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system prepares in advance for potential depth reliability issues by having a fallback hand model and computation pipeline ready. When the hand enters the unreliable near-field zone, the system seamlessly transitions to using the hand model that predicts hand geometry and pose from RGB images, cushioning the user from any degradation in tracking accuracy while maintaining natural close-proximity interaction.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Solution Approach 2:

The hand model acts as an intermediary that translates unreliable depth information and RGB image data into accurate hand pose estimates. This intermediary computation allows users to interact naturally at close distances without suffering from the depth camera's reliability limitations.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Device complexity

If a simple camera system is used, then device complexity is reduced, but measurement precision and functionality are limited

Engineering Contradiction:
Improvecamera system structureVSAvoidpointing vector accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent segments the hand tracking problem into multiple independent components: depth camera processing, RGB camera processing, hand model prediction, and fusion/computation. Each segment can be processed independently and combined to achieve high precision. This segmentation allows the use of simpler individual camera systems while achieving complex measurement goals through coordinated processing of multiple data streams.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The hand model serves multiple functions: it predicts hand geometry, determines pose, estimates position, and provides fallback tracking when depth data is unreliable. This multi-functionality allows a relatively simple camera system to achieve sophisticated measurement capabilities that would otherwise require complex specialized hardware.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables accurate and consistent determination of pointing vectors across a wide range of distances, enhancing user interaction comfort and accuracy in hand gesture systems, particularly in head-mounted displays and portable devices.

Implementation Method 1

Other camera systems use a rangefinder or proximity sensor either for particular points in the image or for the whole image such as a time-of-flight camera

Methodology Applied
Scientific EffectTime of flight: Time of Flight

Implementation Method 2

A method using a 3D camera system that defines a minimum Z distance (minZ) for reliable depth data, employing a 2D convolution filter to isolate the hand and find the fingertip position

Methodology Applied
Scientific EffectImage processing: Image Processing

Implementation Method 3

combining depth data with the camera's location to determine the pointing vector, even when hands are close to the camera by utilizing disparity between images

Methodology Applied
Scientific EffectDisparity: Parallax

Data Source

PatentUS10607069B2Determining a pointing vector for gestures performed before a depth camera
Publication Date: 2020.03.31 INTEL CORP
  • US10607069B2 patent drawing
  • US10607069B2 patent drawing
  • US10607069B2 patent drawing

AI summary

A pointing vector is determined for a gesture that is performed before a depth camera. One example includes receiving a first and a second image of a pointing gesture in a depth camera, the depth camera having a first and a second image sensor, applying erosion and dilation to the first image using a 2D convolution filter to isolate the gesture from other objects, finding the imaged gesture in the filtered first image of the camera, finding a pointing tip of the imaged gesture, determining a position of the pointing tip of the imaged gesture using the second image, and determining a pointing vector using the determined position of the pointing tip.