Head Pose Tracking Using Depth Camera and Inertial Sensors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing head pose tracking methods face challenges in achieving accurate and robust tracking, particularly in environments with limited visual features, where inertial sensors can lead to significant drift and require extensive computational resources, and optical flow tracking may result in drift and ambiguity issues.

Innovation Solution

A system employing a depth sensor apparatus and a conventional color video camera, optionally combined with inertial sensors, uses synchronized color and depth image sequences, keyframe matching, and transformation matrix estimation to correct for drift and ambiguity, employing optical flow tracking and depth correction modules to refine matching points and reduce outliers, with the option of using a recursive Bayesian framework and Extended Kalman Filter for fusion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If inertial sensors are used for head pose tracking, then tracking can be performed without environment instrumentation, but significant drift occurs and extensive computational resources are required

Engineering Contradiction:
Improvetracking without environment instrumentationVSAvoidtracking accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent combines inertial sensors with a depth camera system to merge the advantages of both approaches. The inertial sensors provide continuous tracking data while the depth camera periodically corrects drift through 3D feature matching, achieving reliable tracking without requiring environment instrumentation beyond the depth sensor.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The depth camera acts as an intermediary correction mechanism that periodically calibrates the inertial sensor drift. By using 3D feature matching in depth images, the system resets accumulated errors without requiring continuous complex computations from the inertial sensors alone.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If optical flow tracking is used, then computational resources are reduced, but drift and ambiguity issues occur

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidtracking accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent transitions from 2D optical flow tracking to 3D feature matching using depth images. This dimensional change provides explicit depth information that resolves ambiguities inherent in 2D tracking and reduces drift, while maintaining computational efficiency through selective 3D processing rather than continuous high-computation methods.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If environment instrumentation with markers is used, then head pose tracking accuracy is improved, but the system complexity and setup requirements increase

Engineering Contradiction:
Improvehead pose tracking accuracyVSAvoidenvironment instrumentation
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The depth camera captures 3D information of the natural environment without requiring externally placed markers or instruments. The system uses inherent 3D features in the scene to perform feature matching and track head pose, making the environment self-sufficient for tracking purposes.

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP2813082B1Head pose tracking using a depth camera
Publication Date: 2018.03.28 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP2813082B1 patent drawingFigure 1
  • EP2813082B1 patent drawingFigure 2~3
  • EP2813082B1 patent drawingFigure 4A

AI summary

Head pose tracking technique embodiments are presented that use a group of sensors configured so as to be disposed on a user's head. This group of sensors includes a depth sensor apparatus used to identify the three dimensional locations of features within a scene, and at least one other type of sensor. Data output by each sensor in the group of sensors is periodically input, and each time the data is input it is used to compute a transformation matrix that when applied to a previously determined head pose location and orientation established when the first sensor data was input identifies a current head pose location and orientation. This transformation matrix is then applied to the previously determined head pose location and orientation to identify a current head pose location and orientation.