AI Head Pose Estimation Using 2D-3D Sensor Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies require expensive 3D cameras to recognize a person's pose, necessitating a solution to infer 3D head pose information from 2D images using more accessible and cost-effective methods.

Innovation Solution

An AI apparatus and method that utilize a 2D image sensor to acquire images, a 3D image sensor to gather pose information, and a processor to match and correct data, extracting 3D head pose information by synchronizing and correlating 2D and 3D coordinates, enabling the inference of head rotation and position without the need for expensive 3D cameras.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a 3D camera is used to recognize a person's pose, then measurement precision of head pose is improved, but device cost increases

Engineering Contradiction:
Improvehead pose recognition accuracyVSAvoiddevice cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent creates a mapping relationship between 2D image coordinates and 3D head pose coordinates, allowing the system to copy the functionality of expensive 3D pose recognition into a 2D image processing framework. By training a neural network to learn the correspondence between 2D landmarks and 3D pose parameters, the system replicates 3D camera capabilities using affordable 2D imaging equipment.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces expensive, complex 3D camera systems with inexpensive 2D cameras. The 2D camera system, combined with computational algorithms, serves as a cost-effective alternative that achieves comparable pose estimation accuracy without requiring specialized expensive hardware.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

2Ease of manufacture

If a 2D camera is used to infer 3D head pose information, then device cost is reduced, but measurement precision deteriorates

Engineering Contradiction:
Improvedevice costVSAvoidhead pose recognition accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent transforms the problem by changing the parameter space - instead of directly measuring 3D pose parameters with expensive sensors, the system uses 2D image parameters (landmark coordinates) and transforms them through a learned mapping function. The neural network learns optimal parameter transformations that recover 3D pose information from 2D projections, effectively changing how the measurement is performed rather than changing the hardware.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the mechanical/optical 3D sensing system with a computational approach. Instead of using multiple cameras or specialized 3D sensors to physically capture depth information, the system uses a single 2D camera combined with machine learning algorithms to computationally infer 3D pose, substituting physical measurement mechanisms with intelligent data processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If 3D head pose information is corrected using 2D landmark coordinates, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improve3D head pose accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary actions by pre-training a neural network on large datasets of corresponding 2D images and 3D pose annotations. This pre-training phase learns the complex mapping relationships between 2D landmarks and 3D pose parameters, so that during actual operation, the system can directly apply the learned model without performing complex real-time optimization or calibration procedures.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary neural network model that mediates between 2D image inputs and 3D pose outputs. This intermediary learns the complex transformation relationships during training and serves as a bridge that simplifies the inference process, converting the difficult direct mapping problem into a straightforward forward propagation through a trained network.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11288840B2Artificial intelligence apparatus for estimating pose of head and method for the same
Publication Date: 2022.03.29 LG ELECTRONICS INC
  • US11288840B2 patent drawing
  • US11288840B2 patent drawing
  • US11288840B2 patent drawing

AI summary

Disclosed is an artificial intelligence (AI) apparatus including a two-dimensional (2D) image sensor configured to acquire a 2D image of a head of a person, a three-dimensional (3D) image sensor configured to acquire 3D head pose information of the head, and a processor configured to match the 2D image with the 3D head pose information, to extract 3D head pose information for determining a rotation direction of the head from the 3D head pose information, to extract a 2D image matched with the extracted 3D head pose information, to acquire 3D relative coordinates as a reference for correcting the 3D head pose information based on 2D coordinates of a predetermined landmark point of the extracted 2D image, and to acquire the corrected 3D head pose information of the predetermined landmark point of each 2D image by correcting the 3D head pose information based on the 3D relative coordinates.