Self-Tracking Controller Cameras for Full-Body Pose Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing AR/VR systems struggle to accurately estimate a user's full body pose due to limitations in camera visibility, particularly with HMD cameras unable to capture lower body parts like legs and feet, leading to incomplete pose estimation that affects user experience.

Innovation Solution

Utilizing self-tracking controllers with integrated cameras and IMUs for SLAM to capture images of body parts not visible to HMD cameras, combining keypoints from both HMD and controller cameras through an inverse-kinematic optimizer, and employing muscular-skeletal models and ML models for accurate full body pose estimation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If HMD cameras are used to track body parts, then upper body tracking is improved, but lower body visibility is lost

Engineering Contradiction:
Improveupper body tracking accuracyVSAvoidlower body visibility
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system divides the body tracking task into two segments: HMD cameras handle upper body tracking while controller cameras handle lower body tracking. This segmentation allows each camera system to specialize in the body regions it can best capture, with the HMD optimized for head and upper body visibility and controllers positioned to capture legs and feet that are otherwise invisible to the HMD cameras.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If multiple controllers with cameras are added to capture lower body, then full body pose estimation is improved, but device complexity increases

Engineering Contradiction:
Improvefull body pose estimation accuracyVSAvoidsystem configuration complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The controllers serve multiple functions: they are used for interaction input and simultaneously function as mobile cameras for lower body tracking. By making the controllers multi-functional, the system avoids adding dedicated lower body cameras, thereby improving full body pose estimation without proportionally increasing device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The controllers perform self-tracking using their integrated cameras and IMUs to determine their own pose and location in 3D space. This self-service capability eliminates the need for additional external tracking infrastructure, allowing the controllers to independently provide the imaging function needed for lower body capture while managing their own spatial awareness.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If controller cameras are used for self-localization, then lower body tracking is improved, but processing load increases

Engineering Contradiction:
Improvelower body keypoint detectionVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The controllers perform SLAM (Simultaneous Localization and Mapping) to pre-determine their own pose and location in 3D space before being used for body tracking. This preliminary self-localization action allows the system to establish a known reference frame for the controller cameras, reducing the computational burden during actual body tracking by avoiding the need to simultaneously solve for both controller position and body pose.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250321649A1Body pose estimation using self-tracked controllers
Publication Date: 2025.10.16 META PLATFORMS TECHNOLOGIES LLC
  • US20250321649A1 patent drawing
  • US20250321649A1 patent drawing
  • US20250321649A1 patent drawing

AI summary

In one embodiment, a computing system may determine a pose of a device held by or attached to a hand of a user based on sensor data captured by the device. The system may determine a pose of a headset worn by the user based on sensor data captured by the headset. The system may detemline positions of a first set of keypoints associated with a first portion of a body of the user based on (1) one or more first images captured by one or more cameras of the device, (2) the pose of the device, (3) one or more second images captured by one or more cameras of the headset and (4) the pose of the headset. The system may determine a body pose of the user based at least on the positions of the first set of keypoints.