Egocentric Environmental Embeddings for XR Body Pose Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing XR devices struggle to accurately capture and estimate a user's full-body posture due to limited camera capabilities, often relying on egocentric video that obscures the user's body pose, necessitating techniques that require multiple cameras or direct video capture of body parts.

Innovation Solution

Estimating a user's body pose from egocentric video of the environment using environmental feature maps and embeddings, fused with head-pose estimates from motion sensors, through a transformer-based model that integrates static, dynamic, and interactee environmental data to predict a comprehensive 3D user model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple cameras are used to capture different parts of the user's body, then measurement precision of body pose is improved, but device complexity increases

Engineering Contradiction:
Improvebody pose estimation accuracyVSAvoidnumber of cameras
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

Instead of directly capturing the user's body with cameras, the patent inverts the approach by capturing the environment and inferring body pose from environmental features. The system uses egocentric camera views of the surrounding environment, not direct views of the body, to estimate full-body posture through environmental context analysis.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent introduces environmental feature maps and embeddings as intermediary representations between the camera input and body pose estimation. These environmental embeddings serve as a mediator that connects egocentric video data to body pose predictions, enabling indirect inference of body posture from environmental context.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If egocentric cameras are used in XR headsets, then device simplicity is maintained, but measurement precision of body pose deteriorates due to blind spots

Engineering Contradiction:
Improvecamera system simplicityVSAvoidbody pose estimation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent inverts the traditional approach by not directly observing the body with cameras but instead inferring body pose from environmental features captured by egocentric cameras. This allows the use of simple single-camera systems while maintaining pose estimation capability through environmental context.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The system uses the egocentric camera's inherent environmental capture capability to serve the additional function of body pose estimation. The same camera that captures the user's field of view also provides environmental data that can be processed to infer body posture, making the simple camera system multi-functional.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If direct video capture of body parts is used, then measurement precision is improved, but adaptability to XR device constraints deteriorates

Engineering Contradiction:
Improvebody pose measurement accuracyVSAvoidcompatibility with XR devices
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent makes the solution universal by adapting body pose estimation to work with the constraints of XR devices. The method uses environmental data that is naturally captured by XR head-mounted cameras, making the system compatible with XR devices without requiring specialized body-capturing hardware. The same camera system serves both XR rendering and pose estimation functions.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12613573B2Determining body pose from environmental data
Publication Date: 2026.04.28 SAMSUNG ELECTRONICS CO LTD
  • US12613573B2 patent drawing
  • US12613573B2 patent drawing
  • US12613573B2 patent drawing

AI summary

In one embodiment, a method includes accessing video captured by one or more cameras of a head-mounted device (HMD) worn by a user and determining, from the accessed video, multiple environmental feature maps and multiple corresponding environmental embeddings representing an environment of the user captured in the accessed video. The method further includes fusing, by multiple trained environmental transformer models, the environmental embeddings to create a fused environmental embedding; determining a head pose of the user coincident with the accessed video; and predicting, by a trained decoder and based on the fused environmental embedding and the determined head pose of the user, a body pose of the user.