Egocentric Environmental Embeddings for XR Body Pose Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing XR devices struggle to accurately capture and estimate a user's full-body posture due to limited camera capabilities, often relying on egocentric video that obscures the user's body pose, necessitating techniques that require multiple cameras or direct video capture of body parts.
Innovation Solution
Estimating a user's body pose from egocentric video of the environment using environmental feature maps and embeddings, fused with head-pose estimates from motion sensors, through a transformer-based model that integrates static, dynamic, and interactee environmental data to predict a comprehensive 3D user model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple cameras are used to capture different parts of the user's body, then measurement precision of body pose is improved, but device complexity increases
Solution Approach 1:
Instead of directly capturing the user's body with cameras, the patent inverts the approach by capturing the environment and inferring body pose from environmental features. The system uses egocentric camera views of the surrounding environment, not direct views of the body, to estimate full-body posture through environmental context analysis.
Solution Approach 2:
The patent introduces environmental feature maps and embeddings as intermediary representations between the camera input and body pose estimation. These environmental embeddings serve as a mediator that connects egocentric video data to body pose predictions, enabling indirect inference of body posture from environmental context.
2Device complexity
If egocentric cameras are used in XR headsets, then device simplicity is maintained, but measurement precision of body pose deteriorates due to blind spots
Solution Approach 1:
The patent inverts the traditional approach by not directly observing the body with cameras but instead inferring body pose from environmental features captured by egocentric cameras. This allows the use of simple single-camera systems while maintaining pose estimation capability through environmental context.
Solution Approach 2:
The system uses the egocentric camera's inherent environmental capture capability to serve the additional function of body pose estimation. The same camera that captures the user's field of view also provides environmental data that can be processed to infer body posture, making the simple camera system multi-functional.
3Measurement precision
If direct video capture of body parts is used, then measurement precision is improved, but adaptability to XR device constraints deteriorates
Solution Approach 1:
The patent makes the solution universal by adapting body pose estimation to work with the constraints of XR devices. The method uses environmental data that is naturally captured by XR head-mounted cameras, making the system compatible with XR devices without requiring specialized body-capturing hardware. The same camera system serves both XR rendering and pose estimation functions.
Data Source
AI summary
In one embodiment, a method includes accessing video captured by one or more cameras of a head-mounted device (HMD) worn by a user and determining, from the accessed video, multiple environmental feature maps and multiple corresponding environmental embeddings representing an environment of the user captured in the accessed video. The method further includes fusing, by multiple trained environmental transformer models, the environmental embeddings to create a fused environmental embedding; determining a head pose of the user coincident with the accessed video; and predicting, by a trained decoder and based on the fused environmental embedding and the determined head pose of the user, a body pose of the user.


