Skeletal Tracking via Temporal Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Virtual reality and augmented reality systems face challenges in accurately rendering avatars without depth sensors, leading to increased costs and complexity, and existing skeletal tracking methods are noisy and resource-intensive.
Innovation Solution
The use of machine learning techniques to identify and filter skeletal joints from a single image and predict positions in subsequent frames, allowing for the generation of virtual objects without depth maps, enabling efficient rendering on devices with simple RGB cameras.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If depth sensors are used to capture depth maps for avatar generation, then the accuracy of virtual object rendering is improved, but the device complexity and cost increase
Solution Approach 1:
The patent extracts and removes the depth sensor component from the system, achieving accurate depth information without requiring physical depth sensors. Instead, it uses monocular RGB images combined with machine learning techniques to infer depth and skeletal joint positions, thereby reducing device complexity while maintaining measurement precision.
Solution Approach 2:
The patent replaces the mechanical/optical depth sensing system with a computational approach using machine learning models. The system substitutes physical depth sensors with algorithms that process RGB images to estimate depth and skeletal positions, eliminating the need for complex hardware while achieving similar functional outcomes.
2Reliability
If traditional skeletal tracking methods are used to identify body pose, then the tracking coverage is comprehensive, but the computational resources required increase significantly
Solution Approach 1:
The patent applies preliminary action by pre-training machine learning models offline to perform skeletal joint detection and depth estimation. During runtime, the pre-trained models process RGB images efficiently without requiring complex real-time computations, thereby reducing computational resource consumption while maintaining tracking accuracy.
Solution Approach 2:
The patent uses copying by training machine learning models on extensive datasets of labeled skeletal positions and depth information. The models learn to copy and generalize skeletal tracking patterns from training data, enabling accurate pose estimation without requiring resource-intensive real-time processing during actual use.
3Measurement precision
If machine learning techniques are applied to predict skeletal joint positions from previous frames, then the tracking precision is improved, but the processing time increases
Solution Approach 1:
The patent uses preliminary action by pre-processing and pre-training machine learning models to perform skeletal detection efficiently. The models are trained offline to recognize skeletal patterns quickly, enabling real-time or near-real-time processing during actual use without sacrificing accuracy.
Solution Approach 2:
The patent applies periodic action by using temporal information from previous frames in a sequential manner. The machine learning model leverages periodic patterns in human motion by comparing current frame predictions with previous frame results, improving accuracy through temporal consistency while maintaining efficient processing through frame-by-frame analysis.
Data Source
AI summary
Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing a program and a method for detecting a pose of a user. The program and method include operations comprising receiving a monocular image that includes a depiction of a body of a user; detecting a plurality of skeletal joints of the body based on the monocular image; accessing a video feed comprising a plurality of monocular images received prior to the monocular image; filtering, using the video feed, the plurality of skeletal joints of the body detected based on the monocular image; and determining a pose represented by the body depicted in the monocular image based on the filtered plurality of skeletal joints of the body.


