3D Hand Keypoint Generation for Mixed Reality Occlusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current mixed reality technologies face challenges in accurately capturing and displaying a user's hand movements, especially when multiple hand joints are occluded, leading to a less immersive experience due to inadequate representation of hand actions in virtual environments.
Innovation Solution
A method and system that utilize a trained 3D hand joint generative model to refine and synchronize the keypoints of occluded hand joints with visible hand joints, using random noise and iterative refinement, and render them using a 3D virtual hand modeler to enhance the accuracy and naturalness of hand movement representation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current mixed reality technology captures only visible hand joints, then the system complexity remains low, but the measurement precision of hand movements deteriorates when joints are occluded
Solution Approach 1:
The patent introduces a generative model as an intermediary component that receives visible hand joint data and generates predictions for occluded joints. This mediator processes the available information and produces complete hand pose estimates, resolving the contradiction by adding a sophisticated processing layer that improves measurement precision without requiring complex hardware changes
Solution Approach 2:
The patent replaces direct mechanical/optical capture of all hand joints with a computational approach using generative models. Instead of relying solely on physical sensors to detect every joint, the system uses machine learning models to infer occluded joint positions from visible ones, substituting mechanical detection with intelligent computation to maintain precision while managing system complexity
2Loss of information
If more hand joints are captured, then the completeness of hand movement representation improves, but the difficulty of detecting and measuring increases due to occlusion
Solution Approach 1:
The patent applies preliminary action by using the generative model to predict and fill in occluded joint information before the incomplete data is rendered or further processed. The model proactively generates estimates for missing joints based on visible joint configurations, ensuring information completeness is achieved through advance computational inference rather than attempting to directly detect all joints
Solution Approach 2:
The patent uses copying by creating virtual representations of occluded joints through the generative model. Instead of directly measuring difficult-to-capture joints, the system copies the spatial relationships and movement patterns from visible joints to generate accurate predictions for hidden joints, maintaining information completeness while avoiding the detection difficulties of direct measurement
3Manufacturing precision
If iterative refinement is performed to improve keypoint accuracy, then the manufacturing precision of hand pose estimation improves, but the loss of time increases due to multiple processing iterations
Solution Approach 1:
The patent implements periodic action through iterative refinement, where the generative model processes hand pose estimation through multiple sequential iterations. Each iteration refines the predictions further, progressively improving manufacturing precision. The periodic nature of these iterations allows the system to balance accuracy gains against time consumption, as each pass builds upon the previous one to converge toward more accurate results
Data Source
AI summary
According to one embodiment, a method, computer system, and computer program product for mixed reality is provided. The present invention may include receiving 3D hand keypoints (keypoints) of a user's visible hand joints from the user's capturable hand, and visible hand joints, if any, from the user's uncapturable hand; using random noise sampled with a unit normal distribution as initial keypoints for the uncapturable hand joints from the user's uncapturable hand; inputting the received and the initial keypoints, in a preset order, into a trained 3D hand joint generative model; performing an iterative refinement of the uncapturable hand joints from the user's uncapturable hand using the trained 3D hand joint generative model; identifying whether generated keypoints of the user's uncapturable hand are synchronized with the keypoints of the user's capturable hand; and rendering the generated 3D keypoints for the user's uncapturable hand joints using a 3D virtual hand modeler.


