6D Pose Estimation for AR via User Keypoints and Iterative Refinement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for determining 6D pose estimates of physical 3D objects in augmented reality are overly complex for larger objects and lack efficiency in scaling to thousands of objects, particularly in AR tutorials where precise digital or virtual content placement is required.
Innovation Solution
The method involves user-input keypoints selection on a mobile device to generate an initial 6D pose estimate, which is then refined using a cost function and template matching, allowing for accurate placement of digital or virtual content in an AR scene, with the option to use machine learning for automatic keypoint estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current methods are used for determining 6D pose estimates, then measurement precision can be achieved, but device complexity and computational overhead become excessive for larger objects
Solution Approach 1:
The patent segments the 6D pose estimation process into two distinct phases: (1) an initial estimation phase using a simplified algorithm to obtain rough pose parameters, and (2) a refinement phase using iterative optimization to improve accuracy. This segmentation divides the complex single-step process into manageable stages, reducing overall computational complexity while maintaining precision.
Solution Approach 2:
The patent performs preliminary action by first obtaining an initial 6D pose estimate using a computationally efficient method before applying more intensive refinement algorithms. This preliminary estimation provides a good starting point that reduces the search space and computational burden of subsequent optimization steps, effectively preparing the system in advance to avoid more complex processing later.
2Measurement precision
If current methods are used for determining 6D pose estimates, then measurement precision can be achieved, but productivity and scalability to thousands of objects are reduced
Solution Approach 1:
The patent segments the processing workflow into an initial fast estimation stage and a subsequent refinement stage, allowing systems to adjust the level of refinement based on productivity requirements. For high-volume object processing, the system can rely more on the initial estimation with less refinement, thereby increasing throughput while maintaining acceptable accuracy levels.
Solution Approach 2:
The patent applies partial action by allowing the refinement process to be optional or adjustable in intensity. Instead of always applying full refinement to every object, the system can apply refinement selectively or to a limited degree, achieving sufficient accuracy for many applications without the full computational cost, thus improving overall processing productivity.
3Ease of operation
If user-input keypoints are used, then ease of operation is improved, but loss of time increases due to manual input requirements
Solution Approach 1:
The patent performs preliminary action by automatically detecting and selecting keypoint locations on the object in the image before the user needs to interact with the system. This preliminary automatic keypoint identification prepares the data in advance, allowing the user to simply confirm or make minor adjustments rather than manually specifying each keypoint, thereby reducing the time users spend on this task while maintaining ease of operation.
Data Source
AI summary
Embodiments include systems and methods for determining a 6D pose estimate associated with an image of a physical 3D object captured in a video stream. An initial 6D pose estimate is inferred and then further iteratively refined. The video stream may be frozen to allow the user to tap or touch a display to indicate a location of the user-input keypoints. The resulting 6D pose estimate is used to assist in replacing or superimposing the physical 3D object with digital or virtual content in an augmented reality (AR) frame.


