XR Object Pose Tracking for Training Data Refinement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current extended reality (XR) systems face challenges in accurately and efficiently training object detection algorithms, as initial training processes are time-consuming and often produce lower quality data due to the need for manual image capture and alignment, or they rely on synthetic images lacking real-world features.
Innovation Solution
The method involves using initial training data based on synthetic images or limited real-world data, which is then updated with real-world images captured by the XR device once it successfully tracks the object, allowing for the extraction and incorporation of feature data to enhance the training data accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual training process is used to capture and align images, then training data quality is improved, but training time and effort increase significantly
Solution Approach 1:
The system performs preliminary action by using synthetic images to create initial training data before real object capture. This preliminary training data enables the XR device to begin tracking objects immediately, while real-world images are captured and processed in the background to refine the training data later, eliminating the need for time-consuming manual image capture and alignment.
Solution Approach 2:
The system creates copies of real objects using synthetic 3D models and renders synthetic images that replicate object appearances from multiple viewpoints. These synthetic copies serve as training data substitutes, allowing the system to train without capturing actual real-world images manually, thus reducing training time while maintaining sufficient training data quality.
2Productivity
If fewer images are captured to reduce training time, then training speed is improved, but training data quality deteriorates
Solution Approach 1:
The system transitions from 2D image capture to 3D modeling by creating three-dimensional models of objects and rendering synthetic images from multiple virtual viewpoints. This dimensional change allows comprehensive training data generation without requiring physical capture from numerous angles, maintaining training data quality while improving training speed.
Solution Approach 2:
The system changes parameters by using synthetic image generation instead of physical image capture. This parameter change enables the creation of training data with controlled variations in lighting, viewpoint, and object pose without being constrained by physical capture limitations, thus achieving both high training speed and maintained data quality.
3Loss of time
If synthetic images are used for training, then training time is reduced, but training data accuracy deteriorates due to missing real-world features
Solution Approach 1:
The system merges synthetic images and real-world images into a combined training dataset. Synthetic images provide the structural and geometric information, while real-world images contribute authentic texture, color, and environmental features. This combination maintains training data accuracy while avoiding the time-consuming manual capture process.
Solution Approach 2:
The system implements feedback by using real-world image capture to verify and refine the synthetic training data. The XR device captures real objects during operation, and these captured images are used to update and improve the synthetic models, ensuring the training data accurately reflects real-world appearances while maintaining the efficiency of synthetic generation.
Data Source
AI summary
A method includes acquiring, from the camera, a camera data sequence including a first image frame of a real object in a scene, tracking a pose of the real object with respect to the camera along the camera data sequence, and displaying an XR object on the display by rendering the XR object based at least on the pose. Flag data is set, in a memory area of the at least one memory, indicative of whether or not the displayed XR object is consistent in pose with the real object. The method includes outputting, to a separate computing device having another processor, second image frames in the camera data sequence acquired when the flag data indicates that the displayed XR object is consistent in pose with the real object.


