AR Labeling via Aligned 3D CAD Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning technologies face inefficiencies in collecting pixel-level training data for object detection, relying on tedious manual annotation or complex algorithms, which are costly and time-consuming, and often compromise on precision and accuracy.
Innovation Solution
A system utilizing augmented reality (AR) technology and pre-existing CAD models to align and overlay 3D models onto real-world objects, enabling automatic pixel-level labeling by projecting 2D outlines onto images captured from various angles and conditions, reducing manual labor and increasing data diversity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual annotation is used to obtain pixel-level outlines, then precision and accuracy of labels are improved, but cost and time expenditure increase significantly
Solution Approach 1:
The patent uses 3D model projections to generate 2D outline annotations that are copied and applied to multiple images. Instead of manually annotating each image, the system creates a template from a 3D model and automatically projects it onto numerous images, dramatically reducing annotation time while maintaining consistent precision across all labels
Solution Approach 2:
The system performs preliminary 3D reconstruction or model acquisition before the annotation process. By having the 3D model ready in advance, the patent enables rapid generation of 2D projections for multiple images without requiring repeated manual annotation efforts, thus reducing overall time expenditure while preserving label accuracy
2Productivity
If crowdsourced workers are used for labeling, then cost is reduced and turnaround time is improved, but precision and accuracy of labels are compromised
Solution Approach 1:
The patent generates accurate 2D outline annotations by projecting 3D models onto images, creating standardized labels that can be rapidly applied to multiple images. This automated copying process maintains high precision while enabling fast turnaround, eliminating the need to choose between speed and accuracy that plagues crowdsourced approaches
3Manufacturing precision
If complex algorithms and extensive manual labor are used for 3D segmentation, then pixel-level outlines can be obtained, but device complexity and time expenditure increase
Solution Approach 1:
Instead of starting with 2D images and attempting to reconstruct 3D information through complex algorithms, the patent inverts the approach by starting with pre-existing 3D models and projecting them onto 2D images. This reversal simplifies the process significantly, obtaining pixel-level outlines through straightforward projection geometry rather than complex inverse problem solving
Data Source
AI summary
One embodiment provides a system that facilitates efficient collection of training data for training an image-detection artificial intelligence (AI) engine. During operation, the system obtains a three-dimensional (3D) model of a physical object placed in a scene, generates a virtual object corresponding to the physical object based on the 3D model, and substantially superimposes, in a view of an augmented reality (AR) camera, the virtual object over the physical object. The system can further configure the AR camera to capture a physical image comprising the physical object in the scene and a corresponding AR image comprising the virtual object superimposed over the physical object, and create an annotation for the physical image based on the AR image.


