Synthetic Visual Inspection Data Using Augmented Reality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Deep learning models for computer vision face challenges in generalizing to real-life scenarios due to overfitting when trained with insufficient data, particularly in variations of lighting and orientation, which traditional image manipulation techniques fail to address.
Innovation Solution
The method involves creating synthetic visual inspection data sets using augmented reality to generate a 3D model of an anchor object, allowing for the simulation of various orientations and conditions, thereby expanding the training dataset through data augmentation techniques like moving, rotating, and re-texturing associated 3D models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional image manipulation techniques (crop, flip, rotate, hue, saturation) are used to augment training data, then data diversity is improved, but the model still fails to generalize to real-life scenarios with variations in lighting and orientation
Solution Approach 1:
The patent creates synthetic copies of training images by rendering 3D models of objects in various orientations, lighting conditions, and backgrounds. These synthetic images are then combined with real images to form an augmented training dataset, allowing the model to learn from diverse simulated scenarios without requiring extensive real-world data collection
Solution Approach 2:
The patent systematically varies parameters such as lighting conditions (illumination intensity, direction, color temperature), object orientations (rotation angles, positions), and background environments in the synthetic image generation process. This creates a comprehensive set of training examples that cover the range of real-life variations the model will encounter
2Productivity
If a small training dataset is used during model training, then training time and computational resources are reduced, but the model overfits and performs poorly on real-life scenarios
Solution Approach 1:
The system generates numerous synthetic copies of object images by rendering 3D models in different poses, lighting conditions, and environments. These synthetic images are combined with real images to create a large augmented training dataset, providing sufficient diversity for model generalization without requiring proportional increases in real data collection
Solution Approach 2:
The patent pre-renders synthetic training images covering various conditions (different lighting, orientations, backgrounds) before model training begins. This preliminary preparation of diverse training data allows the model to be trained efficiently on a comprehensive dataset rather than requiring extensive data collection during the training process
3Adaptability or versatility
If synthetic images are created by compositing target objects onto backgrounds, then training data diversity is improved, but image realism may be compromised
Solution Approach 1:
The patent carefully controls rendering parameters including lighting conditions (to match background illumination), shadow placement and intensity, object positioning and scaling, and background selection. These parameter adjustments ensure that synthetic composites appear realistic while maintaining diversity in orientations and configurations
Solution Approach 2:
The patent uses rendered synthetic images as an intermediary between real object photos and the final training dataset. These synthetically generated images serve as a bridge, combining the accuracy of 3D model geometry with diverse real-world lighting and background conditions, producing realistic training examples
Data Source
AI summary
In an approach for creating synthetic visual inspection data sets for training an artificial intelligence computer vision deep learning model utilizing augmented reality, a processor enables a user to capture a plurality of images of an anchor object using a camera on a user computing device. A processor receives the plurality of images of the anchor object from the user. A processor generates a baseline model of an anchor object. A processor generates a training data set. A processor trains the baseline model of the anchor object. A processor creates a trained Artificial Intelligence (AI) computer vision deep learning model. A processor enables the user to interact with the trained AI computer vision deep learning model in an access mode.


