Synthetic Image Generation for ML Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing methods for training machine learning models for object detection are time-consuming and expensive, particularly when custom object types like industrial parts require extensive annotation of images captured in various positions, orientations, and settings, which is exacerbated by the need for pixel-wise detection.
Innovation Solution
The approach generates photorealistic synthetic images using 3D rendering engines with randomized visual parameters such as illumination, textures, and camera effects, allowing for automated creation of training datasets with minimal human input, thereby training machine learning models for robust object detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional methods are used to capture and annotate images for training ML models, then the model can be trained for object detection, but the process is time-consuming and expensive
Solution Approach 1:
The patent uses 3D rendering engines to generate synthetic images that copy and replicate real-world objects, scenes, and visual characteristics without requiring actual physical objects or human annotators. The synthetic images are created by rendering 3D models with randomized parameters, producing photorealistic images that can be used for training ML models without the time-consuming process of capturing and manually annotating real images.
Solution Approach 2:
The patent applies parameter changes by randomly varying multiple visual parameters such as lighting conditions, camera angles, object positions, textures, and environmental factors when generating synthetic images. This allows the system to create diverse training data efficiently by changing parameters rather than capturing each variation physically, thereby increasing productivity while reducing time loss.
2Measurement precision
If extensive manual annotation is performed for custom object types, then pixel-wise detection accuracy is achieved, but the cost and time required increase significantly
Solution Approach 1:
The system performs self-service by automatically generating annotations alongside the synthetic images. The 3D rendering engine inherently knows the ground truth labels, object positions, and pixel-wise segmentation masks because the images are generated from precise 3D models and parameters. This eliminates the need for manual annotation while maintaining high pixel-wise detection accuracy, as the annotations are automatically derived from the rendering process itself.
3Adaptability or versatility
If real images are captured from multiple positions and orientations, then comprehensive training data is obtained, but the complexity and cost of data collection increase
Solution Approach 1:
The patent implements dynamics by dynamically randomizing multiple parameters including camera position, camera orientation, object position, object orientation, lighting conditions, and texture variations when generating synthetic images. This dynamic parameter randomization allows the system to comprehensively cover various viewing angles and conditions without requiring physical movement of cameras or objects, thereby achieving versatile training data with reduced system complexity.
Data Source
AI summary
In one example in accordance with the present disclosure, an electronic device is described. An example electronic device includes a processor and memory storing executable instructions that when executed cause the processor to generate multiple synthetic images of an object based on defined object parameters and randomized visual parameters. The instructions also cause the processor to generate annotations of the object in multiple synthetic images based on the defined object parameters and the randomized visual parameters. The instructions further cause the processor to train a machine-learning (ML) model for detecting the object using the multiple synthetic images and annotations.


