Synthetic Training Image Generation for Robotic Picking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Industrial robotic package conveyance systems face challenges in accurately classifying and picking objects due to the need for extensive and costly training data, particularly in situations where edge cases and complex environments are encountered, leading to inefficiencies in using AI and ML systems.
Innovation Solution
The approach involves synthesizing composite training images by combining background and foreground data, including 3D point cloud data, to generate pre-labelled images that simulate various scenarios, reducing the requirement for extensive real-world data and enabling the training of AI and ML models to determine pickable and unpickable objects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If real-world training data is collected and labeled manually, then training data accuracy is improved, but time consumption and cost increase significantly
Solution Approach 1:
The patent uses synthetic image generation to create copies of training data through computer-generated images instead of manually collecting and labeling real-world images. This copying approach maintains the essential characteristics needed for training while eliminating the time-consuming manual data collection and labeling process
Solution Approach 2:
The patent replaces the mechanical process of manual data collection, labeling, and curation with an automated computational system that generates synthetic training images programmatically. This substitution eliminates human labor from the data preparation pipeline while maintaining data quality
2Measurement precision
If diverse training scenarios are covered, then model classification accuracy is improved, but data complexity and cost increase
Solution Approach 1:
The patent employs dynamic parameter adjustment in synthetic image generation, where training parameters such as object positions, orientations, lighting conditions, and background elements are varied programmatically to create diverse training scenarios. This dynamic approach covers multiple edge cases without manually creating complex datasets
Solution Approach 2:
The patent systematically changes parameters in synthetic image generation including object characteristics, environmental conditions, and occlusion levels to create diverse training data. By controlling parameter variations programmatically, the system achieves comprehensive scenario coverage while managing data complexity through structured parameter spaces
3Reliability
If extensive training data is collected, then model performance on edge cases is improved, but resource requirements and cost increase
Solution Approach 1:
The patent performs preliminary action by generating and preparing synthetic training data before the actual training process. This advance preparation includes creating diverse edge case scenarios programmatically, ensuring that sufficient training examples are available beforehand without requiring extensive ongoing data collection during model development
Solution Approach 2:
The patent uses synthetic copying to generate sufficient training data volumes through programmatic replication and variation of base image templates. This copying approach efficiently produces large datasets with controlled diversity, reducing the need for collecting and storing vast amounts of real-world data while maintaining model performance
Data Source
AI summary
Training images can be synthesized in order to obtain enough data to train a model (e.g., a neural network) to recognize various classifications of a type of object. Images can be synthesized by blending images of objects labeled using those classifications into selected background images. To improve results, one or more operations are performed to determine whether the synthesized images can still be used as training data, such as by verifying one or more objects of interested represented in those images is not occluded, or at least satisfies a threshold level of acceptance. The training images can be used with real world images to train the model.


