Synthetic Data Generation for Machine Vision Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge in training machine vision systems lies in the requirement for large amounts of quality training data, which is often obtained through manual annotation of real-world images, a time-consuming and costly process. Additionally, determining the appropriate data to collect in terms of subject, quality, and distribution is complex.
Innovation Solution
The use of computer graphics to generate synthetic training data, which includes annotations, allows for on-demand creation of datasets with user-configurable size and characteristics, using digital assets that can be captured or imported, and adaptive data generation methods to optimize the training process by iteratively refining the dataset based on performance scores.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual annotation of real-world images is used to gather training data, then the quality and accuracy of training data is improved, but the time consumption and cost increase significantly
Solution Approach 1:
The patent uses computer graphics to generate synthetic copies of real-world images and annotations. Instead of manually annotating actual photographs, the system creates digital replicas with automatically generated annotations, eliminating the time-consuming manual process while maintaining data quality through controlled synthesis parameters
Solution Approach 2:
The system employs automated algorithms to generate synthetic training data with annotations without human intervention. The computer graphics engine automatically creates images, applies transformations, and generates corresponding annotations, making the data generation process self-sufficient and eliminating dependency on manual labor
2Measurement precision
If manual annotation methods are used for data collection, then the accuracy of annotations is improved, but the scalability and efficiency of data generation deteriorates
Solution Approach 1:
Synthetic images are generated as digital copies with programmatically created annotations. The system replicates real-world scenarios through computer graphics while automatically generating corresponding annotations, achieving both accuracy through controlled parameter settings and high efficiency through automated generation
Solution Approach 2:
The system varies parameters such as object positions, orientations, lighting conditions, and background environments in synthetic image generation. By systematically changing these parameters, the system generates diverse training data with accurate annotations at scale, improving both annotation quality and generation efficiency
3Adaptability or versatility
If real-world images are collected for training, then the diversity of training data is improved, but the control over data distribution and quality characteristics deteriorates
Solution Approach 1:
The system controls data diversity by programmatically varying parameters in synthetic image generation, including object types, positions, orientations, lighting conditions, and background environments. This approach provides both diverse training data and precise control over data distribution characteristics
Solution Approach 2:
The system pre-defines the characteristics and distribution of training data before generation by setting parameters in the computer graphics engine. This preliminary configuration ensures desired data diversity and quality characteristics are achieved without complex post-processing or selection
4Quantity of substance
If large amounts of training data are generated manually, then the completeness of training datasets is improved, but the cost and resource requirements increase
Solution Approach 1:
The system generates large volumes of training data through automated computer graphics synthesis rather than manual collection. Digital copies of scenes with programmatically generated annotations are created at scale, achieving high data volume without proportional increases in human resource consumption
Solution Approach 2:
The automated synthetic data generation system operates without continuous human intervention, generating large datasets through self-executing algorithms. This self-service approach reduces resource consumption by eliminating the need for manual data collection and annotation processes
Data Source
AI summary
A computer-implemented method of performing machine vision prediction of digital images using synthetically generated training assets comprises digitally capturing a plurality of assets; configuring each of the assets in the plurality of assets with a plurality of asset attributes; under computer program control, selecting a plurality of different combinations of parameters from among the plurality of asset attributes, and creating a plurality of sets of different synthetic dataset parameters; using computer graphics software, and example parameter values from among the synthetic dataset parameters, creating a synthetic dataset by compiling from a plurality of example images and metadata; configuring a plurality of machine learning trials and executing the trials to train a machine vision model, resulting in creating and storing a trained machine vision model; executing a validation of the trained machine vision model; and inferring a prediction using the trained machine vision model. Trained models are scored against success criteria and re-trained using pseudo-random sampling of different parameters clustered around failure points. As a result, machine vision models may be trained with high accuracy using large datasets of synthesized digital images that are richly parameterized, rather than human captured digital images.


