Synthetic Data Generation for Machine Learning Model Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional approaches for training machine learning models often require large amounts of training data, which may not be readily available, leading to improperly trained models that generate incorrect outputs in real-world scenarios.
Innovation Solution
A computer-implemented method for generating data to train machine learning models involves generating a prompt based on a template and object information, using a first machine learning model to describe the texture or geometry of an object, and then using a second machine learning model to generate the texture or geometry, followed by rendering operations to produce rendered images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional approaches are used to train machine learning models, then the models can be trained with available data, but the models become improperly trained and generate incorrect outputs due to insufficient training data
Solution Approach 1:
The patent generates synthetic training data by copying and transforming existing data through virtual rendering. The system creates virtual representations of objects and scenes, generating training images that replicate real-world data characteristics without requiring actual physical data collection. This allows unlimited generation of training examples while maintaining data diversity and quality.
Solution Approach 2:
The patent introduces a virtual rendering system as an intermediary between the training data generation process and the machine learning model. This intermediary transforms object descriptions and geometries into synthetic images through computer graphics rendering, serving as a bridge that creates realistic training data without direct access to real-world data sources.
2Reliability
If large amounts of diverse training data are generated, then model accuracy improves, but the complexity of the data generation process increases
Solution Approach 1:
The patent segments the data generation process into distinct modular components: object description generation, geometry synthesis, texture generation, and virtual rendering. Each component can be independently developed, optimized, and replaced. The system divides the complex task of creating diverse training data into manageable stages that process information incrementally through specialized modules.
Solution Approach 2:
The patent creates a universal virtual rendering pipeline that can generate multiple types of training data (images, annotations, variations) from a single input object description. The same rendering system handles different object types, scenarios, and data requirements, making the complex generation process reusable and adaptable to various machine learning tasks without requiring separate specialized systems for each data type.
Data Source
AI summary
One embodiment of a method for generating data to train a machine learning model includes generating a prompt based on a template and information associated with an object, generating, via a first machine learning model and based on the prompt, a text description of at least one of a texture or a geometry for the object, generating, via a second machine learning model and based on the text description, the at least one of the texture or the geometry for the object, and performing one or more rendering operations based on the at least one of the texture or the geometry for the object to generate one or more rendered images.


