Personalized Training Data Synthesis for ML Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Obtaining and labeling training data for machine learning systems is labor-intensive and costly, especially due to the sheer volume of data required, and synthetic data alone can struggle to accurately represent real-world scenarios, known as the 'synth to real' gap.
Innovation Solution
A method for generating personalized training data using a computational model based on user image data, combining real-world 3D topology and synthetic elements, which allows for automatic labeling and generation of varied views under different conditions, reducing manual labor and improving data accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual labelling of training data is performed, then data accuracy is improved, but productivity deteriorates due to arduous and time-consuming human involvement
Solution Approach 1:
The patent uses synthetic data generation to create copies of real-world training data scenarios. A computational model generates synthetic images, point clouds, and labels that replicate real driving conditions without requiring manual human labelling, thus maintaining accuracy while dramatically improving productivity
Solution Approach 2:
The system performs self-service by automatically generating labelled training data through computational models without human intervention. The model synthesizes data and assigns labels autonomously, eliminating the need for manual labelling while maintaining data quality
2Productivity
If synthetic data is used for training, then productivity is improved due to easier and cheaper generation, but measurement precision deteriorates due to the 'synth to real' gap
Solution Approach 1:
The patent employs parameter changes by systematically varying environmental conditions (lighting, weather, time of day) and object parameters in synthetic data generation. This creates diverse, realistic training scenarios that bridge the synth-to-real gap while maintaining high productivity through automated generation
Solution Approach 2:
The system creates composite training datasets by combining synthetic generated data with real-world data characteristics. The computational model integrates multiple data types (images, point clouds, labels) with varying degrees of synthetic and real-world properties to achieve both efficiency and accuracy
3Reliability
If large volume of training data is collected, then machine learning model performance is improved, but loss of time increases due to the sheer amount of data required
Solution Approach 1:
The patent applies preliminary action by pre-generating comprehensive training datasets using computational models before actual model training begins. Synthetic data covering various scenarios is created in advance, eliminating the time-consuming process of collecting and labelling large volumes of real-world data
Data Source
AI summary
A method of dataset generation is described. The method comprises steps including receiving user image data and generating personalised training data based on the received user image data. Generating personalised training data comprises generating a computational model based at least in part on the received user data and generating the personalised training data based on the computational model.


