Simulator-Generated Edge-Scene Training Data for Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning systems trained on static or artificial data sets may perform poorly in scenarios not included in the training data, particularly those with low probability of occurrence or deemed dangerous, leading to inadequate responses.
Innovation Solution
A computing device and method that generates training data using a simulator and physics engine to create diverse scenes, including edge scenes, based on parametric attributes, combining simulated and real-world data to enhance training robustness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If static data sets are used for training machine learning systems, then training data availability is improved, but the system's ability to handle scenarios not included in the training data deteriorates
Solution Approach 1:
The patent transforms the static training data approach into a dynamic one by implementing a simulator that can generate training data for any scenario on-demand. The system uses parametric attributes that can be modified to create diverse scenes, including edge cases, allowing the machine learning system to adapt to unseen scenarios while maintaining comprehensive training data availability
Solution Approach 2:
The patent applies preliminary action by pre-configuring the simulator with parametric attributes and scene definitions that can generate edge cases and rare scenarios. This allows the system to prepare training data for potential future scenarios in advance, enabling the machine learning system to handle unseen situations without requiring extensive real-world data collection for every possible case
2Adaptability or versatility
If artificial data sets from simulators are used for training, then data generation flexibility is improved, but the level of detail and realism deteriorates
Solution Approach 1:
The patent uses parameter changes to bridge the gap between artificial and real-world data quality. The simulator employs photo-realistic rendering parameters and physical accuracy parameters that can be adjusted to match real-world conditions. By carefully tuning these parameters, the system generates artificial data with high detail and realism while maintaining the flexibility to create diverse scenarios
Solution Approach 2:
The patent creates a composite training data set that combines artificial simulator-generated data with real-world data characteristics. The system integrates photo-realistic models with physically accurate simulations, combining the benefits of both artificial flexibility and real-world authenticity to produce training data that maintains high detail and realism while offering generation flexibility
3Manufacturing precision
If real-world data is used for training machine learning systems, then data realism and detail are improved, but the coverage of rare and dangerous scenarios deteriorates
Solution Approach 1:
The patent uses copying by creating simulated replicas of real-world scenarios through photo-realistic models. The simulator copies the visual and physical characteristics of real-world environments, objects, and conditions while allowing controlled modification to generate rare and dangerous edge cases that would be impossible or unethical to capture in real-world data collection
4Quantity of substance
If large quantities of data are collected to train machine learning systems, then training comprehensiveness is improved, but the time and resources required deteriorates
Solution Approach 1:
The patent implements self-service by enabling the simulator to automatically generate training data on-demand without requiring manual data collection efforts. The system serves itself by using programmed parametric attributes and scene definitions to create diverse training scenarios, eliminating the need for time-consuming real-world data collection while maintaining comprehensive training data quantity
Data Source
AI summary
A computing device, method and computer program product are provided to generate training data for a machine learning system including training data representative of one or more edge scenes. In the context of a computing device, the computing device includes a simulator configured in accordance with a sampling algorithm to create a plurality of different scenes, including one or more edge scenes, within a scenario that is at least partially defined by one or more parametric attributes. The computing device also includes a physics engine generate training data representative of the plurality of different scenes including the one or more edge scenes. The physics engine is configured to modify the one or more parametric attributes to generate additional and different training data based upon another plurality of different scenes created by the simulator within another scenario that is at least partially defined by one or more parametric attributes, as modified.
