CG-Based Synthetic Training Data Generation for AI Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Collecting training data for machine learning models, especially for specialized scenarios like automated driving and medical imaging, is time-consuming, costly, and prone to human errors, with ethical challenges in simulating difficult situations.
Innovation Solution
Utilizing computer graphics (CG) models to generate training data, including realistic images and annotations, which can simulate various scenarios and conditions, reducing the need for actual data collection and minimizing human intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If training data is collected from actual landscapes and paintings, then the training data reflects real-world scenarios, but the data collection process is time-consuming and labor-intensive
Solution Approach 1:
The patent uses computer-generated images to create copies of real-world training data scenarios. Instead of collecting actual photographs and manual paintings, the system generates synthetic images that replicate real-world conditions, including lighting, textures, and object appearances. This copying approach maintains the realism needed for training while eliminating the time-consuming data collection process.
Solution Approach 2:
The patent replaces the mechanical process of manual data collection (photographing, painting, annotating) with an automated computer-based system. The system uses algorithms to generate training data automatically, substituting human labor with computational processes that can produce large volumes of training data rapidly and consistently.
2Measurement precision
If human annotators add annotations to training data, then the annotations are accurate and meaningful, but human errors occur and the process is labor-intensive
Solution Approach 1:
The patent implements a self-service annotation system where the computer-generated images automatically include embedded annotations and metadata. Instead of requiring human annotators to manually label objects and features, the generation system itself embeds the necessary annotation information during the image creation process, making the system self-sufficient and eliminating human annotation labor.
Solution Approach 2:
The patent replaces the mechanical process of manual annotation with automated computational annotation. Algorithms automatically generate and attach metadata, labels, and annotations to the synthetic training images, substituting human annotators with computational processes that eliminate human errors while maintaining annotation accuracy through consistent algorithmic application.
3Adaptability or versatility
If training data for special situations (e.g., nighttime automated driving) is collected, then the model can handle specialized scenarios, but the data collection becomes costly
Solution Approach 1:
The patent uses parameter changes to efficiently create specialized training scenarios. By modifying parameters such as lighting conditions, time of day, weather conditions, and environmental settings in the computer-generated system, the same base generation process can produce diverse specialized scenarios (nighttime, rain, snow, different locations) without requiring separate expensive data collection campaigns for each condition.
Solution Approach 2:
The patent creates a universal data generation system that can produce training data for multiple specialized scenarios using a single platform. The same computer-generated system can create images for various conditions (different times, weather, locations, objects), making the system multi-functional and eliminating the need for separate specialized data collection processes for each scenario type.
4Adaptability or versatility
If images of difficult-to-encounter situations (e.g., accidents, pathology scenes) are collected, then the model can learn from rare events, but ethical issues arise and data collection is difficult
Solution Approach 1:
The patent uses copying to create synthetic representations of rare and sensitive scenarios without capturing or exploiting real-world harmful events. Instead of collecting actual images of accidents or medical procedures from real patients, the system generates synthetic copies that replicate the visual and structural characteristics needed for training, thereby avoiding ethical violations while still providing training value for rare scenarios.
Solution Approach 2:
The patent introduces a computer-generated intermediary system that mediates between the need for rare scenario training data and ethical constraints. This intermediary generates synthetic representations that serve as a buffer, allowing the model to learn from rare events without directly using or exploiting sensitive real-world data, thus resolving the ethical conflict while maintaining training effectiveness.
Data Source
AI summary
[Object]Training data is acquired using a computer graphics.[Solution]An image generation method includes acquiring a CG model or an artificial image generated based on the CG model and performing, by a processor, processing on the CG model or the artificial image and generating metadata of a processed image used for AI learning used for an image acquired by a sensor or the artificial image.


