Procedural Dataset Synthesis for Machine Vision Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for training machine-learning models for machine vision in self-driving vehicles require extensive and costly data collection across various environmental and weather conditions, leading to inefficiencies and inadequate representation of rare scenarios, which hinders the development of robust object recognition systems.

Innovation Solution

A cloud computing system that procedurally synthesizes diverse datasets of images and associated ground truth data, using a 3D graphics engine to generate realistic scenarios and control the distribution of training data, allowing for efficient creation of large datasets with specific variations and corner cases, thereby reducing the need for extensive real-world data collection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real-world data collection is used to train machine-learning models, then the dataset includes authentic environmental variations, but the cost and time required for data gathering increases significantly

Engineering Contradiction:
Improveauthenticity of training dataVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent creates virtual copies of real-world scenes using 3D graphics engines and procedural generation. Instead of capturing actual images from diverse environments, the system synthesizes realistic training images by rendering virtual scenes with controlled parameters, thereby eliminating time-consuming field data collection while maintaining visual authenticity

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system pre-generates comprehensive training datasets before actual model training begins. By procedurally creating diverse scenarios in advance with all necessary annotations and variations already embedded, the approach eliminates the need for time-consuming data collection during the training phase

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If extensive real-world data collection is performed to cover all environmental conditions, then the dataset diversity improves, but the cost and complexity of data gathering increases

Engineering Contradiction:
Improvedataset diversityVSAvoiddata collection system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The virtual scene generation system serves multiple functions simultaneously: it generates diverse environmental conditions, controls lighting scenarios, creates various weather conditions, and produces ground truth annotations all through a single procedural framework, eliminating the need for separate data collection systems for each condition

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system achieves dataset diversity by systematically varying parameters in virtual scene generation, such as lighting conditions, weather patterns, object positions, and environmental features, rather than physically traveling to diverse locations. This allows comprehensive coverage of all possible scenarios through controlled parameter modification

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If manual annotation is used to create ground truth data, then the accuracy of object detection training improves, but the labor and time required for annotation increases

Engineering Contradiction:
Improveground truth accuracyVSAvoidannotation efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The virtual scene generation system automatically generates its own ground truth data during the rendering process. Since the system creates the scenes programmatically with known object positions, sizes, and attributes, it can directly extract precise annotations without requiring manual intervention, achieving both high accuracy and automated efficiency

Inventive Principle:
Principle #25Self-service

4Reliability

If a vehicle is equipped with image capturing devices for data collection, then real-world training data can be gathered, but the cost of equipment and deployment increases

Engineering Contradiction:
Improvequality of training dataVSAvoiddata collection equipment
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system replaces physical image capturing devices with virtual rendering engines that generate photorealistic images. Instead of deploying cameras and recording equipment in vehicles, the approach creates synthetic copies of real-world images through 3D rendering, eliminating hardware costs while maintaining visual quality

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent substitutes mechanical data collection systems (cameras, sensors, recording equipment) with a software-based virtual generation system. The mechanical process of capturing images in diverse environments is replaced by computational rendering, eliminating the need for physical deployment of expensive equipment

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240370780A1System and method for procedurally synthesizing datasets of objects of interest for training machine-learning models
Publication Date: 2024.11.07 NVIDIA CORP
  • US20240370780A1 patent drawing
  • US20240370780A1 patent drawing

AI summary

The disclosure provides a method of training a machine-learning model employing a cloud computing system, a cloud computing system for training a machine-learning model, and a cloud computing system for synthesizing a training dataset for training a machine-learning model. In one example, the method of training a machine-learning model employing a cloud computing system includes: (1) synthesizing a plurality of images according to one or more training image definitions, (2) procedurally generating, at least partially in parallel with the synthesizing, ground truth data according to the one or more training image definitions, (3) forming a training dataset having the plurality of images and the ground truth data, and (4) training a machine-learning model using the training dataset and the ground truth data, wherein at least the synthesizing and the procedurally generating are performed by the cloud computing system.