Geospatial Region Clustering for Autonomous Training Data Transfer

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Curating large training data sets for autonomous driving systems that include diverse geographic regions is challenging due to physical accessibility issues, high costs, and regulatory constraints, making it impractical to collect sufficient data for machine learning models, especially in regions with limited market access or complex geographical conditions.

Innovation Solution

The use of geospatial clustering techniques with deep neural networks and map data to group similar geographic regions based on perceptual characteristics, allowing for the sharing of training data across clustered regions, thereby reducing the need for extensive data collection in each specific area.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If large training data sets are curated for diverse geographic regions, then model accuracy and precision improve, but data collection costs and complexity increase significantly

Engineering Contradiction:
Improveobject detection accuracyVSAvoiddata collection complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges geographically dispersed training data by clustering regions with similar visual characteristics into semantic regions. Data from multiple geographic locations are combined into unified training sets for each semantic region, allowing diverse data sources to be integrated while reducing overall collection complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates synthetic copies of training data through rendering engines that generate realistic sensor data in virtual environments. These synthetic data copies supplement real-world data, enabling model training without requiring extensive physical data collection from challenging geographic regions.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If data collection is conducted in regions with limited accessibility or regulatory constraints, then model coverage improves, but collection costs and time increase

Engineering Contradiction:
Improvegeographic coverageVSAvoiddata collection time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary clustering of geographic regions into semantic regions before data collection. This pre-organization allows data collection efforts to be focused on representative locations within each semantic region, reducing travel time and regulatory hurdles while maintaining comprehensive geographic coverage.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces synthetic data generation as an intermediary between physical data collection limitations and model training requirements. When physical access to certain regions is restricted, synthetic data serves as a mediator to provide the necessary training examples without requiring actual presence in constrained geographic areas.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If extensive data collection is performed for each specific geographic region, then region-specific model performance improves, but overall data curation costs increase

Engineering Contradiction:
Improveregion-specific model performanceVSAvoiddata curation cost
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent creates universal training data sets for each semantic region that can be applied across multiple geographic locations sharing similar characteristics. A single training data set serves multiple geographic regions simultaneously, reducing redundant data collection while maintaining region-specific model performance through semantic rather than geographic segmentation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20230368079A1Geospatial Clustering of Regions Using Neural Networks for Autonomous Systems and Applications
Publication Date: 2023.11.16 NVIDIA CORP
  • US20230368079A1 patent drawing
  • US20230368079A1 patent drawing
  • US20230368079A1 patent drawing

AI summary

In various examples, a cell model that partitions a geographic region into one or more cells is used to determine clusters of cell which share similarities. Sensor data is provided to one or more machine learning models trained to classify the sensor data to one or more cells of the cell model. Based on classifying sensor data to cells of a cell model, similarities between pairings of cells of the cell model may be determined and used to form clusters of the cell which are sufficiently similar in order to aid in the curation of training data used to train machine learning models in order to aid an autonomous or semi-autonomous machine in a surrounding environment.