Time-Series Ground Truth Generation for Lane-Line Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The process of generating training data for machine learning models, particularly for autonomous driving, is labor-intensive and inefficient due to the need for manual annotation and accurate labeling of features, which limits the performance of deep learning systems.

Innovation Solution

A method that utilizes sensor data from vehicles to create a training dataset by capturing time series elements, such as image and odometry data, to determine ground truth for features like lane lines, which is then used to train machine learning models to predict three-dimensional representations of these features, reducing the reliance on manual annotation and improving data accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation and labeling of training data is performed, then data accuracy is improved, but labor intensity and time consumption increase significantly

Engineering Contradiction:
Improvedata accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by using semi-automatic annotation tools to pre-label training data before final machine learning processing. This preliminary labeling reduces the need for extensive manual annotation later, thereby maintaining data accuracy while reducing overall time consumption and labor intensity.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If manual curation of training data is performed, then data quality is improved, but resource investment increases

Engineering Contradiction:
Improvedata qualityVSAvoidresource efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system implements self-service mechanisms where the machine learning model automatically processes and validates training data with minimal human intervention. The semi-automatic annotation system allows the data to be curated through automated processes that self-correct and self-validate, maintaining high data quality while significantly improving resource efficiency by reducing manual labor requirements.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If extensive manual annotation is performed to improve model performance, then model accuracy is improved, but the complexity of the annotation process increases

Engineering Contradiction:
Improvemodel accuracyVSAvoidannotation process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces semi-automatic annotation tools as an intermediary between raw data and final training datasets. These tools act as mediators that automate routine annotation tasks while allowing human annotators to focus on complex cases, thereby maintaining model accuracy while reducing the overall complexity of the annotation process.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Reliability

If more training data is collected and annotated, then model performance is improved, but the effort required for data curation increases

Engineering Contradiction:
Improvemodel performanceVSAvoiddata curation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system enables self-service data curation where the machine learning model automatically processes large volumes of training data with minimal human intervention. This allows extensive data collection to proceed efficiently, maintaining high model performance while dramatically improving data curation efficiency by reducing the manual effort required per data point.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11748620B2Generating ground truth for machine learning from time series elements
Publication Date: 2023.09.05 TESLA INC
  • US11748620B2 patent drawing
  • US11748620B2 patent drawing
  • US11748620B2 patent drawing

AI summary

Sensor data, including a group of time series elements, is received. A training data set is determined, including by determining for at least a selected time series element in the group of time series elements a corresponding ground truth. The corresponding ground truth is based on a plurality of time series elements in the group of time series elements. A processor is used to train a machine learning model using the training dataset.