Time-Series Ground Truth Generation for Autonomous Driving ML

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The process of creating training data for deep learning systems used in autonomous driving is labor-intensive and requires significant resources, particularly in manually labeling features in the training data.

Innovation Solution

A method is developed to automatically generate training data by using a time series of sensor data from vehicles, including image data and odometry data, to create a training dataset that can be used to train machine learning models for autonomous driving.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labeling is used to create training data, then accuracy of labels is improved, but productivity and time consumption deteriorate

Engineering Contradiction:
Improvelabel accuracyVSAvoiddata curation efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system uses automatically generated ground truth data from sensor fusion and tracking algorithms to label training data without human intervention. The autonomous vehicle's own sensors and processing systems serve to create the training labels, eliminating the need for external manual annotation while maintaining high accuracy through multiple sensor corroboration.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual human labeling with automated computational algorithms. Sensor fusion algorithms, object tracking systems, and ground truth generation algorithms substitute for human annotators, dramatically increasing productivity while maintaining or improving label accuracy through consistent algorithmic application.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If more training data is collected to improve model performance, then machine learning accuracy is improved, but resources and time investment deteriorate

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system continuously collects and processes sensor data during normal autonomous vehicle operation, converting routine operational data into training data. This continuous data generation from ongoing vehicle operation eliminates the need for separate, time-consuming data collection missions, as every moment of vehicle operation contributes to building the training dataset.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The autonomous vehicle system generates its own training data during normal operation without requiring external data collection resources. The vehicle's sensors, processors, and tracking algorithms work together to automatically create labeled training data from the vehicle's own operational experience, eliminating the need for dedicated data collection teams and equipment.

Inventive Principle:
Principle #25Self-service

3Reliability

If difficult-to-label data is collected to address model weaknesses, then model improvement is enabled, but labeling difficulty and resource requirements deteriorate

Engineering Contradiction:
Improvemodel performanceVSAvoidlabeling complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces complex manual labeling processes with automated sensor fusion and algorithmic ground truth generation. Even for difficult cases such as occluded objects or ambiguous scenarios, the system uses multiple sensor modalities and tracking algorithms to automatically generate labels, eliminating the need for human experts to interpret and label challenging data points.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system introduces automated tracking algorithms and sensor fusion processes as intermediaries between raw sensor data and training labels. These intermediary algorithms process and interpret difficult-to-label data automatically, serving as a bridge that converts complex sensor inputs into reliable training labels without requiring direct human intervention in challenging cases.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12223428B2Generating ground truth for machine learning from time series elements
Publication Date: 2025.02.11 TESLA INC
  • US12223428B2 patent drawing
  • US12223428B2 patent drawing
  • US12223428B2 patent drawing

AI summary

Sensor data, including a group of time series elements, is received. A training data set is determined, including by determining for at least a selected time series element in the group of time series elements a corresponding ground truth. The corresponding ground truth is based on a plurality of time series elements in the group of time series elements. A processor is used to train a machine learning model using the training dataset.