Time-Series Ground Truth Generation for Autonomous Driving ML
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The process of creating training data for deep learning systems used in autonomous driving is labor-intensive and requires significant resources, particularly in manually labeling features in the training data.
Innovation Solution
A method is developed to automatically generate training data by using a time series of sensor data from vehicles, including image data and odometry data, to create a training dataset that can be used to train machine learning models for autonomous driving.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual labeling is used to create training data, then accuracy of labels is improved, but productivity and time consumption deteriorate
Solution Approach 1:
The system uses automatically generated ground truth data from sensor fusion and tracking algorithms to label training data without human intervention. The autonomous vehicle's own sensors and processing systems serve to create the training labels, eliminating the need for external manual annotation while maintaining high accuracy through multiple sensor corroboration.
Solution Approach 2:
The patent replaces the mechanical process of manual human labeling with automated computational algorithms. Sensor fusion algorithms, object tracking systems, and ground truth generation algorithms substitute for human annotators, dramatically increasing productivity while maintaining or improving label accuracy through consistent algorithmic application.
2Measurement precision
If more training data is collected to improve model performance, then machine learning accuracy is improved, but resources and time investment deteriorate
Solution Approach 1:
The system continuously collects and processes sensor data during normal autonomous vehicle operation, converting routine operational data into training data. This continuous data generation from ongoing vehicle operation eliminates the need for separate, time-consuming data collection missions, as every moment of vehicle operation contributes to building the training dataset.
Solution Approach 2:
The autonomous vehicle system generates its own training data during normal operation without requiring external data collection resources. The vehicle's sensors, processors, and tracking algorithms work together to automatically create labeled training data from the vehicle's own operational experience, eliminating the need for dedicated data collection teams and equipment.
3Reliability
If difficult-to-label data is collected to address model weaknesses, then model improvement is enabled, but labeling difficulty and resource requirements deteriorate
Solution Approach 1:
The patent replaces complex manual labeling processes with automated sensor fusion and algorithmic ground truth generation. Even for difficult cases such as occluded objects or ambiguous scenarios, the system uses multiple sensor modalities and tracking algorithms to automatically generate labels, eliminating the need for human experts to interpret and label challenging data points.
Solution Approach 2:
The system introduces automated tracking algorithms and sensor fusion processes as intermediaries between raw sensor data and training labels. These intermediary algorithms process and interpret difficult-to-label data automatically, serving as a bridge that converts complex sensor inputs into reliable training labels without requiring direct human intervention in challenging cases.
Data Source
AI summary
Sensor data, including a group of time series elements, is received. A training data set is determined, including by determining for at least a selected time series element in the group of time series elements a corresponding ground truth. The corresponding ground truth is based on a plurality of time series elements in the group of time series elements. A processor is used to train a machine learning model using the training dataset.


