Offline Perception Model for Autonomous Driving Data Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The high cost and inefficiency of annotating training data for autonomous driving systems, particularly for spatiotemporal models that require annotated sequence data, necessitate the development of automated annotation methods to reduce human involvement.

Innovation Solution

The proposed solution involves training an offline perception model using a foundation model to predict missing sensor data, allowing for the annotation of training data without explicit labeling. This model is then fine-tuned using a second training dataset with annotated data to perform specific perception tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human annotation is used for training data, then annotation quality is high, but annotation cost and time are extremely high

Engineering Contradiction:
Improveannotation qualityVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-training a foundation model on vast amounts of unlabeled sensor data to learn temporal patterns and relationships. This pre-trained model is then used to generate annotations for specific perception tasks, eliminating the need for time-consuming human annotation while maintaining high quality through the model's learned understanding of spatiotemporal dynamics.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements self-service through self-supervised learning where the model learns to predict missing sensor data points in sequences. The system annotates its own training data by using the foundation model to generate labels automatically, reducing dependency on human annotators and significantly decreasing annotation time and cost.

Inventive Principle:
Principle #25Self-service

2Quantity of substance

If vast amounts of data are collected for training, then model performance improves, but annotation cost increases proportionally

Engineering Contradiction:
Improvedata quantityVSAvoidannotation time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent enables the system to self-annotate vast amounts of collected sensor data using the foundation model. By leveraging self-supervised learning on temporal sequences, the model automatically generates annotations for all collected data without requiring proportional human annotation resources, thus maintaining high model performance while avoiding proportional increases in annotation time and cost.

Inventive Principle:
Principle #25Self-service

3Loss of time

If automated annotation is implemented, then annotation cost decreases, but annotation quality and accuracy may deteriorate

Engineering Contradiction:
Improveannotation timeVSAvoidannotation quality
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The patent uses preliminary action by pre-training the foundation model on extensive unlabeled data to develop robust temporal understanding. This pre-training enables the automated annotation process to achieve high quality and accuracy by leveraging the model's learned spatiotemporal patterns, thus maintaining annotation quality while reducing annotation time through automation.

Inventive Principle:
Principle #10Preliminary action

4Adaptability or versatility

If deep learning models are used for perception tasks, then system capability improves, but requirement for labeled training data increases

Engineering Contradiction:
Improvesystem capabilityVSAvoidlabeled training data
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent applies self-service by using self-supervised learning to enable deep learning models to learn from unlabeled sensor data. The foundation model automatically generates its own training labels by predicting missing data points in temporal sequences, thus improving system capability for perception tasks without requiring proportional increases in labeled training data.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250139449A1Computer implemented method for providing a perception model for annotation of training data
Publication Date: 2025.05.01 ZENSEACT AB
  • US20250139449A1 patent drawing
  • US20250139449A1 patent drawing
  • US20250139449A1 patent drawing

AI summary

A method for providing an offline perception model for subsequent annotation of training data for use in training of an online perception model and a device thereof is disclosed. The method includes: training a foundation model to predict sensor data pertaining to a physical environment 5 for a time instance of a sequence of time instances, based on sensor data for remaining time instances of the sequence of time instances; forming the offline perception model by adding a task-specific layer to the trained foundation model, wherein the task-specific layer is configured to perform a perception task of the offline perception model; and fine-tuning the offline perception model to perform the perception task, wherein the second training dataset includes 10 training data annotated for the perception task. The invention further relates to a method for annotating data for use in subsequent training of an online perception model, as well as a device thereof.