Autonomous Driving Prediction Using Generative Image Token Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing autonomous driving systems rely heavily on expensive and labor-intensive accurate perception annotations, limiting their deployment cost, scalability, and practical applicability due to the need for time-consuming and costly supervision signals.

Innovation Solution

An information prediction method that uses a generative model to predict control information for vehicles by encoding perception data, including image and driving data, without relying on precise annotations, by incorporating an image generation task as an auxiliary task to improve environmental perception and representation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If accurate perception annotations are used for training autonomous driving systems, then the prediction accuracy and reliability are improved, but the cost and time consumption increase significantly

Engineering Contradiction:
Improveprediction accuracyVSAvoidannotation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs self-supervised learning by automatically generating supervision signals from raw perception data without external annotations. The generative model creates synthetic training data and supervision signals autonomously, allowing the system to train itself using unannotated data, thereby eliminating the need for time-consuming manual annotations while maintaining training effectiveness

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The generative model creates synthetic copies of real-world perception data through image generation tasks. By generating synthetic images and their corresponding annotations, the system obtains virtual supervision signals that mimic real annotated data, enabling training without accessing costly real-world annotated datasets

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If expensive accurate perception annotations are used, then the model training quality is improved, but the deployment cost and scalability are limited

Engineering Contradiction:
Improvemodel training qualityVSAvoiddeployment cost
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The system autonomously generates its own training data and supervision signals through self-supervised learning mechanisms. The generative model creates synthetic annotated data from unannotated perception inputs, enabling the system to achieve high training quality without incurring costs associated with manual annotation services, thereby reducing deployment costs and improving scalability

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system replaces expensive, hard-to-obtain accurate perception annotations with cheap, easily generated synthetic annotations from the generative model. These synthetic annotations serve as disposable training data that can be generated on-demand without financial cost, making the system scalable and economically viable for widespread deployment

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Measurement precision

If precise perception annotations are required, then the supervision signal quality is improved, but the scalability and practical applicability are reduced

Engineering Contradiction:
Improvesupervision signal qualityVSAvoidscalability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system achieves self-supervised learning by automatically generating supervision signals from raw perception data through image generation tasks. This self-service mechanism allows the system to maintain high supervision signal quality while processing unlimited amounts of unannotated data, thereby achieving both precision and scalability simultaneously without being constrained by annotation availability

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250206330A1Information prediction method, method of training autonomous driving model, device, medium, and autonomous driving vehicle
Publication Date: 2025.06.26 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US20250206330A1 patent drawing
  • US20250206330A1 patent drawing
  • US20250206330A1 patent drawing

AI summary

An information prediction method, a method of training an autonomous driving model, a device, a medium, and an autonomous driving vehicle, which relate to a field of artificial intelligence technology, and in particular, to fields of computer vision technology and deep learning technology, which may be applied to scenarios such as autonomous driving. Specific implementation scheme of the information prediction method is: acquiring perception data including image data acquired by a sensor in a vehicle and driving data of the vehicle; encoding the image data to obtain an image token sequence corresponding to the image data; encoding the driving data to obtain a driving feature corresponding to the driving data; and generating, using a generative model, a predicted token sequence corresponding to the image token sequence and a control information for the vehicle based on the driving feature and the image token sequence.