Autonomous Vehicle Perception With On-the-Fly Multimodal Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer aided perception systems in autonomous driving face challenges such as sensor limitations, reliance on manual training data labeling, and the inability to adapt on-the-fly to dynamic situations, especially when dealing with large fleets of vehicles.

Innovation Solution

Implementing a system for in situ perception that automatically labels training data on-the-fly using cross modality and cross temporal validation, enabling local and global model adaptation in autonomous vehicles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual or semi-manual labeling is used for training data, then data accuracy can be ensured, but the process becomes too slow and costly to adapt on-the-fly

Engineering Contradiction:
Improvedata labeling accuracyVSAvoiddata production speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs self-service by automatically generating and labeling training data through cross-validation between multiple sensors and temporal consistency checks, eliminating the need for manual human intervention in the data labeling process while maintaining high accuracy standards

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements feedback mechanisms where detection results from multiple sensors and time points are continuously validated against each other, with successful validations automatically fed back as labeled training data to improve the model iteratively in real-time

Inventive Principle:
Principle #23Feedback

2Measurement precision

If traditional stereo vision is used to estimate depth, then depth information can be obtained, but the computational cost becomes too high and the process is too slow

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system merges data from multiple sensor modalities including 2D cameras, 3D LiDAR sensors, and radar to obtain depth information, combining the advantages of each sensor type to achieve accurate depth estimation with reduced computational burden compared to traditional stereo vision alone

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system uses intermediate representations such as point clouds from LiDAR and radar data as mediators to bridge the gap between 2D image data and 3D depth estimation, enabling faster and more accurate depth calculation without requiring computationally intensive stereo matching

Inventive Principle:
Principle #24Intermediary (Mediator)

3Quantity of substance

If a fleet of vehicles collects data, then rich information for continuous adaptation is available, but traditional approaches cannot label such data on-the-fly

Engineering Contradiction:
Improvedata volumeVSAvoidautomatic labeling capability
Core Design Contradiction:
Quantity of substanceVSExtent of automation

Solution Approach 1:

The system implements a universal automated labeling framework that can process and validate data from multiple sensor types and multiple vehicle sources simultaneously, making the labeling process scalable and applicable to fleet-wide data collection without requiring source-specific processing pipelines

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system performs preliminary validation and filtering of incoming sensor data using cross-modality and cross-temporal consistency checks before labeling, preparing the data in advance for automatic labeling and reducing the computational burden during real-time processing of fleet-wide data

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250259030A1Method and system for distributed learning and adaptation in autonomous driving vehicles
Publication Date: 2025.08.14 PLUSAI INC
  • US20250259030A1 patent drawing
  • US20250259030A1 patent drawing
  • US20250259030A1 patent drawing

AI summary

The present teaching relates to system, method, medium for in-situ perception in an autonomous driving vehicle. A plurality of types of sensor data acquired continuously by a plurality of types of sensors deployed on the vehicle are first received, where the plurality of types of sensor data provide information about surrounding of the vehicle. Based on at least one model, one or more items are tracked from a first of the plurality of types of sensor data acquired by one or more of a first type of the plurality of types of sensors, wherein the one or more items appear in the surrounding of the vehicle. At least some of the one or more items are then automatically labeled on-the-fly via either cross modality validation or cross temporal validation of the one or more items and are used to locally adapt, on-the-fly, the at least one model in the vehicle.