Autonomous Vehicle Perception With On-the-Fly Multimodal Labeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer aided perception systems in autonomous driving face challenges such as sensor limitations, reliance on manual training data labeling, and the inability to adapt on-the-fly to dynamic situations, especially when dealing with large fleets of vehicles.
Innovation Solution
Implementing a system for in situ perception that automatically labels training data on-the-fly using cross modality and cross temporal validation, enabling local and global model adaptation in autonomous vehicles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual or semi-manual labeling is used for training data, then data accuracy can be ensured, but the process becomes too slow and costly to adapt on-the-fly
Solution Approach 1:
The system performs self-service by automatically generating and labeling training data through cross-validation between multiple sensors and temporal consistency checks, eliminating the need for manual human intervention in the data labeling process while maintaining high accuracy standards
Solution Approach 2:
The system implements feedback mechanisms where detection results from multiple sensors and time points are continuously validated against each other, with successful validations automatically fed back as labeled training data to improve the model iteratively in real-time
2Measurement precision
If traditional stereo vision is used to estimate depth, then depth information can be obtained, but the computational cost becomes too high and the process is too slow
Solution Approach 1:
The system merges data from multiple sensor modalities including 2D cameras, 3D LiDAR sensors, and radar to obtain depth information, combining the advantages of each sensor type to achieve accurate depth estimation with reduced computational burden compared to traditional stereo vision alone
Solution Approach 2:
The system uses intermediate representations such as point clouds from LiDAR and radar data as mediators to bridge the gap between 2D image data and 3D depth estimation, enabling faster and more accurate depth calculation without requiring computationally intensive stereo matching
3Quantity of substance
If a fleet of vehicles collects data, then rich information for continuous adaptation is available, but traditional approaches cannot label such data on-the-fly
Solution Approach 1:
The system implements a universal automated labeling framework that can process and validate data from multiple sensor types and multiple vehicle sources simultaneously, making the labeling process scalable and applicable to fleet-wide data collection without requiring source-specific processing pipelines
Solution Approach 2:
The system performs preliminary validation and filtering of incoming sensor data using cross-modality and cross-temporal consistency checks before labeling, preparing the data in advance for automatic labeling and reducing the computational burden during real-time processing of fleet-wide data
Data Source
AI summary
The present teaching relates to system, method, medium for in-situ perception in an autonomous driving vehicle. A plurality of types of sensor data acquired continuously by a plurality of types of sensors deployed on the vehicle are first received, where the plurality of types of sensor data provide information about surrounding of the vehicle. Based on at least one model, one or more items are tracked from a first of the plurality of types of sensor data acquired by one or more of a first type of the plurality of types of sensors, wherein the one or more items appear in the surrounding of the vehicle. At least some of the one or more items are then automatically labeled on-the-fly via either cross modality validation or cross temporal validation of the one or more items and are used to locally adapt, on-the-fly, the at least one model in the vehicle.


