Cross-Modality Object Labeling for Adaptive AV Perception
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer-aided perception systems in autonomous driving face challenges such as sensor limitations, high costs of manual data labeling, and the inability to adapt on-the-fly to dynamic situations, especially when dealing with large fleets of vehicles generating vast amounts of diverse data.
Innovation Solution
The development of an in-situ perception system that continuously acquires and processes data from multiple sensors, enabling automatic on-the-fly labeling and model updates using cross-modality validation and object-centric stereo for enhanced object detection and depth estimation, allowing for local and global model adaptation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual or semi-manual labeling is used for training data, then labeling accuracy can be maintained, but the cost and time required become prohibitively high
Solution Approach 1:
The system performs self-labeling by automatically generating training data labels through cross-modality validation between active sensors (LiDAR) and passive sensors (cameras). The active sensor data serves as ground truth to validate and label objects detected by passive sensors, enabling the system to label its own training data without human intervention.
Solution Approach 2:
Active sensors (LiDAR) act as an intermediary to validate and label objects detected by passive sensors (cameras). The active sensor data serves as a mediator that provides reliable depth information and object validation, enabling automatic labeling of training data while maintaining accuracy.
2Measurement precision
If active sensors like LiDAR are used for depth information, then depth measurement accuracy is improved, but the range and data density are limited
Solution Approach 1:
The system merges data from active sensors (LiDAR) and passive sensors (cameras) through cross-modality validation. The active sensor provides accurate depth measurements for nearby objects, while the passive sensor extends the sensing range to distant objects, creating a complementary system that overcomes the limitations of each individual sensor type.
Solution Approach 2:
The passive camera system serves multiple functions: it provides extended range detection beyond the LiDAR's effective distance, captures visual information for object classification, and validates LiDAR detections. This multi-functionality allows the system to compensate for the limited range of active sensors.
3Productivity
If traditional stereo vision is used for depth estimation, then computational cost is reduced, but the processing speed and depth map density are insufficient
Solution Approach 1:
The system replaces traditional stereo vision computational methods with active sensor-based depth measurement. Instead of performing computationally intensive stereo matching algorithms, the system directly obtains depth information from LiDAR range measurements, significantly reducing computational cost while improving processing speed and depth map density.
4Adaptability or versatility
If models are updated frequently to adapt to dynamic situations, then adaptability is improved, but the computational resources and time required for training increase
Solution Approach 1:
The system performs preliminary labeling of training data during normal operation using cross-modality validation, so that when model updates are needed, pre-labeled data is already available. This eliminates the need for intensive real-time labeling during adaptation, reducing computational resources required for frequent model updates.
Solution Approach 2:
The system continuously generates labeled training data during normal operation through cross-modality validation, maintaining a steady stream of training examples. This continuous generation of training data allows for frequent, low-cost model updates without requiring intensive batch processing or manual intervention.
Data Source
AI summary
The present teaching relates to method, system, medium, and implementation of in-situ perception in an autonomous driving vehicle. A plurality of types of sensor data are acquired continuously via a plurality of types of sensors deployed on the vehicle, where the plurality of types of sensor data provide information about surrounding of the vehicle. One or more items surrounding the vehicle are tracked, based on some models, from a first of the plurality of types of sensor data from a first type of the plurality of types of sensors. A second of the plurality of types of sensor data are obtained from a second type of the plurality of sensors and are used to generate validation base data. Some of the one or more items are labeled, automatically, via validation base data to generate labeled at least some item, which is to be used to generate model updated information for updating the at least one model.


