Cross-Modality Object Labeling for Adaptive AV Perception

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer-aided perception systems in autonomous driving face challenges such as sensor limitations, high costs of manual data labeling, and the inability to adapt on-the-fly to dynamic situations, especially when dealing with large fleets of vehicles generating vast amounts of diverse data.

Innovation Solution

The development of an in-situ perception system that continuously acquires and processes data from multiple sensors, enabling automatic on-the-fly labeling and model updates using cross-modality validation and object-centric stereo for enhanced object detection and depth estimation, allowing for local and global model adaptation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual or semi-manual labeling is used for training data, then labeling accuracy can be maintained, but the cost and time required become prohibitively high

Engineering Contradiction:
Improvelabeling accuracyVSAvoidtime to produce training data
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-labeling by automatically generating training data labels through cross-modality validation between active sensors (LiDAR) and passive sensors (cameras). The active sensor data serves as ground truth to validate and label objects detected by passive sensors, enabling the system to label its own training data without human intervention.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Active sensors (LiDAR) act as an intermediary to validate and label objects detected by passive sensors (cameras). The active sensor data serves as a mediator that provides reliable depth information and object validation, enabling automatic labeling of training data while maintaining accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If active sensors like LiDAR are used for depth information, then depth measurement accuracy is improved, but the range and data density are limited

Engineering Contradiction:
Improvedepth measurement accuracyVSAvoidsensing range
Core Design Contradiction:
Measurement precisionVSLength of stationary object

Solution Approach 1:

The system merges data from active sensors (LiDAR) and passive sensors (cameras) through cross-modality validation. The active sensor provides accurate depth measurements for nearby objects, while the passive sensor extends the sensing range to distant objects, creating a complementary system that overcomes the limitations of each individual sensor type.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The passive camera system serves multiple functions: it provides extended range detection beyond the LiDAR's effective distance, captures visual information for object classification, and validates LiDAR detections. This multi-functionality allows the system to compensate for the limited range of active sensors.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If traditional stereo vision is used for depth estimation, then computational cost is reduced, but the processing speed and depth map density are insufficient

Engineering Contradiction:
Improveprocessing speedVSAvoiddepth map density
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system replaces traditional stereo vision computational methods with active sensor-based depth measurement. Instead of performing computationally intensive stereo matching algorithms, the system directly obtains depth information from LiDAR range measurements, significantly reducing computational cost while improving processing speed and depth map density.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Adaptability or versatility

If models are updated frequently to adapt to dynamic situations, then adaptability is improved, but the computational resources and time required for training increase

Engineering Contradiction:
Improvemodel adaptation to dynamic situationsVSAvoidcomputational resources for training
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary labeling of training data during normal operation using cross-modality validation, so that when model updates are needed, pre-labeled data is already available. This eliminates the need for intensive real-time labeling during adaptation, reducing computational resources required for frequent model updates.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system continuously generates labeled training data during normal operation through cross-modality validation, maintaining a steady stream of training examples. This continuous generation of training data allows for frequent, low-cost model updates without requiring intensive batch processing or manual intervention.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS12039445B2Method and system for on-the-fly object labeling via cross modality validation in autonomous driving vehicles
Publication Date: 2024.07.16 PLUSAI INC
  • US12039445B2 patent drawing
  • US12039445B2 patent drawing
  • US12039445B2 patent drawing

AI summary

The present teaching relates to method, system, medium, and implementation of in-situ perception in an autonomous driving vehicle. A plurality of types of sensor data are acquired continuously via a plurality of types of sensors deployed on the vehicle, where the plurality of types of sensor data provide information about surrounding of the vehicle. One or more items surrounding the vehicle are tracked, based on some models, from a first of the plurality of types of sensor data from a first type of the plurality of types of sensors. A second of the plurality of types of sensor data are obtained from a second type of the plurality of sensors and are used to generate validation base data. Some of the one or more items are labeled, automatically, via validation base data to generate labeled at least some item, which is to be used to generate model updated information for updating the at least one model.