Object-Centric Stereo Depth Estimation With Cross-Modal Validation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer-aided perception systems in autonomous driving face limitations, including sensor sensitivity to lighting conditions, computational inefficiencies in depth estimation, and the need for manual data labeling, which hinders real-time adaptation and scalability in fleet operations.

Innovation Solution

The development of a system that utilizes object-centric stereo for enhanced depth estimation and cross-modal validation to automatically label training data on-the-fly, enabling continuous adaptation and global model updates based on diverse environmental data from a fleet of vehicles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional stereo using multiple cameras is used for depth estimation, then depth information can be obtained, but the computational cost is high and the speed is slow

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the depth estimation process by first detecting objects using 2D cameras, then estimating depth only for detected objects using 3D sensor data. This selective approach avoids the computational burden of traditional stereo matching across the entire image, significantly reducing processing time while maintaining depth accuracy for relevant objects.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges 2D camera data for object detection with 3D sensor data (LiDAR or radar) for depth estimation. By combining the strengths of both sensor types - 2D cameras provide high-resolution object identification while 3D sensors provide accurate depth information - the system achieves both speed and accuracy without relying on computationally expensive traditional stereo algorithms.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If 3D sensors such as LiDAR are used to acquire depth information, then depth data can be obtained, but the range is limited and data density is low

Engineering Contradiction:
Improvedepth information qualityVSAvoidsensing range
Core Design Contradiction:
Measurement precisionVSArea of stationary object

Solution Approach 1:

The patent combines 2D camera data that captures wide-field visual information with 3D sensor data that provides precise depth measurements. The 2D cameras extend the effective sensing range by detecting objects at greater distances, while 3D sensors provide detailed depth information for closer objects, creating a complementary system that overcomes the limited range of individual 3D sensors.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The 2D camera system acts as an intermediary that identifies and localizes objects in the wide field of view, then guides the 3D sensor focus toward these detected objects. This intermediary role allows the system to effectively extend the 3D sensor's useful range by pre-screening the environment and directing depth measurement resources to relevant areas.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If manual or semi-manual labeling is used for training data, then training data can be produced, but the process is slow and costly

Engineering Contradiction:
Improvetraining data qualityVSAvoiddata production speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs self-service by automatically generating its own training data through unsupervised learning from raw sensor data collected during normal operation. The multi-sensor fusion architecture inherently provides labeled data - objects detected by 2D cameras are automatically associated with their 3D depth measurements - eliminating the need for manual annotation while maintaining high data quality and enabling continuous data production at system operational speed.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements feedback loops where detection results from 2D cameras are continuously refined using 3D sensor validation, and these validated results are fed back as training data. This self-reinforcing feedback mechanism allows the system to automatically improve its own performance over time without external intervention, producing reliable training data at the speed of system operation.

Inventive Principle:
Principle #23Feedback

4Adaptability or versatility

If a fleet of vehicles collects data, then rich information for adaptation is available, but traditional approaches cannot label data on-the-fly

Engineering Contradiction:
Improvelearning capabilityVSAvoidautomatic labeling capability
Core Design Contradiction:
Adaptability or versatilityVSExtent of automation

Solution Approach 1:

Each vehicle in the fleet serves itself by automatically generating labeled training data through its own sensor operations. The system requires no external labeling infrastructure - the combination of 2D camera detections and 3D sensor measurements automatically creates labeled datasets during normal driving, enabling continuous adaptation and knowledge sharing across the entire fleet without manual intervention.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11435750B2Method and system for object centric stereo via cross modality validation in autonomous driving vehicles
Publication Date: 2022.09.06 PLUSAI INC
  • US11435750B2 patent drawing
  • US11435750B2 patent drawing
  • US11435750B2 patent drawing

AI summary

The present teaching relates to a method, system, medium, and implementation of processing image data in an autonomous driving vehicle. Sensor data acquired by one or more types of sensors deployed on the vehicle are continuously received to provide different types of information about surrounding of the vehicle. Based on a first data set acquired by a first sensor of a first type of sensors at a specific time, an object is detected, where the first data set provides a first type of information with a first perspective. Depth information of the object is estimated via object centric stereo based on the object and a second data set, acquired at the specific time by a second sensor of the first type of sensors with a second perspective. The estimated depth information is further enhanced based on a third data set acquired by a third sensor of a second type of sensors at the specific time, providing a second type of information about the surrounding of the vehicle.