Object-Centric Stereo Depth Estimation for Autonomous Vehicle Perception

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer-aided perception systems in autonomous driving face limitations due to sensor sensitivity to environmental conditions and the need for manual data labeling, which is time-consuming and costly, making real-time adaptation and scalability challenging, especially when dealing with large fleets of vehicles.

Innovation Solution

The system employs object-centric stereo for depth estimation and cross-modal validation to automatically label training data on-the-fly, enabling continuous local and global model adaptation, using a combination of passive and active sensors to enhance depth estimation accuracy and reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional stereo using multiple cameras is used to estimate depth, then depth information can be obtained, but the computational cost is high and the speed is slow

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent extracts and focuses computational resources only on detected object regions rather than processing the entire image. By identifying object boundaries and limiting stereo matching operations to these specific regions, the system achieves accurate depth estimation for objects while significantly reducing overall computational burden and processing time.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies different processing strategies to different regions of the image. High-precision stereo matching is applied only to detected object regions where depth accuracy is critical, while other regions receive minimal or no processing. This local differentiation maintains measurement precision for objects while improving overall productivity.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If 3D sensors such as LiDAR are used to acquire depth information, then depth data can be obtained, but the range is limited and data density is low

Engineering Contradiction:
Improvedepth information qualityVSAvoidsensing range
Core Design Contradiction:
Measurement precisionVSArea of stationary object

Solution Approach 1:

The patent merges data from multiple sensor types (2D cameras and 3D LiDAR sensors) to overcome the limitations of each individual sensor. The 2D cameras provide wide-area detection and object identification, while the 3D sensors provide accurate depth information for detected objects. This combination extends the effective sensing range and increases data density in critical regions.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system uses 2D camera detection as an intermediary to guide 3D sensor operation. The 2D cameras first identify objects of interest, which then triggers targeted 3D depth measurement only for those specific objects. This intermediary approach allows the system to achieve high data density where needed while maintaining cost-effectiveness and extending effective range.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If manual or semi-manual labeling is used to create training data, then meaningful training data can be generated, but the process is time-consuming and costly

Engineering Contradiction:
Improvetraining data qualityVSAvoiddata preparation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs self-labeling by automatically generating training data from its own sensor inputs and detection algorithms. The multi-sensor fusion system automatically identifies objects, estimates their properties, and creates labeled training datasets without human intervention. This self-service approach maintains high training data quality while eliminating the time-consuming and costly manual labeling process.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system uses feedback from multi-sensor detection results to automatically create and refine training data. Detection outcomes from cameras and LiDAR are fed back into the training pipeline, where they are automatically labeled and used to improve the model. This closed-loop feedback mechanism ensures continuous improvement of training data quality without manual involvement.

Inventive Principle:
Principle #23Feedback

4Adaptability or versatility

If a fleet of vehicles collects data for continuous adaptation, then rich information can be gathered, but traditional approaches cannot label data on-the-fly

Engineering Contradiction:
Improvecontinuous adaptation capabilityVSAvoidautomatic data labeling
Core Design Contradiction:
Adaptability or versatilityVSExtent of automation

Solution Approach 1:

The fleet system performs self-labeling of collected data using its own multi-sensor detection and object recognition algorithms. Each vehicle automatically processes its sensor data, identifies objects, and generates labeled training datasets in real-time without requiring external manual labeling. This enables continuous adaptation as the fleet accumulates experience from diverse driving conditions.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system enables continuous data labeling and model adaptation without interruption. As vehicles collect data during normal operation, the system continuously processes this data, automatically labels it, and updates models in real-time. This continuous action allows the fleet to adapt continuously to new situations without stopping for manual data preparation.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS11392133B2Method and system for object centric stereo in autonomous driving vehicles
Publication Date: 2022.07.19 PLUSAI INC
  • US11392133B2 patent drawing
  • US11392133B2 patent drawing
  • US11392133B2 patent drawing

AI summary

The present teaching relates to a method, system, medium, and implementation of processing image data in an autonomous driving vehicle. Sensor data acquired by one or more types of sensors deployed on the vehicle are continuously received. The sensor data provide different information about surrounding of the vehicle. Based on a first data set acquired by a first sensor of a first type of the one or more types of sensors at a specific time, an object is detected, where the first data set provides a first type of information about the surrounding of the vehicle. Depth information of the object is then estimated via object centric stereo at object level based on the object detected as well as a second data set acquired by a second sensor of the first type of the one or more types of sensors at the specific time. The second data set provides the first type of information about the surrounding of the vehicle with a different perspective as compared with the first data set.