Object-Centric Stereo Depth Estimation for Real-Time Vehicle Perception
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer-aided perception systems in autonomous driving face limitations such as sensor sensitivity to lighting conditions, computational inefficiencies in depth estimation, and the need for manual data labeling, which hinders real-time adaptation and scalability in fleet operations.
Innovation Solution
The implementation of a virtual agent and an in-situ computer-aided perception system that uses object-centric stereo for depth estimation and cross-modal validation to automatically label training data on-the-fly, enabling continuous model adaptation and global model updates based on diverse data from a fleet of vehicles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional stereo using multiple cameras is used to estimate depth, then depth information can be obtained, but it is computationally expensive and slow, and cannot generate a depth map with adequate density
Solution Approach 1:
The patent segments the depth estimation process by first detecting objects using 2D sensors, then applying stereo depth estimation only to detected object regions rather than the entire image. This selective approach reduces computational load while maintaining depth accuracy for relevant objects, resolving the contradiction between measurement precision and productivity.
2Measurement precision
If 3D sensors such as LiDAR are used to acquire depth information, then depth data can be obtained, but the sensing technology has limitations of limited range and low data density
Solution Approach 1:
The patent merges data from multiple sensor types (2D cameras and 3D LiDAR) to compensate for individual sensor limitations. The 2D sensors provide wide coverage and high angular resolution, while the 3D sensor provides accurate depth measurements, creating a complementary system that achieves both extended range and high data density.
3Measurement precision
If manual or semi-manual labeling is used to create training data, then training data can be produced, but it is very difficult, time-consuming, and costly, making real-time adaptation impossible
Solution Approach 1:
The system performs self-service by automatically generating and labeling training data using its own sensor fusion and object detection capabilities. The vehicle labels its own operational data in real-time, eliminating the need for manual annotation and enabling continuous self-improvement without time loss.
4Adaptability or versatility
If a fleet of vehicles collects data from many different dynamics, then rich information source is available for continuous adaptation, but traditional approaches cannot label such data on-the-fly to produce meaningful training data
Solution Approach 1:
The patent implements a universal labeling framework that works across diverse fleet conditions and sensor configurations. The same object detection and stereo processing pipeline automatically labels training data regardless of the specific driving scenario, vehicle type, or environmental conditions, enabling automated adaptation across the entire fleet without requiring scenario-specific labeling approaches.
Data Source
AI summary
The present teaching relates to a method, system, medium, and implementation of processing image data in an autonomous driving vehicle. Sensor data acquired by one or more types of sensors deployed on the vehicle are continuously received. The sensor data provide different information about surrounding of the vehicle. Based on a first data set acquired by a first sensor of a first type of the one or more types of sensors at a specific time, an object is detected, where the first data set provides a first type of information about the surrounding of the vehicle. Depth information of the object is then estimated via object centric stereo at object level based on the object detected as well as a second data set acquired by a second sensor of the first type of the one or more types of sensors at the specific time. The second data set provides the first type of information about the surrounding of the vehicle with a different perspective as compared with the first data set.


