Object-Centric Stereo Depth Estimation for Autonomous Vehicle Perception
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer-aided perception systems in autonomous driving face limitations due to sensor sensitivity to environmental conditions and the need for manual data labeling, which is time-consuming and costly, making real-time adaptation and scalability challenging, especially when dealing with large fleets of vehicles.
Innovation Solution
The system employs object-centric stereo for depth estimation and cross-modal validation to automatically label training data on-the-fly, enabling continuous local and global model adaptation, using a combination of passive and active sensors to enhance depth estimation accuracy and reliability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional stereo using multiple cameras is used to estimate depth, then depth information can be obtained, but the computational cost is high and the speed is slow
Solution Approach 1:
The patent extracts and focuses computational resources only on detected object regions rather than processing the entire image. By identifying object boundaries and limiting stereo matching operations to these specific regions, the system achieves accurate depth estimation for objects while significantly reducing overall computational burden and processing time.
Solution Approach 2:
The system applies different processing strategies to different regions of the image. High-precision stereo matching is applied only to detected object regions where depth accuracy is critical, while other regions receive minimal or no processing. This local differentiation maintains measurement precision for objects while improving overall productivity.
2Measurement precision
If 3D sensors such as LiDAR are used to acquire depth information, then depth data can be obtained, but the range is limited and data density is low
Solution Approach 1:
The patent merges data from multiple sensor types (2D cameras and 3D LiDAR sensors) to overcome the limitations of each individual sensor. The 2D cameras provide wide-area detection and object identification, while the 3D sensors provide accurate depth information for detected objects. This combination extends the effective sensing range and increases data density in critical regions.
Solution Approach 2:
The system uses 2D camera detection as an intermediary to guide 3D sensor operation. The 2D cameras first identify objects of interest, which then triggers targeted 3D depth measurement only for those specific objects. This intermediary approach allows the system to achieve high data density where needed while maintaining cost-effectiveness and extending effective range.
3Reliability
If manual or semi-manual labeling is used to create training data, then meaningful training data can be generated, but the process is time-consuming and costly
Solution Approach 1:
The system performs self-labeling by automatically generating training data from its own sensor inputs and detection algorithms. The multi-sensor fusion system automatically identifies objects, estimates their properties, and creates labeled training datasets without human intervention. This self-service approach maintains high training data quality while eliminating the time-consuming and costly manual labeling process.
Solution Approach 2:
The system uses feedback from multi-sensor detection results to automatically create and refine training data. Detection outcomes from cameras and LiDAR are fed back into the training pipeline, where they are automatically labeled and used to improve the model. This closed-loop feedback mechanism ensures continuous improvement of training data quality without manual involvement.
4Adaptability or versatility
If a fleet of vehicles collects data for continuous adaptation, then rich information can be gathered, but traditional approaches cannot label data on-the-fly
Solution Approach 1:
The fleet system performs self-labeling of collected data using its own multi-sensor detection and object recognition algorithms. Each vehicle automatically processes its sensor data, identifies objects, and generates labeled training datasets in real-time without requiring external manual labeling. This enables continuous adaptation as the fleet accumulates experience from diverse driving conditions.
Solution Approach 2:
The system enables continuous data labeling and model adaptation without interruption. As vehicles collect data during normal operation, the system continuously processes this data, automatically labels it, and updates models in real-time. This continuous action allows the fleet to adapt continuously to new situations without stopping for manual data preparation.
Data Source
AI summary
The present teaching relates to a method, system, medium, and implementation of processing image data in an autonomous driving vehicle. Sensor data acquired by one or more types of sensors deployed on the vehicle are continuously received. The sensor data provide different information about surrounding of the vehicle. Based on a first data set acquired by a first sensor of a first type of the one or more types of sensors at a specific time, an object is detected, where the first data set provides a first type of information about the surrounding of the vehicle. Depth information of the object is then estimated via object centric stereo at object level based on the object detected as well as a second data set acquired by a second sensor of the first type of the one or more types of sensors at the specific time. The second data set provides the first type of information about the surrounding of the vehicle with a different perspective as compared with the first data set.


