Depth Data Model Training with Upsampling and Loss Balancing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sensor systems in vehicles, particularly autonomous vehicles, face challenges with limited range and low density of data, which affects accurate object detection and environment traversal.
Innovation Solution
A machine learning model is trained using stereo image data and depth data to predict depth information, employing losses such as pixel loss, smoothing loss, structural similarity loss, and consistency loss, enabling self-supervised and supervised learning for improved depth data generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If sensors are used to capture sensor data, then object detection capability is enabled, but the range and data density remain limited
Solution Approach 1:
The patent introduces an intermediate processing stage that upsamples low-density sensor data to high-density depth maps. The machine learning model acts as an intermediary that transforms limited sensor inputs into dense depth information, effectively bridging the gap between sparse sensor data and comprehensive environmental understanding.
Solution Approach 2:
The patent changes the density parameter of the sensor data through upsampling operations. By transforming low-density sensor readings into high-density depth maps, the system effectively increases the quantity of usable data without adding physical sensors, resolving the contradiction between data density and measurement precision.
2Measurement precision
If multiple loss functions are used for training, then model accuracy improves, but training complexity increases
Solution Approach 1:
The patent segments the training objective into multiple distinct loss functions (pixel loss, smoothing loss, structural similarity loss, consistency loss), each addressing a specific aspect of depth prediction quality. This segmentation allows the complex training process to be broken down into manageable components that can be weighted and combined systematically.
Solution Approach 2:
The patent merges multiple loss functions into a unified training objective with learnable weights. By combining pixel loss, smoothing loss, structural similarity loss, and consistency loss into a single composite loss that the optimizer can minimize, the system manages training complexity while achieving high accuracy through multi-faceted optimization.
Data Source
AI summary
Techniques for training a machine learned (ML) model to determine depth data based on image data are discussed herein. Training can use stereo image data and depth data (e.g., lidar data). A first (e.g., left) image can be input to a ML model, which can output predicted disparity and/or depth data. The predicted disparity data can be used with second image data (e.g., a right image) to reconstruct the first image. Differences between the first and reconstructed images can be used to determine a loss. Losses may include pixel, smoothing, structural similarity, and/or consistency losses. Further, differences between the depth data and the predicted depth data and/or differences between the predicted disparity data and the predicted depth data can be determined, and the ML model can be trained based on the various losses. Thus, the techniques can use self-supervised training and supervised training to train a ML model.


