Image-Based Depth Data and Bounding Boxes for Sparse Lidar
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Autonomous vehicles face challenges in accurately detecting and navigating through environments due to limited sensor range and low density of data, particularly in areas with sparse lidar data, which can lead to inaccuracies in localization and object detection.
Innovation Solution
A machine-learning model is trained using image data and lidar data as ground truth to generate depth data, enabling the vehicle to determine object locations, relative depths, and trajectories more accurately, even in areas with sparse data, by associating image-based depth data with three-dimensional bounding boxes and using relative depth data to supplement sparse lidar measurements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If sensors are used to capture sensor data for object detection, then object detection capability is provided, but sensor range is limited and data density is low
Solution Approach 1:
The patent introduces an intermediary machine-learned model that processes image data to generate depth data and bounding boxes. This model acts as a mediator between the image sensor data and the localization system, enabling the system to derive depth information and object boundaries without relying solely on sparse lidar measurements, thereby improving measurement precision while working around data density limitations
Solution Approach 2:
The patent creates a copied representation of the environment by generating a depth map from image data through machine learning. This depth map serves as a virtual copy of the real-world depth information that would otherwise require dense lidar scanning, allowing the system to achieve accurate depth perception without the computational and hardware burden of high-density lidar data
2Measurement precision
If lidar data is used for localization, then localization accuracy is improved, but processing power and memory requirements increase
Solution Approach 1:
The patent employs a machine-learned model that processes image data to generate depth information on-demand. Rather than continuously processing and storing large volumes of lidar data, the system uses the more energy-efficient image sensor combined with a trained model that generates necessary depth data only when needed for localization, significantly reducing processing power and memory requirements while maintaining accuracy
Solution Approach 2:
The patent substitutes the mechanical lidar scanning system with an optical image capture system combined with computational processing. The machine-learned model replaces the physical lidar measurements by inferring depth from image data, thereby reducing the energy consumption associated with operating and processing dense lidar point clouds
3Quantity of substance
If machine-learned model generates depth data from image data, then data density is improved, but model training complexity increases
Solution Approach 1:
The patent performs preliminary action by training the machine-learned model offline using paired image and depth data before deployment. During actual operation, the pre-trained model rapidly generates depth data from image input without requiring complex real-time training, thus achieving high data density while keeping operational complexity low. The heavy computational burden is shifted to the offline training phase
4Area of stationary object
If sensors operate in sparse data regions, then coverage area is expanded, but measurement precision decreases
Solution Approach 1:
The patent changes the parameter of depth data generation by using a machine-learned model that can operate effectively with limited image data in sparse regions. The model adapts to varying data conditions by generating plausible depth estimates even when image features are sparse, maintaining acceptable measurement precision across expanded coverage areas where traditional lidar would fail
Data Source
AI summary
A vehicle can use an image sensor to both detect objects and determine depth data associated with the environment the vehicle is traversing. The vehicle can capture image data and lidar data using the various sensors. The image data can be provided to a machine-learned model trained to output depth data of an environment. Such models may be trained, for example, by using lidar data and/or three-dimensional map data associated with a region in which training images and/or lidar data were captured as ground truth data. The autonomous vehicle can further process the depth data and generate additional data including localization data, three-dimensional bounding boxes, and relative depth data and use the depth data and/or the additional data to autonomously traverse the environment, provide calibration/validation for vehicle sensors, and the like.


