Upsampling 3D Point Clouds via CNN and CRF Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Autonomous vehicles face challenges in generating high-resolution 3D point clouds for real-time scene reconstruction due to the high cost and limited availability of high-resolution LIDAR equipment, which is essential for object segmentation, detection, tracking, and classification.
Innovation Solution
A method that combines low-resolution LIDAR data with a calibrated multi-camera system using deep learning techniques to generate high-resolution 3D point clouds, employing convolutional neural networks (CNNs) and conditional random field (CRF) models to upscale and refine depth maps from camera images and LIDAR data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high-resolution LIDAR equipment is used, then measurement precision of 3D point clouds is improved, but device cost increases
Solution Approach 1:
The patent combines low-resolution LIDAR depth data with high-resolution camera images to generate high-resolution 3D point clouds. The CNN model fuses the depth map from LIDAR with the color image from camera, transferring high-frequency texture details from the image to the point cloud, achieving high resolution without expensive LIDAR hardware.
Solution Approach 2:
The patent uses camera images as a proxy to copy high-resolution surface details and transfers them to the LIDAR-generated point cloud structure. The CRF model refines this by copying accurate depth information from LIDAR while preserving the high-resolution appearance from the camera image, creating a high-resolution point cloud from lower-cost sensors.
2Measurement precision
If high-resolution LIDAR data is generated, then object detection precision is improved, but data processing time increases
Solution Approach 1:
The patent performs preliminary downsampling of the LIDAR depth map to a lower resolution that matches the camera image resolution before processing. The CNN and CRF models then efficiently upsample and refine this pre-processed data, reducing the computational burden compared to processing full-resolution LIDAR data directly while maintaining detection precision.
Solution Approach 2:
The patent replaces the mechanical/computational process of generating high-resolution point clouds directly from high-resolution LIDAR with a learning-based approach. The CNN and CRF models substitute traditional point cloud processing algorithms, efficiently generating high-resolution output from lower-resolution input through learned patterns rather than computationally intensive direct processing.
3Ease of manufacture
If low-resolution LIDAR data is used, then device cost is reduced, but measurement precision deteriorates
Solution Approach 1:
The patent introduces camera images as an intermediary medium to bridge the resolution gap. The high-resolution camera image serves as a mediator that provides missing high-frequency details, which are then integrated with the low-resolution LIDAR depth information through the CNN model to produce high-resolution 3D point clouds from low-cost sensors.
Solution Approach 2:
The patent changes the resolution parameter of the LIDAR depth map dynamically through downsampling and upsampling operations controlled by the neural network. The system transforms the low-resolution LIDAR data into multiple resolution stages, using the CRF model to refine the final high-resolution output, effectively changing resolution parameters to overcome the limitations of low-cost LIDAR hardware.
Data Source
AI summary
In one embodiment, a method or system generates a high resolution 3-D point cloud to operate an autonomous driving vehicle (ADV) from a low resolution 3-D point cloud and camera-captured image(s). The system receives a first image captured by a camera for a driving environment. The system receives a second image representing a first depth map of a first point cloud corresponding to the driving environment. The system downsamples the second image by a predetermined scale factor until a resolution of the second image reaches a predetermined threshold. The system generates a second depth map by applying a convolutional neural network (CNN) model to the first image and the downsampled second image, the second depth map having a higher resolution than the first depth map such that the second depth map represents a second point cloud perceiving the driving environment surrounding the ADV.


