360-degree point cloud densification method and device, vehicle and storage medium
By splicing the views of multiple cameras into 360-degree circumference views and combining radar point clouds to iteratively update the dense depth map, the problem of consistency between camera view angles is solved, and the density of 360-degree point clouds is achieved, with good generalization capabilities and economic benefits.
Patent Information
- Application Number
- CN202411810882.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-05-16
AI Technical Summary
The existing depth completion and point cloud density technologies are difficult to effectively solve the consistency problem of overlapping view angles between multiple cameras, and there are bottleneck problems such as scene generalization.
By splicing the single-view views of multiple cameras into 360-degree circumference views, the radar point cloud is projected to the image coordinate system of multiple cameras, cross-guided features and self-guided features are extracted, dense depth maps are iteratively updated, and finally back-projected the dense depth map to the radar coordinate system, realizing the density of 360-degree point clouds.
While achieving point cloud density, it solves the consistency problem of camera viewing angle cross-section, and adopts self-supervised training to save manual standard costs, improve economic benefits, and demonstrates good generalization capabilities under different data sets.
Smart Images

Figure CN120014213A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of vehicle technology, and in particular to a 360-degree point cloud densification method, device, vehicle and storage medium. Background Art
[0002] CBDES (Computing Brain Development System) consists of two parts: CBB (Computing Base Brain) and GAASD (Graphical ADAS-AD Software Developer). It aims to provide basic algorithm components and frameworks for various intelligent driving systems and is the key to future autonomous driving technology. CBDES provides powerful computing power for autonomous driving systems. It builds an integrated end-to-end autonomous driving model through advanced algorithms and deep learning technologies to achieve efficient perception and accurate recognition of complex driving scenarios. At the same time, CBDES transfers application algorithm development capabilities to OEMs, supports OEM engineers to quickly build their own defined intelligent driving systems, and perform function adaptation and parameter adjustment. Compared with existing products and development models, it has the three advantages of "high efficiency / high quality / generative".
[0003] Data closed-loop function: The training of the industry's more classic models (such as BEV+transformer, etc.) and most neural network algorithms still requires a large amount of data, and the vehicle operation scenarios may be somewhat different from the model training scenarios. Therefore, users need to adopt CBDES's pre-training and fine-tuning training strategy to update the weights of functional modules through data-driven. At this time, CBDES builds a data closed-loop system to realize data collection, labeling, simulation and functional module updates. Through the data closed-loop system, users can continuously iterate and optimize module functions from the original data to meet the driving needs of the current environment.
[0004] At present, LiDAR is still one of the most critical sensors in autonomous vehicles. It can scan the surrounding environment with high frequency and high precision, and generate massive amounts of laser point cloud data. However, the single-frame point cloud scanned by LiDAR manufacturers is still relatively sparse at a distance, and cannot meet the requirements of downstream tasks such as detection. Therefore, 360-degree densification of the collected laser point cloud data has become one of the core technologies to ensure the working performance of CBDES. The 360-degree densification of laser point clouds provides a basic guarantee for the vehicle's all-round environmental perception, because if the vehicle wants to understand the surrounding roads, obstacles, pedestrians and other information in real time, it is necessary to obtain enough point clouds reflected by the target to make accurate perception and judgment. At the same time, laser point clouds not only provide real-time perception, but can also be used to build high-precision maps. By processing laser point clouds, detailed three-dimensional maps can be generated to provide more accurate positioning and navigation support for autonomous vehicles.
[0005] In summary, 360-degree point cloud densification technology is an important component of CBDES and autonomous driving systems. However, the existing depth completion and point cloud densification technologies have not effectively solved the consistency problem of overlapping view areas between multiple cameras, and there are bottleneck problems such as scene generalization. Summary of the invention
[0006] The present application provides a 360-degree point cloud densification method, device, vehicle and storage medium to solve the problem that the related technology is difficult to solve the consistency problem of the overlapping areas of the viewing angles of multiple cameras, and there are problems such as scene generalization.
[0007] The first aspect of the present application provides a 360-degree point cloud densification method, comprising the following steps: stitching the single-view views of multiple cameras into a 360-degree surround view; projecting the radar point cloud to the image coordinate system of the multiple cameras to obtain a dense depth map corresponding to the 360-degree surround view, and sparsely analyzing the dense depth map to obtain a sparse depth map; extracting cross-guided features based on the 360-degree surround view and the sparse depth map, and iteratively updating the dense depth map according to the cross-guided features and the self-guided features, wherein the self-guided features are feature maps extracted from the dense depth map after each iterative update; obtaining a 360-degree surround depth map after the iterative update of the dense depth map, and back-projecting the 360-degree surround depth map to the radar coordinate system to obtain a densified 360-degree point cloud.
[0008] Optionally, the dense depth map is iteratively updated according to the cross-guided features and the self-guided features, including: inputting the cross-guided features and the self-guided features into the target dependent update unit of the Transformer, and the target dependent update unit outputs the iterative update to complete the dense depth map to obtain a 360-degree surround depth map.
[0009] Optionally, the target dependent update unit includes a channel dimension network, a first global attention network, a second global attention network and a residual structure network, wherein the channel dimension network fuses cross-guided features and self-guided features to obtain fused features, the first global attention network predicts the weights of the convolution kernel based on the fused features, the second global attention network predicts a spatially deformable convolution kernel based on the fused features, and the residual structure network iteratively updates the dense depth map based on the weights of the convolution kernel and the spatially deformable convolution kernel.
[0010] Optionally, before the cross-guided features and the self-guided features are input into the target dependent update unit of the Transformer, the method further includes: obtaining a training data set, wherein the training data set includes the true value of the dense depth map; using the training data set to train the target dependent update unit of the Transformer, and during the training process, using the true value of the dense depth map to supervise the training process, and using the loss function to back propagate to update the network parameters.
[0011] Optionally, cross-guided features are extracted based on the 360-degree surround view and the sparse depth map, including: inputting the 360-degree surround view and the sparse depth map into an RGB encoder, the RGB encoder extracting multi-scale RGB features from the 360-degree surround view and the sparse depth map; upsampling the multi-scale RGB features and inputting them into a depth encoder, the depth encoder fusing the multi-scale RGB features to obtain cross-guided features, wherein the resolution of the upsampled multi-scale RGB features is the same as the resolution of the dense depth map.
[0012] Optionally, the dense depth map is thinned to obtain a sparse depth map, including: inputting the dense depth map into an AFF network structure, and the AFF network structure obtains the sparse depth map through multiple adaptive downsampling.
[0013] Optionally, the single-view views of multiple cameras are stitched into a 360-degree surround view, including: performing image feature matching on the single-view views of multiple cameras; calculating an image transformation structure between images based on the matched image features; performing image mapping on the single-view views of multiple cameras using the image transformation structure, superimposing the mapped multiple perspective images, aligning feature points on the superimposed images, and selecting stitching seams; and obtaining a 360-degree surround view by stitching the aligned multiple perspective images according to the stitching seams.
[0014] The second aspect of the present application provides a 360-degree point cloud densification device, including: a stitching module, used to stitch the single-view views of multiple cameras into a 360-degree surround view; a projection sparse module, used to project the radar point cloud to the image coordinate system of multiple cameras to obtain a dense depth map corresponding to the 360-degree surround view, and sparse the dense depth map to obtain a sparse depth map; an extraction and update module, which extracts cross-guided features based on the 360-degree surround view and the sparse depth map, and iteratively updates the dense depth map according to the cross-guided features and the self-guided features, wherein the self-guided features are feature maps extracted from the dense depth map after each iterative update; a back-projection module, used to obtain a 360-degree surround depth map after the iterative update of the dense depth map is completed, and the 360-degree surround depth map is back-projected to the radar coordinate system to obtain a densified 360-degree point cloud.
[0015] The third aspect of the present application provides a vehicle, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the 360-degree point cloud densification method of the first aspect.
[0016] The fourth aspect of the present application provides a computer-readable storage medium having a computer program or instructions stored thereon. When the computer program or instructions are executed, the 360-degree point cloud densification method of the first aspect is implemented.
[0017] Therefore, this application includes the following beneficial effects:
[0018] The embodiment of the present application splices the single-view images of multiple cameras into a 360-degree surround view, projects the radar point cloud to the image coordinate system of multiple cameras to obtain a dense depth map corresponding to the 360-degree surround view, and then sparses the image to obtain a sparse depth map. The cross-guided features are extracted based on the 360-degree surround view and the sparse depth map, and the dense depth map is iteratively updated according to the cross-guided features and the self-guided features. After the dense depth map is iteratively updated, the 360-degree surround depth map is obtained, and then projected to the radar coordinate system to obtain a densified 360-degree point cloud. The densification of the point cloud solves the problem of consistency in the overlapping parts of the camera perspectives, and the self-supervised training is adopted to save the standard labor cost, improve the economic benefits, and have good generalization ability for the scene, achieving good performance under different data sets. Thus, the problem that the related technology is difficult to solve the consistency of the overlapping areas of the perspectives of multiple cameras, and there are problems such as scene generalization.
[0019] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0021] Figure 1 A flowchart of a 360-degree point cloud densification method provided according to an embodiment of the present application;
[0022] Figure 2 A flow chart of a 360-degree point cloud densification method according to an embodiment of the present application;
[0023] Figure 3 A schematic diagram of a 360-degree point cloud densification network structure provided according to an embodiment of the present application;
[0024] Figure 4 is an exemplary diagram of a 360-degree point cloud densification device according to an embodiment of the present application;
[0025] Figure 5 It is a schematic diagram of the structure of a vehicle provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0026] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0027] The following describes the 360-degree point cloud densification method, device, vehicle and storage medium of the embodiment of the present application with reference to the accompanying drawings. In view of the fact that the related technologies mentioned in the above background technology are difficult to solve the consistency of the overlapping areas of the perspectives of multiple cameras, and there are problems such as scene generalization, the present application provides a 360-degree point cloud densification method, in which the single-view views of multiple cameras are spliced into a 360-degree surround view, the radar point cloud is projected into the image coordinate system of multiple cameras to obtain a dense depth map corresponding to the 360-degree surround view, and the sparse depth map is obtained by sparseness, and the cross-guided features are extracted based on the 360-degree surround view and the sparse depth map. The dense depth map is iteratively updated according to the cross-guided features and the self-guided features, and the 360-degree surround depth map is obtained after the dense depth map is iteratively updated. The 360-degree surround depth map is projected into the radar coordinate system to obtain the densified 360-degree point cloud, which realizes the densification of the point cloud while solving the problem of consistency of the overlapping parts of the camera perspectives, adopts self-supervised training, saves labor standard costs, improves economic benefits, and has good generalization ability for the scene, and achieves good performance under different data sets. This solves the problem that related technologies are difficult to solve for the consistency of overlapping areas of view between multiple cameras and have problems with scene generalization.
[0028] Specifically, Figure 1 A schematic diagram of a process for densifying a 360-degree point cloud provided in an embodiment of the present application.
[0029] like Figure 1 As shown, the 360-degree point cloud densification method includes the following steps:
[0030] In step S101 , single-view images from multiple cameras are stitched into a 360-degree surround view.
[0031] It can be understood that the embodiment of the present application can capture single-view images through multiple cameras, and stitch the single-view images of multiple cameras into a 360-degree panoramic view. The stitching method will be described in detail below and will not be repeated here.
[0032] In an embodiment of the present application, the single-view views of multiple cameras are stitched into a 360-degree surround view, including: performing image feature matching on the single-view views of multiple cameras; calculating the image transformation structure between the images based on the matched image features; performing image mapping on the single-view views of multiple cameras using the image transformation structure, superimposing the mapped multiple perspective images, aligning feature points of the superimposed images, and selecting stitching seams; and obtaining a 360-degree surround view by stitching the aligned multiple perspective images according to the stitching seams.
[0033] Among them, the image feature matching of the single-view views of multiple cameras is to find the common feature points between different images, so as to determine the relative position relationship between them; the feature point alignment of the superimposed images can be achieved by using the APAP (As-Projective-As-Possible) algorithm.
[0034] It can be understood that the embodiments of the present application can extract image features from a single view captured by cameras from multiple different angles, perform feature matching, find out the common feature points between different images, and thus determine the relative position relationship between them. Based on these matched feature points, the transformation matrix between the images is calculated, and then the calculated transformation matrix is used to map the images from each perspective so that they can be correctly superimposed together in the same reference frame. The APAP algorithm is used to align the feature points of the superimposed images to ensure that the images of each perspective are seamlessly connected at the boundaries, and the aligned multiple images are spliced into a complete 360-degree panoramic view.
[0035] In step S102, the radar point cloud is projected onto the image coordinate systems of multiple cameras to obtain a dense depth map corresponding to the 360-degree surround view, and the dense depth map is thinned to obtain a sparse depth map.
[0036] Among them, a dense depth map refers to a data structure or image representation in which depth information is attached to each pixel of the image.
[0037] It can be understood that the embodiment of the present application can project the radar point cloud into the image coordinate system of multiple cameras to obtain a dense depth map corresponding to the 360-degree surround view to reduce the amount of data, and sparsely obtain a sparse depth map. The sparseness method will be described in detail below and will not be repeated here.
[0038] In an embodiment of the present application, a dense depth map is thinned to obtain a sparse depth map, including: inputting the dense depth map into an AFF network structure, and the AFF network structure obtains the sparse depth map through multiple adaptive downsampling.
[0039] Among them, the AFF (Adaptive Feature Fusion) network structure can quickly simplify the original point cloud into a small number of feature point clouds through multiple adaptive downsampling, and make the point cloud concentrated in the target area and relatively sparse in the background area, thereby achieving the sparseness of the dense depth map.
[0040] It can be understood that the embodiment of the present application can use the AFF network structure through multiple adaptive downsampling to quickly simplify the original point cloud into a small number of feature point clouds, and make the point cloud concentrated in the target area and relatively sparse in the background area, thereby achieving the sparseness of the dense depth map and obtaining a sparse depth map.
[0041] In step S103, a cross-guided feature is extracted based on the 360-degree surround view and the sparse depth map, and the dense depth map is iteratively updated according to the cross-guided feature and the self-guided feature, wherein the self-guided feature is a feature map extracted from the dense depth map after each iterative update.
[0042] Among them, the cross-guided features are features extracted from the 360-degree surround view and the sparse depth map, which are used to guide the update of the dense depth map.
[0043] It can be understood that the embodiments of the present application can extract cross-guided features from the 360-degree surround view and the sparse depth map, which contain rich scene geometry and texture information, and then extract self-guided features from the dense depth map after each iterative update, and iteratively update the dense depth map based on the cross-guided features and the self-guided features.
[0044] In an embodiment of the present application, cross-guided features are extracted based on a 360-degree surround view and a sparse depth map, including: inputting the 360-degree surround view and the sparse depth map into an RGB encoder, the RGB encoder extracting multi-scale RGB features from the 360-degree surround view and the sparse depth map; upsampling the multi-scale RGB features and inputting them into a depth encoder, the depth encoder fusing the multi-scale RGB features to obtain cross-guided features, wherein the resolution of the upsampled multi-scale RGB features is the same as the resolution of the dense depth map.
[0045] Among them, the RGB encoder is a neural network structure used to extract multi-scale RGB features from the input 360-degree surround view and sparse depth map; upsampling is an image processing technology used to increase the resolution of the feature map; the depth encoder is a neural network structure used to fuse the upsampled multi-scale RGB features to generate cross-guided features.
[0046] It can be understood that the embodiment of the present application extracts cross-guided features from the 360-degree surround view and the sparse depth map. First, the 360-degree surround view and the sparse depth map are input into the RGB encoder. The RGB encoder extracts multi-scale RGB features from these two inputs. These features contain visual information at different levels. The extracted multi-scale RGB features are upsampled and their resolution is ensured to be the same as that of the dense depth map. The upsampled multi-scale RGB features are input into the depth encoder. The depth encoder fuses these features to generate cross-guided features.
[0047] In an embodiment of the present application, a dense depth map is iteratively updated according to cross-guided features and self-guided features, including: inputting the cross-guided features and self-guided features into a target-dependent update unit of a Transformer, and the target-dependent update unit outputs an iteratively updated dense depth map to obtain a 360-degree surround depth map.
[0048] Among them, the target-dependent update unit is a TDU (Target-dependent Updated Unit) module, which can realize the position and weight of variable convolution, and further realize the iterative update of the depth map based on variable convolution.
[0049] It can be understood that the embodiment of the present application inputs the cross-guided features and the self-guided features into the target dependency update unit of the Transformer, which can capture the long-distance dependencies between the features, thereby more efficiently updating the dense depth map, and finally generating a high-quality 360-degree surround depth map after multiple iterations.
[0050] In an embodiment of the present application, the target dependent update unit includes a channel dimension network, a first global attention network, a second global attention network and a residual structure network, wherein the channel dimension network fuses cross-guided features and self-guided features to obtain fused features, the first global attention network predicts the weight of the convolution kernel based on the fused features, the second global attention network predicts the spatially deformable convolution kernel based on the fused features, and the residual structure network iteratively updates the dense depth map based on the weight of the convolution kernel and the spatially deformable convolution kernel.
[0051] It can be understood that the target dependent update unit of the embodiment of the present application includes a channel dimension network, a first global attention network, a second global attention network and a residual structure network. The channel dimension network can combine the cross-guided features with the self-guided features to generate fused features; the first global attention network uses the fused features to predict the weights of the convolution kernel; and the second global attention network predicts a spatially deformable convolution kernel based on the same fused features; the residual structure network can iteratively update the depth map based on the predicted convolution kernel weights and their spatially deformable characteristics, thereby gradually improving the accuracy of the depth map.
[0052] In an embodiment of the present application, before the cross-guided features and the self-guided features are input into the target dependent update unit of the Transformer, it also includes: obtaining a training data set, wherein the training data set includes the true value of the dense depth map; using the training data set to train the target dependent update unit of the Transformer, and during the training process, using the true value of the dense depth map to supervise the training process, and using the loss function to back propagate to update the network parameters.
[0053] Among them, the loss functions are L1 distance and L2 distance. L1 distance refers to Manhattan distance, which refers to the sum of the absolute value distances between two points in the coordinate system in the direction of the coordinate axis. L2 refers to Euclidean distance, which refers to the straight-line distance between two points in multidimensional space. Back propagation is an algorithm used to train neural networks. It calculates the gradient of the loss function to the network parameters and then adjusts the parameters to reduce the loss.
[0054] It can be understood that before inputting the cross-guided features and the self-guided features into the target-dependent update unit of the Transformer, the embodiment of the present application also needs to obtain a training data set including the true values of the 360-degree surround view, the sparse depth map, and the corresponding dense depth map, and use the training data set to train the target-dependent update unit of the Transformer. During the training process, the training process is supervised by the true value of the dense depth map, and the loss function is used to measure the difference between the dense depth map predicted by the model and the true value. Through the back-propagation algorithm, the network parameters are updated according to the gradient of the loss function to gradually optimize the performance of the model.
[0055] In step S104, after the dense depth map is iteratively updated, a 360-degree surround depth map is obtained, and the 360-degree surround depth map is back-projected to the radar coordinate system to obtain a densified 360-degree point cloud.
[0056] It can be understood that after completing the iterative update of the dense depth map in the embodiment of the present application, we obtain a panoramic depth map covering a 360-degree field of view, that is, a 360-degree surround depth map, and back-project the depth information in the 360-degree surround depth map into the radar coordinate system to generate a densified 360-degree point cloud.
[0057] According to the 360-degree point cloud densification method proposed in the embodiment of the present application, the single-view views of multiple cameras are spliced into a 360-degree surround view, the radar point cloud is projected to the image coordinate system of multiple cameras to obtain a dense depth map corresponding to the 360-degree surround view, and the sparse depth map is obtained by sparseness. Cross-guided features are extracted based on the 360-degree surround view and the sparse depth map, and the dense depth map is iteratively updated according to the cross-guided features and the self-guided features. After the dense depth map is iteratively updated, the 360-degree surround depth map is obtained, and the 360-degree surround depth map is projected to the radar coordinate system to obtain the densified 360-degree point cloud. The point cloud densification is achieved while solving the consistency problem of the cross-viewing angle of the camera. The self-supervised training method is adopted to save the standard labor cost, improve the economic benefit, and at the same time have a good generalization ability for the scene, and achieve good performance under different data sets.
[0058] The 360-degree point cloud densification method is further described below through a specific embodiment.
[0059] This embodiment designs a method for self-supervising 360° point cloud densification, such as Figure 2 As shown, the multi-view color images are first spliced into a 360° surround image, and the corresponding 360° depth image is constructed at the same time. This method uses the dense depth map of the point cloud projected to the image as the supervision signal of the network, and adopts the idea of AFF to thin the point cloud and then project it to the image coordinate system as the input sparse depth map. In order to make full use of the geometric and semantic information of the color image, this embodiment extracts the fused cross-guided features of the color image and the sparse depth image through the feature pyramid. The sparse depth map is pre-filled and then feature extracted as a self-guided feature. The self-guided features and cross-guided features of different scales are input into the TDU (target dependent update unit) module, and the depth map is iteratively updated step by step to obtain the final depth map estimation result. The final depth map estimation result is back-projected using the completed 360-degree depth map to achieve point cloud densification, where the network structure of the 360-degree point cloud densification is as shown below. Figure 3 As shown, the specific technical solution is as follows:
[0060] Step (1) stitching of multi-view images:
[0061] The single-view images of multiple cameras are stitched into a 360-degree surround view. The method used is: first, feature matching is performed according to the data set; then the transformation structure between images is calculated by matching features, and image mapping is achieved using the image transformation structure; finally, the APAP algorithm is used to align feature points of the superimposed image, and the graph cut method is used to automatically select the stitching seams, and fusion is achieved based on the multi-band blending strategy.
[0062] Step (2) Acquisition of sparse depth map:
[0063] The sparse depth map input to the network is obtained by projecting the radar point cloud onto the image and then sparsely processing it. The point cloud sparseness adopts the AFF network structure. Through multiple adaptive downsampling, the original point cloud can be quickly simplified into a small number of feature point clouds, and the point cloud is concentrated in the target area, while it is relatively sparse in the background area. After obtaining the sparse depth map of a single camera, the image mapping structure obtained by view stitching is used, and then the sparse depth maps of multiple cameras are stitched into a 360-degree surround sparse depth map and sent to the network.
[0064] Step (3) Initialization of dense depth map:
[0065] Through the camera calibration parameters and the lidar calibration parameters, the lidar point cloud is projected onto the image to obtain the depth map corresponding to the color image. The initialization and completion depth map is obtained through the dilation operation and the median Gaussian blur operation.
[0066] Step (4) Cross-feature extraction based on RGB image and sparse depth image:
[0067] Two sub-networks, a depth encoder and an RGB encoder, are used to extract features from the sparse depth map and the corresponding RGB image, respectively. The extracted multi-scale RGB features are injected into the depth encoder to integrate information from different modalities and obtain cross-guided features of different scales.
[0068] Step (5) Iteratively update the initial depth map:
[0069] The cross-guided features are first upsampled to the same resolution as the initial depth map and used in the iterative update process.
[0070] The TDU (Target-dependent Updated Unit) module implements the position and weight of variable convolution, and implements the iterative update of the depth map based on variable convolution. The TDU module is the core module of the algorithm, which integrates the cross features of the depth map features and the RGB map and sparse depth to calculate the position and weight of the spatial deformation kernel. Using TDU, you can learn the spatially varying kernel to calculate the depth map.
[0071] Specific calculation method: The self-guided feature comes from the depth map iterated t times, and is the feature extracted from the depth map; the cross-guided feature comes from the multimodal fusion feature extracted based on RGB and sparse depth map. After the two features are spliced in the channel dimension, they are connected to the Transformer-based deformable convolution generation network. Since the input image is a 360-degree panoramic view, t is based on the Transformer-based attention mechanism, which can learn global features through position encoding and calculate the weights and offsets of the downstream deformable convolution. The specific implementation method is to predict w×h×k in one of the branches of the Transformer output 2 The feature map of size corresponds to the weight of the convolution kernel, and the other branch predicts w×h×(2×k 2 ) size feature map, the channel dimension predicts the offset in the x and y directions respectively, and obtains the spatially deformable convolution kernel. The residual structure is used to update the iterative depth map because the results of the t+1 iteration and the results of the t iteration have a high correlation. Using the residual structure learning can make the network learn the change.
[0072] Step (6) uses the dense depth map as supervision to update the network parameters
[0073] Through multiple iterative updates in step (5), the final depth map result is obtained, the dense depth map is used as the true value for supervision, and the L1 distance and L2 distance are used as the loss function for back propagation to update the network weights.
[0074] Step (7) Surround view dense depth map splitting and point cloud densification
[0075] The dense surround depth map is mapped back to a single view using the image mapping structure obtained by view stitching in step (0). Then, the depth value of each pixel in the single camera depth map is traversed. The depth map is projected into the radar coordinate system using the calibration parameters of the camera and lidar to achieve densification of the point cloud.
[0076] Next, a 360-degree point cloud densification device proposed according to an embodiment of the present application is described with reference to the accompanying drawings.
[0077] Figure 4 It is a block diagram of a 360-degree point cloud densification device according to an embodiment of the present application.
[0078] like Figure 4 As shown, the 360-degree point cloud densification device 10 includes: a stitching module 201 , a projection sparse module 202 , an extraction and updating module 203 and a back-projection module 204 .
[0079] Among them, the stitching module 201 is used to stitch the single-view views of multiple cameras into a 360-degree surround view; the projection sparse module 202 is used to project the radar point cloud to the image coordinate system of multiple cameras to obtain a dense depth map corresponding to the 360-degree surround view, and to thin out the dense depth map to obtain a sparse depth map; the extraction and update module 203 extracts cross-guided features based on the 360-degree surround view and the sparse depth map, and iteratively updates the dense depth map according to the cross-guided features and the self-guided features, wherein the self-guided features are feature maps extracted from the dense depth map after each iterative update; the back-projection module 204 is used to obtain a 360-degree surround depth map after the iterative update of the dense depth map is completed, and the 360-degree surround depth map is back-projected to the radar coordinate system to obtain a densified 360-degree point cloud.
[0080] In an embodiment of the present application, the stitching module 201 is further used to: stitch the single-view views of multiple cameras into a 360-degree panoramic view, including performing image feature matching on the single-view views of multiple cameras; calculating the image transformation structure between the images based on the matched image features; using the image transformation structure to perform image mapping on the single-view views of multiple cameras, superimposing the mapped multiple perspective images, aligning the feature points of the superimposed images, and selecting stitching seams; stitching the aligned multiple perspective images according to the stitching seams to obtain a 360-degree panoramic view.
[0081] In an embodiment of the present application, the projection sparse module 202 is further used to: thin out the dense depth map to obtain a sparse depth map, including: inputting the dense depth map into the AFF network structure, and the AFF network structure obtains the sparse depth map through multiple adaptive downsampling.
[0082] In an embodiment of the present application, the extraction and update module 203 is further used to iteratively update the dense depth map according to the cross-guided features and the self-guided features, including inputting the cross-guided features and the self-guided features into the target dependent update unit of the Transformer, and the target dependent update unit outputs the iterative update to complete the dense depth map to obtain a 360-degree surround depth map.
[0083] In an embodiment of the present application, the target dependent update unit includes a channel dimension network, a first global attention network, a second global attention network and a residual structure network, wherein the channel dimension network fuses cross-guided features and self-guided features to obtain fused features, the first global attention network predicts the weight of the convolution kernel based on the fused features, the second global attention network predicts the spatially deformable convolution kernel based on the fused features, and the residual structure network iteratively updates the dense depth map based on the weight of the convolution kernel and the spatially deformable convolution kernel.
[0084] In an embodiment of the present application, before the cross-guided features and the self-guided features are input into the target dependent update unit of the Transformer, it also includes: obtaining a training data set, wherein the training data set includes the true value of the dense depth map; using the training data set to train the target dependent update unit of the Transformer, and during the training process, using the true value of the dense depth map to supervise the training process, and using the loss function to back propagate to update the network parameters.
[0085] In an embodiment of the present application, the extraction and update module 203 is further used to: extract cross-guided features based on the 360-degree surround view and the sparse depth map, including: inputting the 360-degree surround view and the sparse depth map into an RGB encoder, and the RGB encoder extracts multi-scale RGB features from the 360-degree surround view and the sparse depth map; upsampling the multi-scale RGB features and inputting them into a depth encoder, and the depth encoder fuses the multi-scale RGB features to obtain cross-guided features, wherein the resolution of the upsampled multi-scale RGB features is the same as the resolution of the dense depth map.
[0086] It should be noted that the above explanation of the embodiment of the 360-degree point cloud densification method is also applicable to the 360-degree point cloud densification device of this embodiment, and will not be repeated here.
[0087] According to the 360-degree point cloud densification device proposed in the embodiment of the present application, through the coordinated action of the stitching module, the projection sparse module, the extraction and update module and the back-projection module, the single-view views of multiple cameras can be stitched into a 360-degree surround view, and the radar point cloud can be projected into the image coordinate system of multiple cameras to obtain a dense depth map corresponding to the 360-degree surround view, and the sparse depth map can be obtained by thinning, and cross-guided features are extracted based on the 360-degree surround view and the sparse depth map. The dense depth map is iteratively updated according to the cross-guided features and the self-guided features. After the dense depth map is iteratively updated, the 360-degree surround depth map is obtained, and the 360-degree surround depth map is projected to the radar coordinate system to obtain a densified 360-degree point cloud, which realizes the densification of the point cloud while solving the problem of consistency of the cross-section of camera perspectives, adopts self-supervised training, saves labor standard costs, improves economic benefits, and has good generalization ability for scenes, and achieves good performance under different data sets.
[0088] Figure 5 A schematic diagram of the structure of a vehicle provided in an embodiment of the present application. The vehicle may include:
[0089] A memory 301 , a processor 302 , and a computer program stored in the memory 301 and executable on the processor 302 .
[0090] When the processor 302 executes the program, the 360-degree point cloud densification method provided in the above embodiment is implemented.
[0091] Furthermore, the vehicle also includes:
[0092] The communication interface 303 is used for communication between the memory 301 and the processor 302 .
[0093] The memory 301 is used to store computer programs that can be run on the processor 302 .
[0094] The memory 301 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.
[0095] If the memory 301, the processor 302 and the communication interface 303 are implemented independently, the communication interface 303, the memory 301 and the processor 302 can be connected to each other through a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0096] Optionally, in a specific implementation, if the memory 301, the processor 302 and the communication interface 303 are integrated on a chip, the memory 301, the processor 302 and the communication interface 303 can communicate with each other through an internal interface.
[0097] The processor 302 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.
[0098] An embodiment of the present application also provides a computer-readable storage medium having a computer program or instruction stored thereon. When the computer program or instruction is executed, the above-mentioned 360-degree point cloud densification method is implemented.
[0099] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0100] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0101] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0102] It should be understood that the various parts of the present application can be implemented in hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, the steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array, a field programmable gate array, etc.
[0103] A person of ordinary skill in the art may understand that all or part of the steps carried by the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the above-mentioned program may be stored in a computer-readable storage medium, which, when executed, includes one of the steps of the method embodiment or a combination thereof.
[0104] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A 360-degree point cloud densification method, characterized in that: The following steps are involved: Stitching the single-view images from multiple cameras into a 360-degree surround view; Projecting the radar point cloud to the image coordinate systems of multiple cameras to obtain a dense depth map corresponding to the 360-degree surround view, and thinning the dense depth map to obtain a sparse depth map; Extracting cross-guided features based on the 360-degree surround view and the sparse depth map, and iteratively updating the dense depth map according to the cross-guided features and the self-guided features, wherein the self-guided features are feature maps extracted from the dense depth map after each iterative update; After the dense depth map is iteratively updated, a 360-degree surround depth map is obtained, and the 360-degree surround depth map is back-projected to the radar coordinate system to obtain a densified 360-degree point cloud.
2. The 360-degree point cloud densification method according to claim 1, characterized in that: The iterative updating of the dense depth map according to the cross-guidance feature and the self-guidance feature comprises: The cross-guided features and the self-guided features are input into the target-dependent update unit of the Transformer, and the target-dependent update unit outputs an iteratively updated dense depth map to obtain a 360-degree surround depth map.
3. The 360-degree point cloud densification method according to claim 2, characterized in that: The target dependent update unit includes a channel dimension network, a first global attention network, a second global attention network and a residual structure network, wherein: The channel dimension network fuses the cross-guided features and the self-guided features to obtain a fused feature, the first global attention network predicts the weight of the convolution kernel based on the fused feature, the second global attention network predicts a spatially deformable convolution kernel based on the fused feature, and the residual structure network iteratively updates a dense depth map based on the weight of the convolution kernel and the spatially deformable convolution kernel.
4. The 360-degree point cloud densification method according to claim 2, characterized in that: Before the cross-guided feature and the self-guided feature are input into the target dependency update unit of the Transformer, the method further includes: Acquire a training data set, wherein the training data set includes a true value of a dense depth map; The target dependent update unit of the Transformer is trained using the training data set, and during the training process, the true value of the dense depth map is used to supervise the training process, and the network parameters are updated using the loss function back propagation.
5. The 360-degree point cloud densification method according to claim 1, characterized in that: The extracting of cross-guidance features based on the 360-degree surround view and the sparse depth map comprises: Inputting the 360-degree surround view and the sparse depth map into an RGB encoder, wherein the RGB encoder extracts multi-scale RGB features from the 360-degree surround view and the sparse depth map; The multi-scale RGB features are upsampled and input into a depth encoder, and the depth encoder fuses the multi-scale RGB features to obtain cross-guided features, wherein the resolution of the upsampled multi-scale RGB features is the same as the resolution of the dense depth map.
6. The 360-degree point cloud densification method according to claim 1, characterized in that: The step of thinning the dense depth map to obtain a sparse depth map includes: The dense depth map is input into the AFF network structure, and the AFF network structure obtains a sparse depth map through multiple adaptive downsampling.
7. The 360-degree point cloud densification method according to claim 1, characterized in that: The step of stitching the single-view images of multiple cameras into a 360-degree surround view includes: performing image feature matching on the single-view images of the multiple cameras; Calculate the image transformation structure between images based on the matched image features; Performing image mapping on the single-view images of the multiple cameras using the image transformation structure, superimposing the mapped multiple perspective images, aligning feature points of the superimposed images, and selecting a stitching seam; A 360-degree panoramic view is obtained by splicing and aligning the multiple perspective images at the splicing seams.
8. A 360-degree point cloud densification device, characterized in that: include: A stitching module, used to stitch the single-view images of multiple cameras into a 360-degree surround view; A projection sparse module is used to project the radar point cloud to the image coordinate system of multiple cameras to obtain a dense depth map corresponding to the 360-degree surround view, and to thin out the dense depth map to obtain a sparse depth map; an extraction and updating module, which extracts cross-guided features based on the 360-degree surround view and the sparse depth map, and iteratively updates the dense depth map according to the cross-guided features and self-guided features, wherein the self-guided features are feature maps extracted from the dense depth map after each iterative update; The back-projection module is used to obtain a 360-degree surround depth map after iteratively updating the dense depth map, and back-project the 360-degree surround depth map to the radar coordinate system to obtain a densified 360-degree point cloud.
9. A vehicle, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the 360-degree point cloud densification method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed, the 360-degree point cloud densification method described in any one of claims 1-7 is implemented.