Point cloud completion method and device, equipment and storage medium

By combining the information of images and point clouds in the point cloud completion process, the mapping relationship between images and point clouds is used to merge features to generate dense point clouds, which solves the problem of low point cloud accuracy caused by the environment of lidar, and achieves higher point cloud completion accuracy and completeness.

CN120259094APending Publication Date: 2025-07-04BEIJING XINGYUN DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410020396.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-04
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the point cloud completion process is affected by the environment by the lidar, resulting in low accuracy of the point cloud after completion.

Method used

By combining the information of the image and point cloud, the mapping relationship between pixel points in the image and points in the point cloud is used to fuse the image features and point cloud features to generate a more dense target point cloud.

Benefits of technology

It improves the accuracy and completeness of the point cloud completion process, provides more accurate and complete reference information, and enhances the feature dimensions of point cloud features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259094A_ABST
    Figure CN120259094A_ABST
Patent Text Reader

Abstract

The invention provides a point cloud completion method and device, equipment and a storage medium, relates to the technical field of computer vision, and is used for solving the problem of low accuracy of a point cloud after completion, and the method comprises the steps: obtaining an initial image feature and an initial point cloud feature, the initial image feature being an image feature of an initial image of a to-be-detected region, and the initial point cloud feature being an image feature of the to-be-detected region; the initial point cloud feature is the point cloud feature of the initial point cloud of the to-be-detected area. According to the mapping relation between the points in the initial point cloud and the pixel points in the initial image, the initial image features and the initial point cloud features are fused, and target point cloud features are obtained. And generating a target point cloud according to the target point cloud features, the density of the target point cloud being greater than the density of the initial point cloud.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular, to a point cloud completion method, apparatus, device, and storage medium. Background Art

[0002] A point cloud is a set of sampled points obtained by a lidar, and the point cloud carries point cloud features such as single-point features, local features, and global features. Therefore, the point cloud is often applied to the field of visual recognition, for example, scenarios such as object detection, pose estimation, scene analysis, and virtual reality.

[0003] Currently, due to the defects of the lidar itself or the occlusion of objects, the point cloud may be incomplete. In order to increase the integrity of the point cloud, the traditional point cloud completion process usually needs to restore the missing part of the point cloud based on the point cloud features in the point cloud, and then obtain the completed point cloud, realizing the expansion of the sparse point cloud to the dense point cloud.

[0004] However, in the above technical solution, the point cloud completion process only realizes the expansion of the sparse point cloud to the dense point cloud based on the point cloud features in the point cloud collected by the lidar. Under the influence of the acquisition environment (such as light, humidity, etc.) of the lidar, the point cloud features in the point cloud collected by the lidar may be inaccurate, reducing the accuracy of the completed point cloud. Summary of the Invention

[0005] Embodiments of the present disclosure provide a point cloud completion method, apparatus, device, and storage medium, which are used to solve the problem of low accuracy of the completed point cloud.

[0006] On the one hand, a point cloud completion method is provided. The method includes: a point cloud completion apparatus obtains an initial image feature and an initial point cloud feature, where the initial image feature is an image feature of an initial image of a region to be detected, and the initial point cloud feature is a point cloud feature of an initial point cloud of the region to be detected. The point cloud completion apparatus fuses the initial image feature and the initial point cloud feature according to the mapping relationship between the points in the initial point cloud and the pixel points in the initial image to obtain a target point cloud feature. The point cloud completion apparatus generates a target point cloud according to the target point cloud feature, and the density of the target point cloud is greater than that of the initial point cloud.

[0007] On the other hand, a point cloud completion apparatus is provided. The apparatus includes: an acquisition module and a processing module.

[0008] An acquisition module is configured to acquire an initial image feature and an initial point cloud feature. The initial image feature is an image feature of an initial image of a region to be detected, and the initial point cloud feature is a point cloud feature of an initial point cloud of the region to be detected. A processing module is configured to fuse the initial image feature and the initial point cloud feature according to a mapping relationship between points in the initial point cloud and pixel points in the initial image, so as to obtain a target point cloud feature. The processing module is further configured to generate a target point cloud according to the target point cloud feature, and the density of the target point cloud is greater than that of the initial point cloud.

[0009] In another aspect, provided is an electronic device, including: a memory and a processor. The memory and the processor are coupled. The memory is configured to store a computer program. When the processor executes the computer program, the point cloud completion method according to any one of the above embodiments is implemented.

[0010] In another aspect, provided is a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the point cloud completion method according to any one of the above embodiments is implemented.

[0011] In the embodiments of the present disclosure, according to the mapping relationship between pixel points in an image and points in a point cloud, the image feature and the point cloud feature are fused, and according to the fused point cloud feature, the point cloud is completed to obtain a point cloud with a greater density. In this way, through the multi-modal fusion between the point cloud and the image, the feature dimension of the point cloud feature is increased, and further the feature information of the point cloud feature is increased, providing more accurate and complete reference information for the point cloud completion process and improving the accuracy of the completed point cloud. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the present disclosure, the drawings required to be used in some embodiments of the present disclosure will be briefly introduced below. Obviously, the drawings in the following description are only the drawings of some embodiments of the present disclosure, and those of ordinary skill in the art can also obtain other drawings according to these drawings.

[0013] Figure 1 FIG. 18 is a schematic diagram of a communication system provided for some embodiments of the present disclosure;

[0014] Figure 2 FIG. 22 is a schematic flowchart of a point cloud completion method provided for some embodiments of the present disclosure;

[0015] Figure 3 FIG. 26 is a schematic diagram of an example of point cloud feature extraction provided for some embodiments of the present disclosure;

[0016] Figure 4 FIG. 30 is a schematic flowchart of another point cloud completion method provided for some embodiments of the present disclosure;

[0017] Figure 5 An example schematic diagram of point cloud completion provided by some embodiments of the present disclosure;

[0018] Figure 6 A structural schematic diagram of a point cloud completion device provided by some embodiments of the present disclosure;

[0019] Figure 7 A structural schematic diagram of a point cloud completion device provided by some embodiments of the present disclosure. Detailed implementation manners

[0020] Next, the technical solutions in the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.

[0021] It should be noted that in the present disclosure, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present disclosure should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner.

[0022] Hereinafter, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0023] In the description of the present disclosure, unless otherwise specified, " / " means "or". For example, A / B may represent A or B. The "and / or" herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, "at least one" means one or more, and "a plurality" means two or more.

[0024] Before introducing the point cloud completion method provided by the embodiments of the present disclosure in detail, the implementation environment and application scenarios of the embodiments of the present disclosure will be introduced first.

[0025] First, the application scenarios of the embodiments of the present disclosure will be introduced.

[0026] In traditional point cloud completion work, the usual learning objective is to map the incomplete point cloud X to the complete point cloud Y' through a mapping function F(X) = Y'. However, this task is ill-posed in theory because an incomplete point cloud can correspond to multiple distributions of complete point clouds and there is no unique solution.

[0027] The purpose of introducing images as auxiliary input information is to improve the interpretability of point cloud completion and avoid the problem of finding one solution from countless solutions. By learning the mapping function F(X, I) = Y', the prediction process can be better understood. Images not only help improve the prediction results as auxiliary information but also serve as constraint information to reduce the possible solution space. Methods based solely on images are vulnerable to illumination, and there are significant differences between photos taken at different times and seasons. Methods based solely on point clouds are limited by the relatively small amount of information in the point clouds. By combining the information of images and point clouds, the advantages can be complementary. The introduction of such image information is feasible in practical scenarios. Therefore, introducing images as auxiliary input information into the point cloud completion task helps improve interpretability, reduce the solution space, and overcome the limitations faced by single methods that use only images or point clouds.

[0028] In recent years, the multi-modal information fusion of images and point clouds has been widely studied in various fields. For example, Barea et al. utilized the point cloud data collected by images and 3D lidar in vehicle detection and positioning tasks. Lu et al. designed a fusion network for object position recognition. In the 3D point cloud semantic segmentation task, fusing the point cloud data captured by images and lidar can improve the learning effect. In the field of autonomous driving, due to the drive of application requirements, the fusion of images and point clouds has become a research trend and has spawned a large number of related works. However, currently, no research work has been found in the field of point cloud completion that applies the fusion technology of images and point clouds to this task. Therefore, how to combine the fusion of images and point clouds in the point cloud completion task to improve the accuracy and integrity of point cloud completion has become a technical problem to be solved urgently.

[0029] To solve the above problems, the embodiments of the present disclosure provide a point cloud completion method. The point cloud completion method provided by the embodiments of the present disclosure is applied to the completion scenario or field from incomplete point clouds to complete point clouds, including:

[0030] Computer vision and graphics: The point cloud completion method can be used to process and reconstruct 3D objects, scenes, and environments. This is of great significance for computer vision tasks such as object detection, pose estimation, scene analysis, and virtual reality.

[0031] Autonomous Driving and Robotics: In autonomous driving and robotics, point cloud completion can be used to perceive and understand the environment. By completing the missing point cloud data, the accuracy and robustness of key tasks such as obstacle detection, path planning, and navigation can be improved. Autonomous vehicles need to accurately perceive and understand the environment, including generating a complete 3D point cloud map in real time. Point cloud completion methods based on cross-modal information fusion can help fill in the missing and noisy parts in the point cloud and improve the perception ability of the autonomous driving system. For example, Waymo (a subsidiary of Alphabet) uses point cloud completion technology in its autonomous vehicles to improve the accuracy and robustness of environmental perception.

[0032] Medical Image Processing: Point cloud data in medical images can be reconstructed and repaired through completion methods. This has potential applications in medical image analysis, lesion detection, and surgical planning. Through point cloud completion methods based on cross-modal information fusion, 3D reconstruction and repair of medical images can be performed, improving doctors' diagnostic capabilities and treatment effects. For example, a research team used point cloud completion methods to reconstruct brain MRI scan data to assist in neurosurgical planning.

[0033] Cultural Relic Protection and Cultural Heritage: Point cloud completion methods can be used to reconstruct and repair damaged cultural relics and cultural heritage. By completing the missing point cloud data, the integrity of historical relics can be restored and digitally protected. In the field of cultural relic protection and cultural heritage, point cloud completion methods based on cross-modal information fusion can be used to reconstruct damaged or incomplete cultural relics. By completing the missing point cloud data, the integrity of the cultural relics can be restored and digitally protected. For example, a research team used point cloud completion methods to repair and reconstruct damaged sculptures and buildings to protect and display cultural heritage.

[0034] Augmented Reality and Virtual Reality: In augmented reality and virtual reality applications, point cloud completion can be used to generate realistic virtual environments and object models. By completing the point cloud data, a more real and immersive user experience can be provided.

[0035] These are some practical cases where point cloud completion methods based on cross-modal information fusion have been applied. With the continuous development of technology, it is expected that this method will be adopted in more fields and applications.

[0036] This disclosure uses the information combining images and point clouds to complete the defective point cloud to a complete point cloud. According to the mapping relationship between the pixel points in the image and the points in the point cloud, the image features and the point cloud features are fused, and based on the fused point cloud features, the point cloud is completed to obtain a point cloud with a greater density. In this way, through the multi-modal fusion between the point cloud and the image, the feature dimension of the point cloud features is increased, and thus the feature information of the point cloud features is increased, providing more accurate and complete reference information for the point cloud completion process and improving the accuracy of the completed point cloud.

[0037] That is to say, by combining the information of images and point clouds, the rich visual features of images and the geometric information of point clouds can be used to complement each other, so as to better restore the missing part of the point cloud. Such a fusion method may provide more context information and constraints, which helps to generate a more realistic and accurate complete point cloud.

[0038] The implementation environment of the embodiments of this disclosure will be introduced below.

[0039] As Figure 1 shown, it is a schematic diagram of a communication system provided by an embodiment of this disclosure. The communication system may include: a point cloud completion device 101, an image acquisition device 102, and a point cloud acquisition device 103. Among them, the acquisition object (or area) of the image acquisition device 102 and the acquisition object of the point cloud acquisition device 103 are the same object, and the point cloud completion device 101 communicates with the image acquisition device 102 and the point cloud acquisition device 103 in a wired / wireless manner respectively.

[0040] Among them, the image acquisition device 102 can acquire images and send the acquired images to the point cloud completion device 101.

[0041] Similarly, the point cloud acquisition device 103 can acquire point clouds and send the acquired point clouds to the point cloud completion device 101.

[0042] The point cloud completion device 101 can extract features from the images from the image acquisition device 102 and the point clouds from the point cloud acquisition device 103 to obtain image features and point cloud features, and through the feature fusion of the point cloud and the image, obtain the fused point cloud features. Then, the point cloud completion device 101 can expand the point clouds acquired from the point cloud acquisition device 103 based on the fused point cloud features to obtain a point cloud with a greater density, thereby realizing the completion of the point cloud.

[0043] In some embodiments, a trained point cloud completion model 104 is deployed in the point cloud completion device 101, and the trained point cloud completion model 104 may include: a point cloud encoder 105, an image encoder 106, a feature fusion module 107, and a point cloud decoder 108.

[0044] Among them, the point cloud encoder 105 is used to extract point cloud features, the image encoder 106 is used to extract image features, the feature fusion module 107 is used to fuse the point cloud features and the image features, and the point cloud decoder 108 is used to expand the point cloud.

[0045] Specifically, the point cloud completion device 101 inputs the image from the image acquisition device 102 and the point cloud from the point cloud acquisition device 103 into the trained point cloud completion model 104, so that the point cloud encoder 105 extracts features from the point cloud from the point cloud acquisition device 103 to obtain point cloud features. Similarly, the image encoder 106 extracts features from the image from the image acquisition device 102 to obtain image features. Then, the point cloud encoder 105 and the image encoder 106 can input the extracted features into the feature fusion module 107, so that the feature fusion module 107 fuses the point cloud features extracted by the point cloud encoder 105 and the image features extracted by the image encoder 106 to obtain the fused point cloud features, and inputs the fused point cloud features into the point cloud decoder 108. After that, the point cloud decoder 108 can expand the point cloud collected by the point cloud acquisition device 103 based on the fused point cloud features to obtain a point cloud with a greater density.

[0046] That is to say, the image encoder extracts features from the input image, the point cloud encoder extracts features from the input incomplete point cloud, the feature fusion module combines the two features, respectively utilizes the complementary advantages of the point cloud and the image to improve the performance and effect of the point cloud completion task, and by fusing information of different modalities, better restores the missing part of the point cloud. The point cloud decoder then predicts the sparse point cloud from the fused features and expands it on this basis to generate a dense and complete point cloud.

[0047] Optionally, the image encoder 106 can also be deployed in the image acquisition device 102. After the image acquisition device 102 acquires an image, it can extract features from the acquired image through the image encoder 106 to obtain image features, and send the extracted image features to the point cloud completion model 104.

[0048] Similarly, the point cloud encoder 105 can also be deployed in the point cloud acquisition device 103. After the point cloud acquisition device 103 acquires a point cloud, it can extract features from the acquired point cloud through the point cloud encoder 105 to obtain point cloud features, and send the extracted point cloud features to the point cloud completion model 104.

[0049] That is to say, the acquisition device (such as the image acquisition device 102, the point cloud acquisition device 103) can directly extract the feature data and upload the extracted feature data to the point cloud completion model 104.

[0050] It should be noted that the image acquisition device 102 in the embodiments of the present disclosure is not limited. For example, the image acquisition device 102 can be a driving recorder. For another example, the image acquisition device 102 can be a mobile phone with a video shooting function. For another example, the image acquisition device 102 can be a surveillance camera.

[0051] Similarly, the point cloud acquisition device 103 in the embodiments of the present disclosure is not limited. For example, the point cloud acquisition device 103 can be a laser scanner. For another example, the point cloud acquisition device 103 can be a laser scanner. For another example, the point cloud acquisition device 103 can be a vehicle-mounted millimeter-wave radar.

[0052] In the embodiments of the present disclosure, the point cloud completion device 101 can be a terminal, or the point cloud completion device 101 can be a server.

[0053] Among them, the terminal can be a device such as a mobile phone, a tablet computer, a desktop type, a laptop, a handheld computer, a notebook computer, an Ultra-mobile Personal Computer (UMPC), a netbook, etc. with transceiver functions. The embodiments of the present disclosure do not impose special restrictions on the specific form of the terminal. It can perform human-computer interaction with the user through one or more methods such as a keyboard, a touchpad, a touch screen, a remote control, voice interaction, or a handwriting device.

[0054] The server can be a single physical server, or it can also be a server cluster composed of multiple servers. Or, the server cluster can also be a distributed cluster. Or, the server can be a cloud server. The embodiments of the present disclosure do not limit the specific implementation manner of the server.

[0055] After introducing the application scenarios and implementation environments of the embodiments of the present disclosure, the point cloud completion method provided by the embodiments of the present disclosure will be described in detail below in combination with the above implementation environments.

[0056] The embodiments of the present disclosure provide a point cloud completion method, as Figure 2 shown, the point cloud completion method can include: S201 - S203.

[0057] S201. The point cloud completion device acquires initial image features and initial point cloud features.

[0058] Among them, the initial image features are the image features of the initial image of the area to be detected, and the initial point cloud features are the point cloud features of the initial point cloud of the area to be detected.

[0059] In an embodiment of the present disclosure, before the point cloud completion device obtains the initial image feature and the initial point cloud feature, the point cloud completion device can obtain the initial image and the initial point cloud by receiving the initial image from the image acquisition device and receiving the initial point cloud from the point cloud acquisition device. Then, the point cloud completion device can respectively perform feature extraction on the initial image and the initial point cloud to obtain the initial image feature of the initial image and the initial point cloud feature of the initial point cloud, and further obtain the initial image feature and the initial point cloud feature.

[0060] As another possible implementation, the point cloud completion device may include an image encoder. The point cloud completion device can receive the initial image from the image acquisition device and input the initial image into the image encoder to obtain the initial image feature.

[0061] Among them, the image encoder is a Residual Networks (ResNet) 34 structure, and the image encoder includes: a 3*3 convolutional kernel, a batch normalization layer, a transposed convolution layer, and there is a residual connection between levels in the image encoder. The number of feature channels in each level of the image encoder is greater than a preset number threshold.

[0062] It should be noted that the preset number threshold can be the number of feature channels in the levels of the existing ResNet34 structure. That is to say, the number of feature channels in each level of the image encoder is greater than the number of feature channels in the corresponding level of the existing ResNet34 structure.

[0063] In an embodiment of the present disclosure, the image encoder is an improved ResNet34 feature encoder. The image encoder removes the last activation function in the ResNet34 feature encoder and modifies the output of the last fully connected layer to 1024 nodes. At the same time, the image encoder can perform feature extraction at different levels and perform scale transformation.

[0064] That is to say, the improved image encoder can obtain a series of image features with different scales and constructs an image feature pyramid.

[0065] Exemplarily, the specific implementation of the image encoder is as follows: First, the image encoder uses a 3x3 convolutional kernel for feature extraction and is usually used in conjunction with a batch normalization layer. This helps to enhance the expressive ability of the features. Second, the image encoder introduces a residual connection to solve the problems of gradient disappearance and expressive ability in the training of deep networks. The residual block consists of three convolutional layers and contains a skip connection that adds the input directly to the output, enabling the network to learn the residual mapping. In addition, the image encoder also uses a transposed convolution layer for upsampling operations to restore the size of the feature map to the size of the original image. Finally, the image encoder increases the number of feature channels at different levels of the network to enhance the expressive ability of the feature encoder.

[0066] It can be understood that through the above embodiments, the image data collected by the acquired camera can be utilized, and the modified ResNet34 feature encoder can be used for feature extraction. Meanwhile, an image feature pyramid is constructed to capture image information at different scales. The implementation method of the image encoder, including the use of convolutional kernels, the introduction of residual connections, the upsampling operation of the deconvolution layer, and the increase in the number of feature channels, can enhance the expression ability and scale perception ability of the image encoder. These steps together contribute to the implementation of the subsequent point cloud completion task.

[0067] Similarly, the point cloud completion device may further include a point cloud encoder. The point cloud completion device can receive the initial point cloud from the point cloud acquisition device and input the initial point cloud into the point cloud encoder to obtain the initial point cloud features.

[0068] Among them, the point cloud encoder is a point cloud processing network (point networks, PointNet) structure. The point cloud encoder is used to select representative points from the point cloud, extract multi-scale geometric features for the representative points, and construct a point cloud feature pyramid (i.e., sub-point cloud features at different scales). The point cloud encoder includes: two fully connected layers, a max pooling layer, and a feature splicing module;

[0069] Among them, the feature splicing module is used to, when the feature dimension output by the max pooling layer is less than the preset dimension threshold, re-input the output of the max pooling layer into the two fully connected layers in sequence until the feature dimension output by the max pooling layer is greater than or equal to the preset dimension threshold, and use the output of the max pooling layer as the output result of the point cloud encoder.

[0070] Exemplarily, as Figure 3 shown, first, the point cloud encoder uses two fully connected layers with shared parameters to extract features from the N×3 point cloud data. These two fully connected layers independently operate on each point of the point cloud to extract the features of each point. This process generates an N×256 feature matrix. Next, the point cloud encoder applies the max pooling layer to process the feature matrix. The max pooling layer selects the maximum value in each column and generates a 1×256 global feature vector. Subsequently, the point cloud encoder splices the global feature vector with the previous feature matrix to form an N×512 feature matrix. Finally, the point cloud encoder performs the same feature extraction, global feature extraction, and feature splicing operations on the spliced feature matrix again. This step generates a 1×

[0071] 1024 point cloud global feature vector fx. Through the above steps, the feature extraction and integration of the N×3 point cloud data are performed to obtain a higher-dimensional point cloud global feature vector fx.

[0072] It can be understood that by using fully connected layers, max pooling layers, and feature concatenation operations that share two parameters, local and global features in the point cloud can be effectively extracted and fused. Such a feature representation has richer information and representational capabilities for subsequent tasks and analyses.

[0073] Optionally, for the feature extraction process of the initial image and the initial point cloud, reference can also be made to the introductions of image feature extraction and point cloud feature extraction in conventional technologies, which will not be elaborated here.

[0074] S202. The point cloud completion device fuses the initial image features and the initial point cloud features according to the mapping relationship between the points in the initial point cloud and the pixel points in the initial image to obtain target point cloud features.

[0075] As a possible implementation, the initial image features include sub-image features of multiple different scales, the initial point cloud features include first sub-point cloud features of multiple different scales, and the scales in the initial image features correspond one-to-one with the scales in the initial point cloud features. The point cloud completion device can convert the initial image features into first point cloud features according to the mapping relationship between the points in the initial point cloud and the pixel points in the initial image. The first point cloud features include second sub-point cloud features corresponding to the sub-image features of each scale, and the points in the first point cloud features correspond one-to-one with the points in the initial point cloud features. Then, the point cloud completion device can fuse the sub-point cloud features of the same scale among all the first sub-point cloud features and all the second sub-point cloud features to obtain target point cloud features. Among them, the target point cloud features include third sub-point cloud features of multiple different scales, and the third sub-point cloud features are composed of all the sub-point cloud features with the same scale as the third sub-point cloud features among multiple first sub-point cloud features and multiple second sub-point cloud features.

[0076] It should be noted that the scale in the embodiments of the present disclosure is the scaling ratio. That is to say, the sub-point cloud features of different scales are the point cloud features of the initial point cloud at different scaling ratios, and the sub-image features of different scales are the image features of the initial image at different scaling ratios.

[0077] It can be understood that by fusing the point cloud image features at different scales, richer fused point cloud features can be obtained. In this way, more valuable references can be provided for subsequent point cloud completion, and the accuracy of the point cloud completion result can be improved.

[0078] Optionally, for the process in which the point cloud completion device fuses the initial image features and the initial point cloud features according to the mapping relationship between the points in the initial point cloud and the pixel points in the initial image to obtain target point cloud features, reference can also be made to the introduction of feature fusion of point cloud images in conventional technologies, which will not be elaborated here.

[0079] As a possible implementation, the image acquisition device can be a camera, and the point cloud acquisition device can be a lidar. The initial image is a visual image of the area to be detected captured by the camera, and the initial point cloud is the lidar point cloud of the area to be detected collected by the lidar. After the point cloud completion device obtains the initial image and the initial point cloud, the point cloud completion device can obtain the internal and external parameters between the camera and the lidar, and determine the mapping relationship between the points in the initial point cloud and the pixel points in the initial image according to the internal and external parameters between the camera and the lidar.

[0080] The following combines specific examples to introduce the process of the point cloud completion device obtaining the initial image and the initial point cloud, as well as obtaining the internal and external parameters between the camera and the lidar. The specific process is as follows:

[0081] 1. Lidar Scanning and Data Acquisition:

[0082] (a) Install the lidar device at an appropriate position in the required scene and connect it to a computer (i.e., the point cloud completion device).

[0083] (b) Use the lidar device to scan the scene, emit laser beams, and measure the reflection time of the laser beams with the objects in the scene to obtain the raw point cloud data.

[0084] (c) The lidar device will automatically store the raw point cloud data in the storage device of the computer.

[0085] 2. Camera Image Acquisition:

[0086] (a) Install the camera device at an appropriate position in the scene collected by the lidar and connect it to the computer.

[0087] (b) Use the camera device to take an image of the scene, ensuring that the camera's field of view range matches the lidar's scanning range.

[0088] (c) The image data can be obtained and saved by directly connecting the camera to the computer.

[0089] 3. Obtaining the Calibration Parameters is as follows:

[0090] (1) Prepare the calibration board: Use a checkerboard calibration board with known dimensions and place it in the scene, ensuring that the calibration board is within the field of view of the lidar and the camera.

[0091] (2) Collect calibration data:

[0092] (a) Simultaneously use the lidar and the camera to collect the point cloud data and image data of the calibration board.

[0093] (b) Keep the lidar and camera stationary, and ensure that the calibration board is captured by the lidar and camera at different positions and poses.

[0094] (c) Obtain the three-dimensional coordinate information of the points on the calibration board through the lidar, and at the same time obtain the two-dimensional coordinate information of the calibration board in the image through the camera.

[0095] (3) Calculation of calibration parameters:

[0096] (a) Using the known dimensions on the calibration board and the corresponding three-dimensional point cloud data and two-dimensional image data, adopt a calibration algorithm to calculate the internal parameters and external parameters between the lidar and the camera. Among them, the internal parameters include the focal length of the camera, the coordinates of the principal point, and the distortion coefficients, etc., and the external parameters include the position and pose information of the camera in the world coordinate system.

[0097] That is to say, the point cloud completion device can obtain the original point cloud collected by the lidar and the image collected by the camera that are synchronized in time and space, as well as the calibration parameters of the lidar and the camera.

[0098] The following combines specific examples to introduce the process of the point cloud completion device determining the mapping relationship between the points in the initial point cloud and the pixel points in the initial image. The specific process is as follows:

[0099] The point cloud completion device represents the first sub-point cloud feature in the initial point cloud feature as where n is the number of points, c q is the dimension of the feature vector of each point, 3 is related to the three-dimensional coordinates of each point; the first three elements of each point in the feature are associated with its three-dimensional coordinates. In the point cloud feature, each point is usually represented by three values indicating its position in three-dimensional space, and these three values correspond to the x, y, and z coordinates of the point respectively. Therefore, when it is mentioned that "3 is related to the three-dimensional coordinates of each point", it means that the first three elements in the feature are related to the three-dimensional coordinate information of the point and are used to represent the position of the point in three-dimensional space. And the following cq elements are other feature vectors related to each point, which are used to represent other attributes or features of the point (such as color, intensity value, normal, curvature, etc. information). Therefore, each row in the first sub-point cloud feature Q can be regarded as a feature vector of a point, where the first three elements represent the three-dimensional coordinates of the point, and the following cq elements represent other features of the point.

[0100] The point cloud completion device records the first sub-image feature with the same scale as the first sub-point cloud feature Q in the initial image feature as where w and h are the width and height of the first sub-image feature respectively, c g is the dimension of the feature vector of each pixel in the first sub-image feature.

[0101] To complete the fusion of the first sub-image feature G and the first sub-point cloud feature Q, it is first necessary to determine the alignment relationship between each point and each pixel (i.e., the mapping relationship between the points in the initial point cloud and the pixel points in the initial image). By using the extrinsic parameter matrix between the pre-calibrated lidar and the camera and the camera projection matrix (i.e., the intrinsic and extrinsic parameters between the lidar and the camera), each point in the first sub-point cloud feature Q can be projected into the coordinate system of the first sub-image feature G, as shown in Equation (1):

[0102]

[0103] where, [x i y i z i 1] T is the three-dimensional homogeneous coordinate of the i-th point q i in Q; [x i y i 1] T is the projected homogeneous coordinate of the point q i projected onto the image feature map coordinate system; d i is the depth scale factor, which serves to ensure that the value of the last dimension of the projected homogeneous coordinate is 1; s is the scale factor of the image feature map G compared to the original image; K is a geometric transformation matrix that maps the three-dimensional point coordinates into the camera coordinate system; from the above Equation (1), the projected coordinates of the point q i can be obtained, thereby determining the aligned pixel of q i , denoted as g i , as shown in Equation (2):

[0104]

[0105] where, is the floor function, and g i does not represent the i-th pixel in the first sub-image feature G, but represents the aligned pixel of the i-th point q i in the first sub-image feature G.

[0106] As a possible implementation, the point cloud completion device may further include a feature fusion module. The point cloud completion device can input the initial point cloud feature output by the point cloud encoder and the initial image feature output by the image encoder into the feature fusion module, thereby realizing the fusion of the initial image feature and the initial point cloud feature to obtain the target point cloud feature.

[0107] Next, taking the feature fusion module as an example and combining specific examples, the process of fusing the initial image feature and the initial point cloud feature to obtain the target point cloud feature will be introduced. The specific process is as follows:

[0108] Feature fusion from image to point cloud is to convert the appearance texture features in the image features into the representation form of the point cloud to fuse with the original point cloud features. During the fusion process, taking q i as an example, considering that each point has a single aligned pixel, the feature fusion module regards the alignment operation as a conversion from pixel features to point features. In the top-level features of the feature pyramid (i.e., features at different scales), due to the low resolution of the image features, the actual area represented by each pixel is large, and usually, one pixel corresponds to multiple points. Inspired by the improvement of ROI-Align compared to ROI-Pooling, the present disclosure proposes to use bilinear interpolation to obtain the transformed features. The feature fusion module can restore the specific features of the projected points in the continuous texture feature space through bilinear interpolation. For each point in the point cloud, its high-dimensional transformed features are calculated to form the transformed features. Then, the feature fusion module uses a multi-layer perceptron (MLP) for further feature extraction and dimensionality compression to obtain the point cloud features Q' obtained from the image. Finally, the feature fusion module concatenates Q and Q', and outputs the fused global feature f (i.e., the target point cloud feature).

[0109] That is to say, in the top-level features of the feature pyramid, due to the low resolution of the image features, the actual area represented by each pixel is large, so one pixel may correspond to multiple points. To convert the image features into point cloud features, the method of bilinear interpolation is used to obtain the transformed features. In this way, the specific features of the projected points in the continuous texture feature space can be restored. For each point in the point cloud, its high-dimensional transformed features are calculated to form the transformed features. This means that for each point, the point features are obtained from the corresponding image features through the transformation operation. To extract useful features from the transformed features and compress them to a lower dimension to reduce the computational complexity and improve the expressive ability of the features, a multi-layer perceptron (MLP) is used for further feature extraction and dimensionality compression to obtain the point cloud features Q' obtained from the image. The original point cloud features Q are concatenated with the point cloud features Q' obtained from the image. In this way, the original point cloud features can be fused with the features extracted from the image, and the fused global feature f is output.

[0110] The concatenation scheme is as follows:

[0111] When concatenating the original point cloud features Q and the point cloud features Q' obtained from the image, it is usually performed on the channel dimension of the features. The shape of the original point cloud features Q is (n, d1), where n is the number of points and d1 is the dimension of the feature vector of each point. The shape of the point cloud features Q' obtained from the image is (n, d2), where d2 is the dimension of the image feature vector of each point.

[0112] When splicing, the feature fusion module can concatenate Q and Q' along the channel dimension of the features to form new features. The shape of the spliced features will be (n, d1 + d2), where the feature vector of each point will contain information about the original point cloud features and the point cloud features obtained from the image.

[0113] It can be understood that when splicing two features along the channel dimension, the point cloud features and the point cloud features obtained from the image have the same structure in the spatial dimension, and the number of their points is the same. Therefore, splicing along the channel dimension can preserve the spatial structure of the point cloud and can fuse the information of the two modal features at the same time.

[0114] It should be noted that after the point cloud completion device obtains the target point cloud features, the point cloud completion device can complete the supplementation of the initial point cloud by executing S203.

[0115] S203. The point cloud completion device generates a target point cloud according to the target point cloud features.

[0116] Among them, the density of the target point cloud is greater than the density of the initial point cloud.

[0117] As a possible implementation, the point cloud completion device can generate multiple neighborhood point sets according to the global features and single-point features in the target point cloud features, and generate a target point cloud according to the initial point cloud and the multiple neighborhood point sets. Among them, a point in the initial point cloud corresponds to a neighborhood point set, and the neighborhood point set includes multiple extended points adjacent to the corresponding point in the initial point cloud.

[0118] Optionally, the point cloud completion device may further include a point cloud decoder. The point cloud completion module can input the target point cloud features output by the feature fusion module into the point cloud decoder to complete the supplementation of the initial point cloud, and then obtain a target point cloud with a greater density.

[0119] The following takes the point cloud decoder as an example and combines specific examples to introduce the point cloud completion process. The specific process is as follows:

[0120] The point cloud decoder performs two-stage processing on the fused global features output by the feature fusion module. First, it directly predicts and generates a sparse point cloud through a multi-layer perceptron network; then, using the Folding mechanism, it expands the local points by a certain multiple according to the global features and the sparse point cloud to obtain a dense point cloud, which specifically includes the following sub-steps:

[0121] 1. The point cloud decoder directly predicts and generates a sparse point cloud through a multi-layer perceptron network. The point cloud decoder inputs the fused global features into the MLP network, and the network can learn the structure and geometric information of the point cloud and generate the corresponding sparse point cloud:

[0122] 2. The point cloud decoder uses the Folding mechanism to expand local points by a certain multiple to obtain a dense point cloud. For each sparse point, the point cloud decoder increases the number of points by copying its eigenvalue to its neighboring points. In this way, each point in the sparse point cloud is expanded into multiple points with similar features, thus generating a dense point cloud;

[0123] First, for the fused global feature f, the point cloud decoder uses a multi-layer perceptron network to directly predict and generate a sparse point cloud. This sparse point cloud contains the information expressed by the global feature f.

[0124] Next, the point cloud decoder uses the Folding mechanism to process the global feature and the sparse point cloud to expand the number of local points. For any point x in the sparse point cloud, the point cloud decoder generates an n×n two-dimensional grid, and the points in it are centered at the origin of coordinates and arranged in a checkerboard pattern. The N = n×n points in this two-dimensional grid actually represent the coordinates for offsetting point x. Their coordinate values are initialized as (x, y) values in the two-dimensional plane arranged in a checkerboard pattern, where i = 1, 2,..., N.

[0125] Then, the point cloud decoder concatenates this two-dimensional coordinate with the three-dimensional coordinate of point x and the global feature, and predicts the coordinate information of the expanded points through a multi-layer perceptron network with shared parameters. This network structure uses a point-wise shared multi-layer perceptron to learn the distribution information of local points using the global feature. By learning this distribution information, the point cloud decoder can expand a single point to generate the coordinates of the surrounding N points. Performing the same operation on all points in the sparse point cloud can expand the number of points in the sparse point cloud by N times to form a dense point cloud.

[0126] It can be understood that according to the mapping relationship between the pixel points in the image and the points in the point cloud, the image features and the point cloud features are fused, and based on the fused point cloud features, the point cloud is complemented to obtain a point cloud with a greater degree of density. In this way, through the multi-modal fusion between the point cloud and the image, the feature dimension of the point cloud features is increased, and then the feature information of the point cloud features is increased, providing more accurate and complete reference information for the point cloud complementation process and improving the accuracy of the complemented point cloud.

[0127] In some embodiments, after the point cloud completion device obtains the first point cloud feature, the point cloud completion device may acquire a target earth mover's distance (EMD) and a target uniformity index value. The target EMD is the EMD between the first point cloud feature and the initial point cloud feature, and the target uniformity index value is used to represent the global distribution non-uniformity and local distribution aggregation degree of the points in the first point cloud feature. Then, the point cloud completion device may determine whether to fuse the sub-point cloud features of the same scale among all the first sub-point cloud features and all the second sub-point cloud features by respectively comparing the target EMD and the target uniformity index value with thresholds.

[0128] It should be noted that EMD is an index commonly used to evaluate the similarity between point clouds. For two point clouds X1 and X2 with the same number of points, their EMD is defined as shown in Formula 3:

[0129]

[0130] where φ: X1→X2 is a bijective function. For a given point set, there is a unique optimal bijective function φ, and a (1+∈) approximation strategy is used. The solution found within a fixed time by this strategy has an approximation error of about 1% with respect to the optimal solution. EMD is only applicable to point clouds with the same number of points. Therefore, for point clouds with different numbers of points, if EMD needs to be used, the point set with more points needs to be randomly sampled to have the same number of points as the point set with fewer points.

[0131] In addition, an evaluation index for point cloud uniformity is introduced. For a point set X with N points (i.e., the first point cloud feature). First, use the farthest point sampling algorithm to select M seed points from it. After selection, let each seed point be the center of a sphere with a set radius r, and find the points within this radius to form a subset of M points, defined as S i , where i = 1, 2,..., M. Since the set r is small enough, it can be considered that S i is located within a small area of area πr 2 , while for the entire point set X, it is located within a unit sphere with a radius of 1 and an area of π. Therefore, for all point sets S i , the expected percentage p of points contained is p = r 2 , and the expected number of points n = Np. The global distribution non-uniformity of points is defined by the following Formula 4:

[0132]

[0133] The meaning of Formula 4 is that for each local point set in the global, the number of points should be approximately close to the expected value, which measures the non-uniformity of the point cloud distribution globally. An index is also needed to measure the local non-uniformity. For the local point set S i, find the distance from each point to its nearest point, and denote the nearest distance of the k-th point as d i,k If the points are evenly distributed, assuming that in the plane S i and the neighboring points are arranged in a hexagonal pattern, the expected distance can be expressed by Equation (5):

[0134]

[0135] Define the local distribution aggregation of points using the following Equation (6):

[0136]

[0137] Combine the above two metrics and accumulate the sum for each seed point, that is, measure the point cloud uniformity using the metric shown in Equation (7):

[0138]

[0139] As a possible implementation, if the point cloud completion device determines that the target EMD is less than the preset EMD threshold and the target uniformity metric value is greater than the preset uniformity metric threshold, the point cloud completion device can fuse the sub-point cloud features of the same scale in the first point cloud feature and the initial point cloud feature to obtain the target point cloud feature.

[0140] As another possible implementation, if the point cloud completion device determines that the target EMD is greater than or equal to the preset EMD threshold and / or the target uniformity metric value is less than or equal to the preset uniformity metric threshold, the point cloud completion device re-converts the initial image feature until the EMD corresponding to the point cloud feature after the conversion of the initial image feature is less than the preset EMD threshold and the corresponding uniformity metric value is greater than the preset uniformity metric threshold.

[0141] It can be understood that by comparing the thresholds of the similarity between point clouds, the global distribution non-uniformity, and the local distribution aggregation degree, the accuracy of the feature conversion result from the image to the point cloud can be determined, ensuring the accuracy of the features used in the point cloud image feature fusion, thereby improving the accuracy of the point cloud image feature fusion result and providing an accurate feature reference for subsequent point cloud completion.

[0142] In the embodiments of the present disclosure, the point cloud completion device is deployed with a point cloud completion model, and the point cloud completion model may include: a point cloud encoder, an image encoder, a feature fusion module, and a point cloud decoder. After obtaining the initial image and the initial point cloud, the point cloud completion device can input the initial image and the initial point cloud into the trained point cloud completion model, and based on the processing of the initial image and the initial point cloud by the point cloud encoder, image encoder, feature fusion module, and point cloud decoder in the point cloud completion model, obtain a target point cloud with a greater density.

[0143] It should be noted that for the process of the point cloud completion device training the point cloud completion model, the point cloud completion device can calculate the loss function of the point cloud completion model, calculate the gradient of the loss function with respect to the model parameters of the point cloud completion model using the backpropagation algorithm, and use the optimization algorithm to update the model parameters of the point cloud completion model. Repeat the above steps until all training rounds are completed, and then obtain the trained point cloud completion model.

[0144] As a possible implementation, in the case where the above-mentioned target EMD is greater than or equal to the preset EMD threshold and / or the target uniformity index value is less than or equal to the preset uniformity index threshold, the point cloud completion device can retrain the point cloud completion model until the trained point cloud completion model can ensure that the EMD corresponding to the point cloud features after the transformation of the initial image features is less than the preset EMD threshold and the corresponding uniformity index value is greater than the preset uniformity index threshold.

[0145] Next, the point cloud completion method provided by the embodiments of the present disclosure will be introduced with specific examples. As Figure 4 shown, the point cloud completion device acquires the original point cloud collected by the lidar, the original image collected by the camera, and the calibration parameters of the lidar and the camera, and inputs the original point cloud, the original image, and the calibration parameters (i.e., internal parameters and external parameters) into the three-dimensional point cloud completion model based on cross-modal information fusion (i.e., the point cloud completion model). After being processed by the image encoder, the point cloud encoder, and the feature fusion module in the point cloud completion model, the point cloud decoder can expand to complete the sparse point cloud into a dense point cloud.

[0146] Exemplarily, as Figure 5 shown, the point cloud completion model can receive the initial image 501 and the initial point cloud 502. Then, the point cloud completion model can input the initial image 501 into the image encoder to obtain the image feature pyramid 503. Similarly, the point cloud completion model can input the initial point cloud 502 into the point cloud encoder to obtain the point cloud feature pyramid 504. After that, the feature fusion module can perform feature fusion on the features of the same scale in the image feature pyramid 503 and the point cloud feature pyramid 504 to obtain the fused point cloud feature pyramid 505 (i.e., the target point cloud feature). Finally, the feature decoder processes the point cloud feature pyramid 505 to obtain the sparse point cloud 506, and further processes the sparse point cloud 506 to obtain the dense point cloud 507, thereby realizing the completion of the point cloud.

[0147] It can be understood that, in order to implement the above functions, the point cloud completion device includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the algorithm steps of each example described in the embodiments of the present disclosure, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0148] The embodiments of the present disclosure can perform functional module division on the point cloud completion device according to the above method embodiments. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one functional module. The above integrated module can be implemented in the form of hardware or software. It should be noted that the division of modules in the embodiments of the present disclosure is illustrative, only a logical function division, and there can be other division methods in actual implementation. The following takes the example of dividing each functional module corresponding to each function for illustration.

[0149] Figure 6 It is a schematic structural diagram of a point cloud completion device provided by an embodiment of the present disclosure. The point cloud completion device 600 can execute the Figure 2 point cloud completion method shown in the above method embodiment. As Figure 6 shown, the point cloud completion device 600 includes: an acquisition module 601 and a processing module 602.

[0150] The acquisition module 601 is used to acquire an initial image feature and an initial point cloud feature. The initial image feature is the image feature of the initial image of the area to be detected, and the initial point cloud feature is the point cloud feature of the initial point cloud of the area to be detected. The processing module 602 is used to fuse the initial image feature and the initial point cloud feature according to the mapping relationship between the points in the initial point cloud and the pixel points in the initial image to obtain a target point cloud feature. The processing module 602 is further used to generate a target point cloud according to the target point cloud feature, and the density of the target point cloud is greater than that of the initial point cloud.

[0151] Optionally, the initial image features include sub-image features of multiple different scales, the initial point cloud features include first sub-point cloud features of multiple different scales, and the scales in the initial image features correspond one-to-one with the scales in the initial point cloud features. The processing module 602 is specifically configured to convert the initial image features into first point cloud features according to the mapping relationship between the points in the initial point cloud and the pixel points in the initial image. The first point cloud features include second sub-point cloud features corresponding to the sub-image features of each scale, and the points in the first point cloud features correspond one-to-one with the points in the initial point cloud features. The processing module 602 is further configured to fuse the sub-point cloud features of the same scale among all the first sub-point cloud features and all the second sub-point cloud features to obtain target point cloud features, and the target point cloud features include third sub-point cloud features of multiple different scales.

[0152] Optionally, the obtaining module 601 is specifically configured to input the initial image into an image encoder to obtain initial image features. The image encoder is a ResNet34 structure of a residual neural network. The image encoder includes: a 3*3 convolutional kernel, a batch normalization layer, a transposed convolution layer, and there are residual connections between the levels in the image encoder. The number of feature channels of each level in the image encoder is greater than a preset number threshold.

[0153] Optionally, the obtaining module 601 is specifically configured to input the initial point cloud into a point cloud encoder to obtain initial point cloud features. The point cloud encoder is a PointNet structure of a point cloud processing network. The point cloud encoder includes: two fully connected layers, a max pooling layer, and a feature concatenation module. Among them, the feature concatenation module is configured to, when the feature dimension output by the max pooling layer is less than a preset dimension threshold, input the output of the max pooling layer into the two fully connected layers in sequence until the feature dimension output by the max pooling layer is greater than or equal to the preset dimension threshold, and use the output of the max pooling layer as the output result of the point cloud encoder.

[0154] Optionally, the obtaining module 601 is further configured to obtain a target Earth Mover's Distance (EMD) and a target uniformity index value. The target EMD is the EMD between the first point cloud features and the initial point cloud features. The target uniformity index value is used to indicate the global distribution non-uniformity and the local distribution aggregation degree of the points in the first point cloud features. The processing module 602 is specifically configured to, when the target EMD is less than a preset EMD threshold and the target uniformity index value is greater than a preset uniformity index threshold, fuse the sub-point cloud features of the same scale among all the first sub-point cloud features and all the second sub-point cloud features to obtain target point cloud features.

[0155] Optionally, the processing module 602 is specifically configured to generate a plurality of neighborhood point sets according to the global features and single-point features in the target point cloud features. One point in the initial point cloud corresponds to one neighborhood point set, and the neighborhood point set includes a plurality of extended points adjacent to the corresponding point in the initial point cloud. The processing module 602 is further configured to generate a target point cloud according to the initial point cloud and the plurality of neighborhood point sets.

[0156] Optionally, the obtaining module 601 is further configured to obtain an initial image and an initial point cloud. The initial image is a visual image of the area to be detected captured by a camera, and the initial point cloud is a lidar point cloud of the area to be detected collected by a lidar. The processing module 602 is further configured to determine the mapping relationship between the points in the initial point cloud and the pixel points in the initial image according to the internal parameters and external parameters between the camera and the lidar.

[0157] Optionally, the point cloud completion device 600 is applied to a point cloud completion model, and the point cloud completion model includes: a point cloud encoder, an image encoder, a feature fusion module, and a point cloud decoder. Among them, the point cloud encoder is used to extract point cloud features, the image encoder is used to extract image features, the feature fusion model is used to fuse the point cloud features and image features, and the point cloud decoder is used to expand the point cloud.

[0158] In the case of implementing the functions of the above integrated modules in the form of hardware, the embodiments of the present disclosure provide another possible network device structure of the point cloud completion device involved in the above embodiments. As Figure 7 shown, the point cloud completion device 700 includes: a processor 702, a bus 704. Optionally, the point cloud completion device 700 may further include a memory 701; optionally, the point cloud completion device 700 may further include a communication interface 703.

[0159] The processor 702 may be used to implement or execute various exemplary logical blocks, modules, and circuits described in conjunction with the embodiments of the present disclosure. The processor 702 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logical blocks, modules, and circuits described in conjunction with the embodiments of the present disclosure. The processor 702 may also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0160] The communication interface 703 is used to connect to other devices through a communication network. The communication network may be an Ethernet, a wireless access network, a wireless local area network (WLAN), etc.

[0161] The memory 701 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0162] As a possible implementation, the memory 701 can exist independently of the processor 702. The memory 701 can be connected to the processor 702 through the bus 704 and is used to store instructions or program code. When the processor 702 calls and executes the instructions or program code stored in the memory 701, the point cloud completion method provided by the embodiments of the present disclosure can be implemented.

[0163] In another possible implementation, the memory 701 can also be integrated with the processor 702.

[0164] The bus 704 can be an extended industry standard architecture (EISA) bus, etc. The bus 704 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 7 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0165] Some embodiments of the present disclosure provide a computer-readable storage medium (for example, a non-transitory computer-readable storage medium). Computer program instructions are stored in the computer-readable storage medium. When the computer program instructions run on a computer, the computer is caused to execute the point cloud completion method described in any one of the above embodiments.

[0166] Exemplarily, the above computer-readable storage medium may include, but is not limited to: magnetic storage devices (such as hard disks, floppy disks, or magnetic tapes, etc.), optical discs (such as Compact Discs (CDs), Digital Versatile Discs (DVDs), etc.), smart cards, and flash memory devices (such as Erasable Programmable Read-Only Memories (EPROMs), cards, sticks, or key drives, etc.). The various computer-readable storage media described in this disclosure may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data).

[0167] An embodiment of the present disclosure provides a computer program product containing instructions. When the computer program product runs on a computer, it causes the computer to execute the point cloud completion method described in any one of the above embodiments.

[0168] As described above, the above are only the specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present disclosure should be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A point cloud completion method, characterized in that, The method includes: Obtaining an initial image feature and an initial point cloud feature, where the initial image feature is an image feature of an initial image of a region to be detected, and the initial point cloud feature is a point cloud feature of an initial point cloud of the region to be detected; Fusing the initial image feature and the initial point cloud feature according to a mapping relationship between points in the initial point cloud and pixel points in the initial image to obtain a target point cloud feature; Generating a target point cloud according to the target point cloud feature, where the density of the target point cloud is greater than the density of the initial point cloud.

2. The method according to claim 1, wherein The initial image feature includes sub-image features of multiple different scales, the initial point cloud feature includes first sub-point cloud features of multiple different scales, and the scales in the initial image feature correspond one-to-one with the scales in the initial point cloud feature; The step of fusing the initial image feature and the initial point cloud feature according to a mapping relationship between points in the initial point cloud and pixel points in the initial image to obtain a target point cloud feature includes: Converting the initial image feature into a first point cloud feature according to a mapping relationship between points in the initial point cloud and pixel points in the initial image, where the first point cloud feature includes second sub-point cloud features corresponding to sub-image features of each scale, and the points in the first point cloud feature correspond one-to-one with the points in the initial point cloud feature; Fusing the first sub-point cloud features and the second sub-point cloud features of the same scale among all the first sub-point cloud features and all the second sub-point cloud features to obtain the target point cloud feature, where the target point cloud feature includes third sub-point cloud features of multiple different scales.

3. The method according to claim 1, characterized in that, Obtaining the initial image feature includes: Inputting the initial image into an image encoder to obtain the initial image feature, where the image encoder is a ResNet34 structure of a residual neural network, and the image encoder includes: a 3*3 convolutional kernel, a batch normalization layer, a deconvolution layer, and there are residual connections between levels in the image encoder, and the number of feature channels of each level in the image encoder is greater than a preset number threshold.

4. The method according to claim 1, characterized in that, Obtaining the initial point cloud feature includes: Inputting the initial point cloud into a point cloud encoder to obtain the initial point cloud feature, where the point cloud encoder is a PointNet structure of a point cloud processing network, and the point cloud encoder includes: two fully connected layers, a max pooling layer, and a feature splicing module; Wherein, the feature splicing module is used to, when the feature dimension output by the max pooling layer is less than a preset dimension threshold, input the output of the max pooling layer into the two fully connected layers in sequence until the feature dimension output by the max pooling layer is greater than or equal to the preset dimension threshold, and use the output of the max pooling layer as the output result of the point cloud encoder.

5. The method according to claim 2, characterized in that, After obtaining the first point cloud feature, the method further includes: Obtaining a target Earth Mover's Distance (EMD) and a target uniformity index value, where the target EMD is the EMD between the first point cloud feature and the initial point cloud feature, and the target uniformity index value is used to indicate the global distribution non-uniformity and local distribution aggregation degree of points in the first point cloud feature. Fusing the sub-point cloud features of the same scale among all the first sub-point cloud features and all the second sub-point cloud features to obtain the target point cloud features includes: When the target EMD is less than a preset EMD threshold and the target uniformity index value is greater than a preset uniformity index threshold, fusing the sub-point cloud features of the same scale among all the first sub-point cloud features and all the second sub-point cloud features to obtain the target point cloud features.

6. The method according to claim 1, characterized in that, Generating a target point cloud according to the target point cloud features includes: Generating a plurality of neighborhood point sets according to the global features and single-point features in the target point cloud features. One point in the initial point cloud corresponds to one neighborhood point set, and the neighborhood point set includes a plurality of extended points adjacent to the corresponding point in the initial point cloud; Generating the target point cloud according to the initial point cloud and the plurality of neighborhood point sets.

7. The method according to claim 1, wherein Before obtaining the initial image features and the initial point cloud features, the method further includes: Obtaining the initial image and the initial point cloud, where the initial image is a visual image of the area to be detected captured by a camera, and the initial point cloud is a laser point cloud of the area to be detected collected by a lidar; Determining the mapping relationship between the points in the initial point cloud and the pixel points in the initial image according to the internal parameters and external parameters between the camera and the lidar.

8. The method according to claim 1, wherein Applied to a point cloud completion model, the point cloud completion model includes: a point cloud encoder, an image encoder, a feature fusion module, and a point cloud decoder; Among them, the point cloud encoder is used to extract point cloud features, the image encoder is used to extract image features, the feature fusion model is used to fuse point cloud features and image features, and the point cloud decoder is used to expand the point cloud.

9. An electronic device, characterized in that, Includes: A memory and a processor; The memory and the processor are coupled; The memory is used to store instructions executable by the processor; When the processor executes the instructions, it executes the method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, A computer instruction is stored on the computer-readable storage medium. When the computer instruction runs on a computer, the computer is caused to execute the method according to any one of claims 1-8.