A Point Cloud Completion Method and Related Devices Based on Cross-Modal and Deep Inpainting
By introducing cross-modal and deep repair technologies into the point cloud completion method, the problem of poor point cloud completion accuracy in the existing technology is solved, and a more efficient and accurate point cloud completion effect is achieved.
Patent Information
- Application Number
- CN202510336974.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The existing point cloud completion methods have problems such as insufficient cross-modal information processing capabilities, insufficient geometric details understanding capabilities, and failure to make full use of multimodal depth information, resulting in poor accuracy of point cloud completion.
The point cloud completion method based on cross-modal and deep repair is adopted. By obtaining the target point cloud data and auxiliary data, the data is encoded and decoded using the encoder and decoder, the point cloud coordinate characteristics under multiple density scales are obtained, and the auxiliary data is deeply repaired. Finally, the comprehensive loss function optimization encoder and decoder are constructed to improve the accuracy of point cloud completion.
The information richness of point cloud encoding features and the accuracy of point cloud coordinate features are improved, the accuracy of point cloud completion is enhanced, and the performance of encoder and decoder is improved.
Smart Images

Figure CN119850886B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of point cloud completion, and particularly relates to a point cloud completion method and related devices based on cross-modal and depth repair. Background Art
[0002] With the rapid development and progress in fields such as autonomous driving and robot navigation, three-dimensional point cloud technology has been widely applied. Point cloud data, as an efficient three-dimensional space expression method, provides an important basis for environmental perception, object recognition, and three-dimensional reconstruction. However, in practical applications, due to hardware limitations of devices (such as limited lidar scanning resolution) and the complexity of the acquisition environment (such as occlusion, lighting conditions, etc.), the directly obtained point cloud data is often sparse, incomplete, and has a lot of noise. Such defects seriously affect the application of point clouds in high-precision tasks, such as high-precision map construction, three-dimensional object detection, instance segmentation, object surface detection, and fine three-dimensional reconstruction, making it difficult for these downstream tasks to achieve good results.
[0003] Currently, existing point cloud completion methods mainly rely on deep learning technology. These methods directly infer the missing geometric information from sparse point clouds to generate dense point clouds. However, there are still the following limitations in relying solely on sparse point clouds for completion:
[0004] (1) Insufficient ability to process cross-modal information: The local geometric information contained in sparse point clouds is limited, making it difficult to accurately infer the details of complex shapes. However, most existing point cloud fusion completion methods are fusions at the result level, or directly choose to complete multi-modal data. Such methods are relatively rough and difficult to effectively combine point clouds and multi-modal information for collaborative completion.
[0005] (2) Insufficient ability to understand geometric details: How to learn local and global information is a very crucial task in point cloud completion. For existing methods, this is still a great challenge. For example, when facing situations where the shape structure of an object is complex and the data has extreme density distributions, existing point cloud completion methods are difficult to efficiently extract useful features.
[0006] (3) Insufficient utilization of multi-modal information: Most completion methods do not fully explore the depth information of multi-modalities and fail to effectively use depth information to assist point cloud completion.
[0007] In summary, it can be seen that the current point cloud completion methods have the problem of poor accuracy in point cloud completion. Summary of the Invention
[0008] The present application provides a point cloud completion method and related devices based on cross-modal and depth repair, which can solve the problem of poor accuracy in point cloud completion.
[0009] In a first aspect, an embodiment of the present application provides a point cloud completion method based on cross-modal and depth repair. The point cloud completion method includes:
[0010] Obtain target point cloud data and auxiliary data corresponding to the target point cloud data;
[0011] Use an encoder to encode the target point cloud data and the auxiliary data to obtain point cloud encoding features. The point cloud encoding features are used to describe the local structure information and global skeleton information of the target point cloud data;
[0012] Use a decoder to decode the point cloud encoding features to obtain point cloud coordinate features at multiple density scales, and perform depth repair on each point cloud coordinate feature according to the auxiliary data to obtain a completed point cloud at each density scale. The point cloud coordinate features are used to describe the coordinates of multiple data points at the corresponding density scale;
[0013] Construct a comprehensive loss function based on all the completed point clouds, and use the comprehensive loss function to optimize the encoder and the decoder to obtain an optimized encoder and an optimized decoder. The comprehensive loss function is used to describe the accuracy of all the completed point clouds;
[0014] Use the optimized encoder and the optimized decoder to complete the point cloud data to be completed to obtain a final completed point cloud.
[0015] Optionally, the encoder includes a multi-layer feature extraction model and a multi-layer feature fusion model;
[0016] The input data of the first input end of the multi-layer feature extraction model is the auxiliary data, the input data of the second input end of the multi-layer feature extraction model is the target point cloud data, and the output end of the multi-layer feature extraction model is connected to the first input end of the multi-layer feature fusion model;
[0017] The input data of the second input end of the multi-layer feature fusion model is the target point cloud data, and the output data of the output end of the multi-layer feature fusion model is the point cloud encoding features.
[0018] Optionally, the multi-layer feature extraction model includes: a first residual neural module, a second residual neural module, a third residual neural module, a first upsampling module, a second upsampling module, a first enhanced continuous convolution module, a second enhanced continuous convolution module, and a third enhanced continuous convolution module;
[0019] The input end of the first residual neural module is the first input end of the multi-layer feature extraction model, the input end of the first enhanced continuous convolution module is the second input end of the multi-layer feature extraction model, and the output ends of the first enhanced continuous convolution module, the second enhanced continuous convolution module, and the third enhanced continuous convolution module are all the output ends of the multi-layer feature extraction model;
[0020] The output end of the first residual neural module is respectively connected to the input end of the second residual neural module and the input end of the first enhanced continuous convolution module. The output end of the second residual neural module is respectively connected to the input end of the third residual neural module and the input end of the first upsampling module. The output end of the first upsampling module is connected to the input end of the second enhanced continuous convolution module. The output end of the third residual neural module is connected to the input end of the second upsampling module. The output end of the second upsampling module is connected to the input end of the third enhanced continuous convolution module. The output end of the first enhanced continuous convolution module is connected to the input end of the second enhanced continuous convolution module. The output end of the second enhanced continuous convolution module is connected to the input end of the third enhanced continuous convolution module.
[0021] Optionally, the multi-layer feature fusion model includes: a first multi-core edge convolution module, a second multi-core edge convolution module, a third multi-core edge convolution module, a fourth multi-core edge convolution module, a first cascading module, a second cascading module, a third cascading module, a first global structure perception module, a second global structure perception module, a clustering module, a max pooling module, a first dilation module, and a second dilation module;
[0022] The input end of the first multi-core edge convolution module is the second input end of the multi-layer feature fusion model. The first input end of the first cascading module is the first input end of the multi-layer feature fusion model. The output end of the third cascading module is the output end of the multi-layer feature fusion model;
[0023] The output end of the first multi-core edge convolution module is respectively connected to the input end of the second multi-core edge convolution module and the input end of the second cascading module. The output end of the second multi-core edge convolution module is respectively connected to the input end of the third multi-core edge convolution module and the input end of the second cascading module. The output end of the third multi-core edge convolution module is respectively connected to the second input end of the first cascading module and the input end of the second cascading module. The output end of the first cascading module is respectively connected to the input end of the fourth multi-core edge convolution module, the input end of the first global structure perception module, and the input end of the clustering module. The output end of the fourth multi-core edge convolution module is connected to the input end of the third cascading module. The output end of the clustering module is connected to the input end of the first global structure perception module. The output end of the first global structure perception module is connected to the input end of the second global structure perception module. The output end of the second global structure perception module is connected to the input end of the first dilation module. The output end of the first dilation module is connected to the input end of the third cascading module. The output end of the second cascading module is connected to the input end of the max pooling module. The output end of the max pooling module is connected to the input end of the second dilation module. The output end of the second dilation module is connected to the input end of the third cascading module.
[0024] Optionally, both the first global structure perception module and the second global structure perception module are global structure perception modules;
[0025] The global structure perception module includes a convolutional layer, a first non-linear transformation layer, a second non-linear transformation layer, a third non-linear transformation layer, a multi-head interactive attention layer, a first residual normalization layer, a feature adaptive selection layer, and a second residual normalization layer;
[0026] The input end of the convolutional layer is the input end of the global structure perception module, and the output end of the second residual normalization layer is the output end of the global structure perception module;
[0027] The output end of the convolutional layer is respectively connected to the input ends of the first non-linear transformation layer, the second non-linear transformation layer, the third non-linear transformation layer, and the first residual normalization layer. The input ends of the multi-head interactive attention layer are respectively connected to the output ends of the first non-linear transformation layer, the second non-linear transformation layer, and the third non-linear transformation layer. The output end of the multi-head interactive attention layer is connected to the input end of the first residual normalization layer. The output end of the first residual normalization layer is respectively connected to the input ends of the second residual normalization layer and the feature adaptive selection layer. The output end of the feature adaptive selection layer is connected to the input end of the second residual normalization layer.
[0028] Optionally, the decoder includes a multi-layer feature decoding model, a registration model, and a depth completion model;
[0029] The output ends of the multi-layer feature decoding model and the registration model are both connected to the input end of the depth completion model. The input data of the input end of the multi-layer feature decoding model is the point cloud encoded feature. The output data of the output end of the multi-layer feature decoding model is the point cloud coordinate features under multiple density sizes. The input data of the input end of the registration model is the auxiliary data. The output data of the output end of the depth completion model is the completed point clouds under multiple density scales.
[0030] Optionally, the multi-layer feature decoding model includes a first mapping module, a second mapping module, a third mapping module, a first addition module, a second addition module, a third addition module, and a multi-layer perception module;
[0031] The input ends of the first mapping module, the second mapping module, the third mapping module, the first addition module, the second addition module, and the third addition module are all the input ends of the multi-layer feature decoding model. The output end of the multi-layer perception module is the output end of the multi-layer feature decoding model;
[0032] The output end of the first mapping module is connected to the input end of the first addition module. The output end of the second mapping module is connected to the input end of the second addition module. The output end of the third mapping module is connected to the input end of the third addition module.
[0033] Optionally, the comprehensive loss function is:
[0034]
[0035] Among them, represents the value of the comprehensive loss function, and represent hyperparameters, , represents the multi-scale completion loss, represents the overall loss:
[0036]
[0037]
[0038] Among them, represents the weighting parameter, represents the discrimination probability of the discriminator for the generated completed point cloud, represents the completed point cloud, represents the discrimination probability of the discriminator for the real point cloud, represents the real point cloud corresponding to the completed point cloud, represents the completion loss at the first density scale, represents the completion loss at the second density scale, represents the completion loss at the third density scale:
[0039]
[0040] Among them, , represents the th sampling point set of the completed point cloud at the density scale, represents the number of sampling points in, represents the th sampling point set of the real point cloud at the density scale, represents the number of sampling points in.
[0041] In a second aspect, an embodiment of the present application provides a point cloud completion device based on cross-modal and depth repair, including:
[0042] An acquisition module, configured to acquire target point cloud data and auxiliary data corresponding to the target point cloud data;
[0043] An encoding module, configured to encode the target point cloud data and the auxiliary data by using an encoder to obtain point cloud encoding features; the point cloud encoding features are used to describe the local structure information and the global skeleton information of the target point cloud data;
[0044] A decoding module, configured to use a decoder to decode the point cloud encoded features to obtain point cloud coordinate features at multiple density scales, and perform depth repair on each point cloud coordinate feature according to the auxiliary data to obtain a completed point cloud at each density scale; the point cloud coordinate features are used to describe the coordinates of multiple data points at the corresponding density scale;
[0045] An optimization module, configured to construct a comprehensive loss function according to all the completed point clouds, and use the comprehensive loss function to optimize the encoder and the decoder to obtain an optimized encoder and an optimized decoder; the comprehensive loss function is used to describe the accuracy of all the completed point clouds;
[0046] A completion module, configured to use the optimized encoder and the optimized decoder to complete the point cloud data to be completed to obtain a final completed point cloud.
[0047] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned point cloud completion method based on cross-modal and depth repair is implemented.
[0048] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned point cloud completion method based on cross-modal and depth repair is implemented.
[0049] The above solution of the present application has the following beneficial effects:
[0050] In the embodiment of the present application, by obtaining the target point cloud data and the auxiliary data corresponding to the target point cloud data, then using the encoder to encode the target point cloud data and the auxiliary data to obtain point cloud encoded features, and then using the decoder to decode the point cloud encoded features to obtain point cloud coordinate features at multiple density scales, and performing depth repair on each point cloud coordinate feature according to the auxiliary data to obtain a completed point cloud at each density scale, then constructing a comprehensive loss function according to all the completed point clouds, and using the comprehensive loss function to optimize the encoder and the decoder to obtain an optimized encoder and an optimized decoder, and finally using the optimized encoder and the optimized decoder to complete the point cloud data to be completed to obtain a final completed point cloud. Among them, obtaining the point cloud encoded features based on the target point cloud data and the auxiliary data can improve the information richness in the point cloud encoded features, thereby improving the accuracy of the point cloud coordinate features obtained by decoding the point cloud encoded features. The accuracy of the completed point cloud obtained according to the accurate point cloud coordinate features is improved. Constructing a comprehensive loss function to optimize the encoder and the decoder can improve the performance of the encoder and the decoder, and further improve the accuracy of point cloud completion.
[0051] Other beneficial effects of the present application will be described in detail in the following specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0053] Figure 1 It is a flowchart of a point cloud completion method based on cross-modal and depth inpainting provided by an embodiment of the present application;
[0054] Figure 2 It is a schematic structural diagram of an encoder provided by an embodiment of the present application;
[0055] Figure 3 It is a schematic structural diagram of a global structure perception module provided by an embodiment of the present application;
[0056] Figure 4 It is a schematic structural diagram of a decoder provided by an embodiment of the present application;
[0057] Figure 5 It is a schematic structural diagram of a discriminator provided by an embodiment of the present application;
[0058] Figure 6 It is a schematic diagram of the specific process of a point cloud completion method based on cross-modal and depth inpainting provided by an embodiment of the present application;
[0059] Figure 7 It is a schematic structural diagram of a point cloud completion device based on cross-modal and depth inpainting provided by an embodiment of the present application;
[0060] Figure 8 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. SPECIFIC EMBODIMENTS
[0061] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented in order to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0062] It should be understood that, as used in the specification of this application and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups.
[0063] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0064] As used in the specification of this application and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.
[0065] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for differential description and cannot be understood as indicating or implying relative importance.
[0066] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0067] In view of the problem of poor accuracy in existing point cloud completion, the embodiments of the present application provide a point cloud completion method based on cross-modal and depth repair. The point cloud completion method obtains target point cloud data and auxiliary data corresponding to the target point cloud data, then uses an encoder to encode the target point cloud data and the auxiliary data to obtain point cloud encoding features, and then uses a decoder to decode the point cloud encoding features to obtain point cloud coordinate features at multiple density scales, and performs depth repair on each point cloud coordinate feature according to the auxiliary data to obtain a completed point cloud at each density scale. Then, a comprehensive loss function is constructed based on all the completed point clouds, and the encoder and the decoder are optimized using the comprehensive loss function to obtain an optimized encoder and an optimized decoder. Finally, the optimized encoder and the optimized decoder are used to complete the point cloud data to be completed, and a final completed point cloud is obtained. Among them, obtaining point cloud encoding features based on the target point cloud data and the auxiliary data can improve the information richness in the point cloud encoding features, and further improve the accuracy of the point cloud coordinate features obtained by decoding the point cloud encoding features. The accuracy of the completed point cloud obtained based on the accurate point cloud coordinate features is improved. Constructing a comprehensive loss function to optimize the encoder and the decoder can improve the performance of the encoder and the decoder, and further improve the accuracy of point cloud completion.
[0068] Next, an exemplary description will be given of the point cloud completion method based on cross-modal and depth repair provided by the present application.
[0069] As Figure 1 shown, the point cloud completion method based on cross-modal and depth repair provided by the present application includes the following steps:
[0070] Step 11, obtain target point cloud data and auxiliary data corresponding to the target point cloud data.
[0071] The above-mentioned target point cloud data can be the surrounding environment of an autonomous vehicle, the point cloud data of an object, and the above-mentioned auxiliary data is used to provide relevant information corresponding to the target point cloud data (such as color, texture information, etc.). For example, the auxiliary data is an image of the surrounding environment of an autonomous vehicle, 3D modeling, etc.
[0072] In some embodiments of the present application, devices such as a laser scanner and a depth camera can be used to obtain the target point cloud data and the corresponding auxiliary data.
[0073] Exemplarily, after obtaining the target point cloud data, data preprocessing is an essential step. The main purpose is to clean the data, including removing noise points and outliers, performing coordinate transformation, unifying the scale, etc., so as to obtain reliable, enhanced, and standardized point cloud data. The steps of preprocessing include:
[0074] 1. Noise reduction processing
[0075] The purpose of noise reduction is to remove the noise points in the point cloud caused by sensor errors, environmental interference or the acquisition process, so as to improve the quality of the point cloud and the effect of subsequent algorithms. Typical algorithms include statistical filtering, radius filtering, smoothing filtering, etc. In this paper, statistical filtering is adopted to denoise the point cloud data.
[0076] (1) Finding the neighborhood: For each point , determine its spherical neighborhood within a given radius , and is the corresponding number of neighborhood points.
[0077]
[0078]
[0079] Among them, is the set of data points in the target point cloud data, and the value of the radius is estimated by calculating the average distance between points in the point cloud.
[0080]
[0081] Among them, is usually a constant between 1 and 2.
[0082] (2) Calculating the mean neighborhood distance: Calculate the local neighborhood average distance of each point.
[0083] ;
[0084] (3) Global statistical analysis: Calculate the global mean and standard deviation of the neighborhood distance means of all points:
[0085] ;
[0086]
[0087] (4) Removing outliers: Set the outlier threshold . If the neighborhood distance mean of a certain point meets the following conditions, then remove it. The finally retained point cloud is denoted as .
[0088] ( Generally take 1 or 2)
[0089] 2. Normalization processing
[0090] The main purpose of normalization is to standardize the point cloud data into a unified coordinate system and scale, which helps the model converge and stabilize quickly. This module includes two steps: centroid normalization and scale normalization. The specific process is as follows:
[0091] (1)Centroid normalization
[0092] Calculate the centroid: ;
[0093] Centroid translation: ;
[0094] (2)Scale normalization
[0095] Calculate the maximum distance to the centroid: ;
[0096] Scaling operation: .
[0097] It is worth mentioning that the preprocessed high-quality data is an important prerequisite and guarantee for subsequent effective analysis and processing in the model to obtain high-precision output.
[0098] Step 12: Use the encoder to encode the target point cloud data and auxiliary data to obtain point cloud encoding features.
[0099] The above point cloud encoding features are used to describe the local structure information and global skeleton information of the target point cloud data.
[0100] As Figure 2 shown, the encoder includes a multi-layer feature extraction model and a multi-layer feature fusion model;
[0101] The input data of the first input end of the multi-layer feature extraction model is auxiliary data, the input data of the second input end of the multi-layer feature extraction model is target point cloud data, and the output end of the multi-layer feature extraction model is connected to the first input end of the multi-layer feature fusion model;
[0102] The input data of the second input end of the multi-layer feature fusion model is target point cloud data, and the output data of the output end of the multi-layer feature fusion model is point cloud encoding features.
[0103] The multi-layer feature extraction model includes: the first residual neural module, the second residual neural module, the third residual neural module, the first upsampling module, the second upsampling module, the first enhanced continuous convolution module, the second enhanced continuous convolution module, and the third enhanced continuous convolution module;
[0104] The input end of the first residual neural module is the first input end of the multi-layer feature extraction model. The input end of the first enhanced continuous convolution module is the second input end of the multi-layer feature extraction model. The output ends of the first enhanced continuous convolution module, the second enhanced continuous convolution module, and the third enhanced continuous convolution module are all the output ends of the multi-layer feature extraction model;
[0105] The output end of the first residual neural module is respectively connected to the input end of the second residual neural module and the input end of the first enhanced continuous convolution module. The output end of the second residual neural module is respectively connected to the input end of the third residual neural module and the input end of the first upsampling module. The output end of the first upsampling module is connected to the input end of the second enhanced continuous convolution module. The output end of the third residual neural module is connected to the input end of the second upsampling module. The output end of the second upsampling module is connected to the input end of the third enhanced continuous convolution module. The output end of the first enhanced continuous convolution module is connected to the input end of the second enhanced continuous convolution module. The output end of the second enhanced continuous convolution module is connected to the input end of the third enhanced continuous convolution module.
[0106] The multi-layer feature fusion model includes: a first multi-core edge convolution module, a second multi-core edge convolution module, a third multi-core edge convolution module, a fourth multi-core edge convolution module, a first cascading module, a second cascading module, a third cascading module, a first global structure perception module, a second global structure perception module, a clustering module, a max pooling module, a first dilation module, and a second dilation module;
[0107] The input end of the first multi-core edge convolution module is the second input end of the multi-layer feature fusion model. The first input end of the first cascading module is the first input end of the multi-layer feature fusion model. The output end of the third cascading module is the output end of the multi-layer feature fusion model;
[0108] The output ends of the first multi-core edge convolution module are respectively connected to the input ends of the second multi-core edge convolution module and the input end of the second cascading module. The output ends of the second multi-core edge convolution module are respectively connected to the input ends of the third multi-core edge convolution module and the input end of the second cascading module. The output ends of the third multi-core edge convolution module are respectively connected to the second input end of the first cascading module and the input end of the second cascading module. The output ends of the first cascading module are respectively connected to the input ends of the fourth multi-core edge convolution module, the input end of the first global structure perception module, and the input end of the clustering module. The output end of the fourth multi-core edge convolution module is connected to the input end of the third cascading module. The output end of the clustering module is connected to the input end of the first global structure perception module. The output end of the first global structure perception module is connected to the input end of the second global structure perception module. The output end of the second global structure perception module is connected to the input end of the first dilation module. The output end of the first dilation module is connected to the input end of the third cascading module. The output end of the second cascading module is connected to the input end of the max pooling module. The output end of the max pooling module is connected to the input end of the second dilation module. The output end of the second dilation module is connected to the input end of the third cascading module.
[0109] It should be noted that the branch composed of the first global structure perception module, the second global structure perception module, the clustering module, and the first dilation module is used to extract global skeleton information. The output data at the output end of the first dilation module is data describing the global skeleton information. The fourth multi-core edge convolution module is used to extract local structure information, and the output data at its output end is data describing the local structure information.
[0110] The above-mentioned first global structure perception module and second global structure perception module are both global structure perception modules;
[0111] As Figure 3 shown, the global structure perception module includes a convolutional layer, a first non-linear transformation layer, a second non-linear transformation layer, a third non-linear transformation layer, a multi-head interactive attention layer, a first residual normalization layer, a feature adaptive selection layer, and a second residual normalization layer;
[0112] The input end of the convolutional layer is the input end of the global structure perception module, and the output end of the second residual normalization layer is the output end of the global structure perception module;
[0113] The output ends of the convolutional layers are respectively connected to the input ends of the first non-linear transformation layer, the second non-linear transformation layer, the third non-linear transformation layer, and the input end of the first residual normalization layer. The input ends of the multi-head interactive attention layer are respectively connected to the output ends of the first non-linear transformation layer, the second non-linear transformation layer, and the third non-linear transformation layer. The output end of the multi-head interactive attention layer is connected to the input end of the first residual normalization layer. The output end of the first residual normalization layer is respectively connected to the input end of the second residual normalization layer and the input end of the feature adaptive selection layer. The output end of the feature adaptive selection layer is connected to the input end of the second residual normalization layer.
[0114] It should be noted that the above first residual neural module, second residual neural module, and third residual neural module can all perform operations of the residual neural network for feature extraction of the input data. The first upsampling module and the second upsampling module are used for upsampling the input data, and the upsampling scale of the second upsampling module is larger than that of the first upsampling module. The first enhanced continuous convolution module, the second enhanced continuous convolution module, and the third enhanced continuous convolution module are all used to perform operations of enhanced continuous convolution for feature enhancement of the input data. The first multi-core edge convolution module, the second multi-core edge convolution module, the third multi-core edge convolution module, and the fourth multi-core edge convolution module are all multi-core edge convolutions for capturing the geometry and feature distribution of the input data. The first cascading module, the second cascading module, and the third cascading module are all used to cascade the corresponding input data. The first global structure perception module and the second global structure perception module are both used for global feature extraction of the input data. The clustering module is used for clustering the input data. The max pooling module is used for max pooling of the input data. The first dilation module and the second dilation module are both used for dimensional dilation of the input data.
[0115] The above convolutional layer is used for performing convolutional operations on the input data. The first non-linear transformation layer, the second non-linear transformation layer, and the third non-linear transformation layer are respectively used for calculating the key vector, query vector, and value vector in the multi-head interactive attention mechanism. The multi-head interactive attention layer is used for performing operations of the multi-head attention mechanism and interaction on the input data. The first residual normalization layer and the second residual normalization layer are used for performing addition and normalization operations on the input data. The feature adaptive selection layer is used for feature selection of the input data.
[0116] Exemplarily, the operation process of the enhanced continuous convolution is as follows:
[0117] For each point in the point cloud with enhanced continuous convolution, called the query point, the k nearest points are found using the K-Nearest Neighbors (KNN) algorithm to form the neighborhood of the point, and then the neighborhood points are projected onto the image plane through the internal and external parameter matrices of the camera. To enhance the expressive power of the projected point features, a feature aggregation strategy based on a 3×3 grid is designed. Specifically, through the mean operation, the pixel features within the 3×3 grid area around the projected point are aggregated into a feature vector. If the area is less than 3×3 in size, it is filled with 0. After this operation, the image feature set of the query point is obtained. .
[0118]
[0119] where is the pixel coordinate of point , are the height and width of the image.
[0120] Next, the image features of the k neighborhood points are fused using a Multi-Layer Perceptron (MLP). The input of the MLP consists of three parts, namely: image features, geometric offsets between the query point and the neighborhood points, and semantic offsets between the query point and the neighborhood points.
[0121] ,
[0122] where is the neighborhood of i (including i), is the image feature of neighborhood point j, is the geometric offset between point j and point i, is the semantic offset between the two, and concat(·) is the concatenation of multiple vectors.
[0123] Finally, the features output by the MLP are aggregated by summation to obtain the enhanced features output by the enhanced continuous convolution.
[0124] The operation expression of the above multi-core edge convolution is:
[0125]
[0126]
[0127]
[0128]
[0129] where, represents the output data of the multi-core edge convolution, represents the concatenation operation. Denote the -th output data, , denote the number of MLPs in the multi-core edge convolution, denote the density of points in the input data, and denote the weight coefficient, denote the feature vector of point i, denote the feature vector of point j, denote the distance offset between two points, denote point 's neighborhood, denote the number of points in the neighborhood, denote point and point 's number of common neighborhood points, denote point 's number of neighboring points.
[0130] The operation process in the above global structure perception module is as follows. The global structure perception module GCP receives the input data X, where X is a matrix of size N×C, corresponding to the feature vectors of N points. First, a 1×1 convolution is used to reduce the dimension of the points to obtain , so as to eliminate redundant information and reduce the computational complexity. In order to capture complex relationships, an MLP is used for non-linear transformation instead of linear projection to obtain the Q, K, and V matrices. Then, multi-head interactive attention in residual form is used to learn features.
[0131]
[0132]
[0133] Among them, denote the normalized data, use vector attention to adjust different feature channels, and adopt self-attention to capture the interaction between heads, 's operation is:
[0134]
[0135]
[0136] Among them is the standard self-attention operation. Through this operation, the model can learn the relationship between attention heads and improve the expression ability of features, denote the output of the i-th point after the vector attention mechanism, denotes the query vector corresponding to the $i$-th point under the current attention head, denotes the key vector corresponding to the $j$-th point under the current attention head, denotes the value vector corresponding to the $j$-th point under the current attention head, denotes the output of the first attention head, denotes the output of the second attention head, denotes the output of the eighth attention head, denotes the output data after the multi-head self-attention operation.
[0137] To enhance the fitting and selection capabilities of the model, a Feature Adaptive Selection layer FAS is added to adaptively adjust the contributions of channels, thereby improving the overall performance. Finally, the output of GCP can be expressed as:
[0138]
[0139]
[0140] where denotes global average pooling, is a 1×1 convolution, is the sigmoid function, denotes the data after being processed by the result FAS layer.
[0141] Step 13: Use the decoder to decode the point cloud encoded features to obtain point cloud coordinate features at multiple density scales, and perform depth repair on each point cloud coordinate feature according to the auxiliary data to obtain the completed point cloud at each density scale.
[0142] The above point cloud coordinate features are used to describe the coordinates of multiple data points at the corresponding density scale. The data points are the points in the point cloud data. In the related technology, point cloud completion refers to adding points to the point cloud at a certain density scale (such as having 10 points per unit volume) through a certain algorithm or method to complete the missing part. In this step, point cloud completion at multiple density scales is performed respectively to meet the completion requirements of point clouds at different density scales in actual tasks.
[0143] Such as Figure 4 shown, the above decoder includes a multi-layer feature decoding model, a registration model, and a depth completion model;
[0144] The output end of the multi-layer feature decoding model and the output end of the registration model are both connected to the input end of the depth completion model. The input data of the input end of the multi-layer feature decoding model is the point cloud encoding feature, the output data of the output end of the multi-layer feature decoding model is the point cloud coordinate features under multiple density dimensions, the input data of the input end of the registration model is the auxiliary data, and the output data of the output end of the depth completion model is the completed point clouds under multiple density scales.
[0145] The multi-layer feature decoding model includes a first mapping module, a second mapping module, a third mapping module, a first addition module, a second addition module, a third addition module, and a multi-layer perception module;
[0146] The input end of the first mapping module, the input end of the second mapping module, the input end of the third mapping module, the input end of the first addition module, the input end of the second addition module, and the input end of the third addition module are all the input end of the multi-layer feature decoding model, and the output end of the multi-layer perception module is the output end of the multi-layer feature decoding model;
[0147] The output end of the first mapping module is connected to the input end of the first addition module, the output end of the second mapping module is connected to the input end of the second addition module, and the output end of the third mapping module is connected to the input end of the third addition module.
[0148] It should be noted that the registration model is used to generate and register point clouds for the auxiliary data. The registration model uses the Iterative Closest Point (ICP) algorithm. The depth completion model is used to implement the step of depth patching for each point cloud coordinate feature according to the auxiliary data to obtain the completed point clouds under each density scale. The multi-layer feature decoding model is used to implement the step of decoding the point cloud encoding feature to obtain the point cloud coordinate features under multiple density scales. The above first mapping module, second mapping module, and third mapping module are all used to map the input data to the corresponding dimensions, and the dimensions corresponding to the first mapping module, second mapping module, and third mapping module are all different. The multiple dimensions correspond one-to-one with the multiple density scales. The first addition module, second addition module, and third addition module are used to perform addition operations on the input data, and the multi-layer perception module is used to perform operations of a multi-layer perceptron on the input data.
[0149] Exemplarily, the operation process in the above multi-layer feature decoding model is as follows:
[0150] First, perform feature expansion on three branches. For branch 1 (the first mapping module and the first addition module), the input feature After being processed by the MLP, it outputs a feature of the same size, and then adds it to the original feature Point-by-point addition is performed on the point dimension. After a reshape operation, the output is .
[0151] For branch 2 (the second mapping module and the second addition module) and branch 3 (the third mapping module and the third addition module), the input features are first mapped to dimension and through an MLP, and at the same time, the original features also need to be expanded by directly copying (r times and times respectively). Then, the expanded mapped features and the original features are added pointwise, and the added features are reshaped, and finally and are output respectively.
[0152] The calculation process of the above branches can be expressed as:
[0153]
[0154] where represents the input data of the input branch, .
[0155] Then, a shared MLP is used to regress the 3D coordinates of the point cloud. The MLP is designed with 5 layers. Except for the input layer, the dimensions of each layer are [1024 - 512 - 64 - 3], and a Relu activation function is added after each layer.
[0156]
[0157] Through this MLP structure, the input high-dimensional features will be mapped to 3D coordinates. For each branch, point clouds with different densities will be generated respectively, providing multi-scale results for the sparse point cloud completion task.
[0158] The above depth completion model can be a multi-layer perceptron MLP. For each point cloud coordinate feature, the process of simultaneously completing the point cloud by the depth completion model and the registration model is as follows:
[0159] (1) The registration model generates a depth point cloud
[0160] The depth map (i.e., auxiliary data) is transformed into 3D space according to the camera's internal parameter matrix to generate a point cloud in the camera coordinate system . For each pixel in the depth map and the corresponding depth value , its coordinates in 3D space can be calculated in the following way
[0161]
[0162] where is the pixel coordinate in the depth map, is the coordinate of the principal point of the camera, parameter is the length of the focal length in the x-axis direction described in pixels, is the length of the focal length in the y-axis direction described in pixels, is the depth value of this pixel in the depth map.
[0163] Then, with the help of the extrinsic parameter matrix, the points in the camera coordinate system are transformed into the original 3D point cloud space. Through the rotation matrix R and translation vector t of the camera, the point set can be rotated and translated in the following way to map the points in the camera coordinate system to the global coordinate system:
[0164]
[0165] where, is the coordinate of the pixel point in the camera coordinate system, R is the rotation matrix of the camera, t is the translation vector, is the coordinate transformed into the point cloud coordinate system.
[0166] (2) Perform point cloud registration in the registration model
[0167] Use the ICP algorithm to register the point cloud generated from the depth map and the point cloud regressed from the 3D coordinates (i.e., the point cloud generated according to the point cloud coordinate characteristics) to prepare for subsequent detailed correction. The ICP algorithm is used to find the optimal rigid transformation (rotation and translation) between two point clouds to minimize the point-to-point error between the two point clouds. The detailed operation steps are as follows:
[0168] ① Select the initial point cloud: Determine the point cloud regressed from the 3D coordinates as the target point cloud , and the point cloud generated from the depth map as the source point cloud .
[0169] ② Match the nearest neighbor points: For each source point , calculate its distance to each point in the target point cloud and select the point with the minimum distance to form a set of paired points .
[0170] ;
[0171] where, represents the Euclidean distance between two points.
[0172] ③ Calculate the transformation matrix: Assume that the source point is transformed into through rotation R and translation t, then the optimal R and t are obtained by minimizing the following objective function:
[0173]
[0174] Among them, represents the Euclidean distance.
[0175] ④ Apply transformation: For each point in the source point cloud Apply this transformation to obtain a new aligned source point cloud, which makes the source point cloud as close as possible to the target point cloud.
[0176]
[0177] ⑤ Repeat iteration: Continuously calculate a new transformation matrix and update the position of the source point cloud. When the change in the loss function for each iteration, that is, the mean square error function based on the above Euclidean distance, is lower than the set threshold or the maximum number of iterations is reached, the algorithm terminates.
[0178] (3)Perform three-dimensional coordinate adjustment in the depth completion model
[0179] After obtaining the depth point cloud, the point cloud after three-dimensional coordinate regression can be refined according to the depth point cloud to make it more accurate. Here, an MLP structure can be used to adjust the coordinates of the point cloud. The specific steps are as follows:
[0180] ① Find neighborhoods: For each point in the point cloud after three-dimensional coordinate regression , select k nearest neighbor points from and the depth point cloud respectively according to K-nearest neighbors to form two neighborhoods and .
[0181] ② Coordinate correction: Use the MLP to fuse the information from these K nearest points to "correct" the coordinates of each point in. The input of the MLP consists of two parts, namely the point cloud coordinates and the coordinate offsets of the corresponding nearest neighbor points to . Generally speaking, for each point in the point cloud after three-dimensional coordinate regression, the MLP outputs the final corrected coordinates by taking the mean of the MLP outputs of all its neighbors :
[0182]
[0183] Among them, among them is the coordinate offset of the corresponding points in the two neighborhoods, is the concatenation operation of multiple vectors.
[0184] It is worth mentioning that obtaining the point cloud encoding features based on the target point cloud data and the auxiliary data can improve the information richness in the point cloud encoding features, and further improve the accuracy of the point cloud coordinate features obtained by decoding the point cloud encoding features. The accuracy of the completed point cloud obtained based on the accurate point cloud coordinate features is improved.
[0185] Step 14: Construct a comprehensive loss function according to all the completed point clouds, and use the comprehensive loss function to optimize the encoder and the decoder to obtain an optimized encoder and an optimized decoder.
[0186] The above comprehensive loss function is used to describe the accuracy of all the completed point clouds.
[0187] Specifically, the comprehensive loss function is:
[0188]
[0189] Among them, represents the value of the comprehensive loss function, and represent hyperparameters, , represents the multi-scale completion loss, represents the overall loss:
[0190]
[0191]
[0192] Among them, represents the weighting parameter, represents the judgment probability of the discriminator for the generated completed point cloud, represents the completed point cloud, represents the judgment probability of the discriminator for the real point cloud, represents the real point cloud corresponding to the completed point cloud, represents the completion loss at the first density scale, represents the completion loss at the second density scale, represents the completion loss at the third density scale:
[0193]
[0194] Among them, , represents the th set of sampling points of the completed point cloud at the density scale, represents the number of sampling points in represents the th set of sampling points of the real point cloud at the density scale, denotes the number of sampling points in
[0195] It should be noted that after constructing the comprehensive loss function, the completed point clouds obtained from multiple target point cloud data for training through the above steps are respectively substituted into the comprehensive loss function to obtain the comprehensive loss function values corresponding to each target point cloud data. Then, the average value of these comprehensive loss function values is calculated, and it is determined whether the average value is less than or equal to the preset loss value. If so, the encoder is used as the optimized encoder and the decoder is used as the optimized decoder. Otherwise, the parameters in the encoder and decoder are adjusted, and the step of encoding the target point cloud data and the auxiliary data using the encoder to obtain the point cloud encoding features is returned.
[0196] In some embodiments of the present application, a discriminator can be used to analyze the difference between each generated completed point cloud and the real point cloud corresponding to the completed point cloud (with the same density scale as the completed point cloud) to improve the accuracy of the completed point cloud (the discriminator can give the judgment probability that the completed point cloud is a real point cloud, and the encoder and decoder can be trained based on the judgment probability and the completed point cloud can be regenerated). As Figure 5 shown, the discriminator includes a serial mapping layer, a connection layer, a mapping layer, a max pooling layer, and a fully connected layer connected in sequence. The input data at the input ends of the serial mapping layer and the connection layer is the completed point cloud, and the output data at the output end of the fully connected layer is the probability that the completed point cloud is a real point cloud. The calculation expression of the discriminator is:
[0197] ;
[0198] ;
[0199] ;
[0200] where denotes the output data of the mapping layer, denotes the completed point cloud, denotes the operation of the serial mapping layer, denotes the operation of the connection layer, denotes the operation of the mapping layer, denotes the maximum value in the j-th channel, denotes the global vector after max pooling, denotes the operation of the fully connected layer, denotes the activation function, denotes the probability that the completed point cloud is a real point cloud.
[0201] Step 15: Use the optimized encoder and the optimized decoder to complete the point cloud data to be completed, and obtain the final completed point cloud.
[0202] The above point cloud data to be completed is the point cloud data that needs to be completed.
[0203] Specifically, the optimized encoder is used to encode the point cloud data to be completed and the auxiliary data corresponding to the point cloud data to be completed, obtaining point cloud encoding features. Then, the optimized decoder is used to decode the point cloud encoding features, obtaining point cloud coordinate features at multiple density scales, and depth repair is performed on each point cloud coordinate feature according to the auxiliary data, obtaining the final completed point cloud at each density scale.
[0204] The method of the present application will be exemplarily described below with a specific example.
[0205] The process of the method of the present application is as Figure 6 shown. The lidar collects point cloud data, and at the same time the camera collects image data (i.e., auxiliary data). Then, image enhancement feature learning (i.e., the processing process of the multi-scale feature extraction model in the encoder) is performed: continuous convolution is performed on the image data to obtain a feature map, and then point cloud structure information learning (i.e., the processing process of the multi-scale feature fusion model in the encoder) is performed on the point cloud data: on the one hand, global skeleton extraction is performed, and multi-head interactive attention processing is performed on the multi-modal data; on the other hand, local structure extraction is performed, and multi-core edge convolution processing is performed on the multi-modal data. Then, point cloud coordinate reconstruction (i.e., the processing process of the multi-layer feature decoding model in the decoder) is performed: feature expansion and three-dimensional coordinate regression processing are performed to obtain point cloud coordinate features, and then depth detail correction (i.e., the processing process of the registration model and the depth completion model in the decoder) is performed: the depth map is corrected according to the point cloud after projection and ICP registration to obtain the completed point cloud, and finally the difference between the completed point cloud and the real point cloud is obtained through the discriminator module.
[0206] It is worth mentioning that obtaining point cloud encoding features based on the target point cloud data and the auxiliary data can improve the information richness in the point cloud encoding features, thereby improving the accuracy of the point cloud coordinate features obtained by decoding the point cloud encoding features. The accuracy of the completed point cloud obtained according to the accurate point cloud coordinate features is improved. Constructing a comprehensive loss function to optimize the encoder and the decoder can improve the performance of the encoder and the decoder, and further improve the accuracy of point cloud completion.
[0207] In addition, the method of the present application enhances the information extraction of auxiliary data through enhanced continuous convolution, captures the point cloud structure through multi-core edge convolution, and then performs early fusion of the features of the two modalities. Then, with the help of two independent branch networks, the skeleton information and local structure of the point cloud are extracted and fused respectively, so as to enhance the feature expression ability and achieve local and global consistency. At the same time, the point cloud generated by the auxiliary data is used to repair and adjust the details of the multi-scale point cloud, thereby improving the completion accuracy and providing technical support for high-precision three-dimensional perception in fields such as unmanned driving and intelligent manufacturing.
[0208] An exemplary description of the point cloud completion device based on cross-modal and depth repair provided by the present application will be given below.
[0209] As Figure 7 shown, an embodiment of the present application provides a point cloud completion device based on cross-modal and depth repair. The point cloud completion device 700 based on cross-modal and depth repair includes:
[0210] An acquisition module 701, configured to acquire target point cloud data and auxiliary data corresponding to the target point cloud data;
[0211] An encoding module 702, configured to encode the target point cloud data and the auxiliary data by using an encoder to obtain point cloud encoding features; the point cloud encoding features are used to describe the local structure information and global skeleton information of the target point cloud data;
[0212] A decoding module 703, configured to decode the point cloud encoding features by using a decoder to obtain point cloud coordinate features at multiple density scales, and perform depth repair on each point cloud coordinate feature according to the auxiliary data to obtain a completed point cloud at each density scale; the point cloud coordinate features are used to describe the coordinates of multiple data points at the corresponding density scale;
[0213] An optimization module 704, configured to construct a comprehensive loss function according to all the completed point clouds, and use the comprehensive loss function to optimize the encoder and the decoder to obtain an optimized encoder and an optimized decoder; the comprehensive loss function is used to describe the accuracy of all the completed point clouds;
[0214] A completion module 705, configured to complete the to-be-completed point cloud data by using the optimized encoder and the optimized decoder to obtain a final completed point cloud.
[0215] It should be noted that for the information interaction, execution process, etc. between the above-mentioned device / unit, since it is based on the same concept as the method embodiment of the present application, its specific functions and the technical effects brought can be specifically referred to in the method embodiment part, and will not be elaborated here.
[0216] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.
[0217] As Figure 8 shown, an embodiment of the present application provides a terminal device. The terminal device D10 in this embodiment includes: at least one processor D100 ( Figure 8 only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, the steps in any of the foregoing method embodiments are implemented.
[0218] Specifically, when the processor D100 executes the computer program D102, by obtaining a point cloud encoding feature based on the target point cloud data and the auxiliary data, the information richness in the point cloud encoding feature can be improved, and then the accuracy of the point cloud coordinate feature obtained by decoding the point cloud encoding feature can be improved. The accuracy of the completed point cloud obtained based on the accurate point cloud coordinate feature is improved. Constructing a comprehensive loss function to optimize the encoder and the decoder can improve the performance of the encoder and the decoder, and further improve the accuracy of point cloud completion.
[0219] The so-called processor D100 may be a central processing unit (CPU), and this processor D100 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0220] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as the hard disk or memory of the terminal device D10. In some other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk equipped on the terminal device D10, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory D101 may also include both the internal storage unit and the external storage device of the terminal device D10. The memory D101 is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program, etc. The memory D101 may also be used to temporarily store data that has been output or is to be output.
[0221] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above various method embodiments can be implemented.
[0222] An embodiment of the present application provides a computer program product, and when the computer program product runs on a terminal device, the terminal device is enabled to execute the steps in the above various method embodiments.
[0223] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the device / terminal equipment based on the cross-modal and depth-inpainting point cloud completion method, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc.
[0224] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0225] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0226] The above is the preferred implementation manner of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle described in this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A point cloud completion method based on cross-modality and deep inpainting, characterized in that: include: Acquire target point cloud data and auxiliary data corresponding to the target point cloud data; Encoding the target point cloud data and the auxiliary data using an encoder to obtain point cloud coding features; the point cloud coding features are used to describe local structural information and global skeleton information of the target point cloud data; Decoding the point cloud coding features using a decoder to obtain point cloud coordinate features at multiple density scales, and performing deep patching on each of the point cloud coordinate features according to the auxiliary data to obtain a completed point cloud at each density scale; the point cloud coordinate features are used to describe the coordinates of multiple data points at the corresponding density scale; Constructing a comprehensive loss function according to all completed point clouds, and optimizing the encoder and the decoder using the comprehensive loss function to obtain an optimized encoder and an optimized decoder; the comprehensive loss function is used to describe the accuracy of all completed point clouds; The optimized encoder and the optimized decoder are used to complete the point cloud data to obtain the final completed point cloud; Wherein, the encoder includes a multi-layer feature extraction model and a multi-layer feature fusion model; The input data of the first input end of the multi-layer feature extraction model is the auxiliary data, the input data of the second input end of the multi-layer feature extraction model is the target point cloud data, and the output end of the multi-layer feature extraction model is connected to the first input end of the multi-layer feature fusion model; The input data of the second input end of the multi-layer feature fusion model is the target point cloud data, and the output data of the output end of the multi-layer feature fusion model is the point cloud coding feature; The decoder includes a multi-layer feature decoding model, a registration model, and a depth completion model; The output end of the multi-layer feature decoding model and the output end of the registration model are both connected to the input end of the depth completion model, the input data of the input end of the multi-layer feature decoding model is the point cloud coding feature, the output data of the output end of the multi-layer feature decoding model is the point cloud coordinate feature under multiple density scales, the input data of the input end of the registration model is the auxiliary data, and the output data of the output end of the depth completion model is the completed point cloud under multiple density scales.
2. The point cloud completion method according to claim 1, characterized in that: The multi-layer feature extraction model includes: a first residual neural module, a second residual neural module, a third residual neural module, a first upsampling module, a second upsampling module, a first enhanced continuous convolution module, a second enhanced continuous convolution module and a third enhanced continuous convolution module; The input end of the first residual neural module is the first input end of the multi-layer feature extraction model, the input end of the first enhanced continuous convolution module is the second input end of the multi-layer feature extraction model, and the output end of the first enhanced continuous convolution module, the output end of the second enhanced continuous convolution module and the output end of the third enhanced continuous convolution module are all output ends of the multi-layer feature extraction model; The output end of the first residual neural module is respectively connected to the input end of the second residual neural module and the input end of the first enhanced continuous convolution module, the output end of the second residual neural module is respectively connected to the input end of the third residual neural module and the input end of the first upsampling module, the output end of the first upsampling module is connected to the input end of the second enhanced continuous convolution module, the output end of the third residual neural module is connected to the input end of the second upsampling module, the output end of the second upsampling module is connected to the input end of the third enhanced continuous convolution module, the output end of the first enhanced continuous convolution module is connected to the input end of the second enhanced continuous convolution module, and the output end of the second enhanced continuous convolution module is connected to the input end of the third enhanced continuous convolution module.
3. The point cloud completion method according to claim 1, characterized in that: The multi-layer feature fusion model includes: a first multi-core edge convolution module, a second multi-core edge convolution module, a third multi-core edge convolution module, a fourth multi-core edge convolution module, a first cascade module, a second cascade module, a third cascade module, a first global structure perception module, a second global structure perception module, a clustering module, a maximum pooling module, a first expansion module, and a second expansion module; The input end of the first multi-core edge convolution module is the second input end of the multi-layer feature fusion model, the first input end of the first cascade module is the first input end of the multi-layer feature fusion model, and the output end of the third cascade module is the output end of the multi-layer feature fusion model; The output end of the first multi-core edge convolution module is respectively connected to the input end of the second multi-core edge convolution module and the input end of the second cascade module, the output end of the second multi-core edge convolution module is respectively connected to the input end of the third multi-core edge convolution module and the input end of the second cascade module, the output end of the third multi-core edge convolution module is respectively connected to the second input end of the first cascade module and the input end of the second cascade module, the output end of the first cascade module is respectively connected to the input end of the fourth multi-core edge convolution module, the input end of the first global structure perception module, and the input end of the clustering module, the output end of the fourth multi-core edge convolution module is respectively connected to the output end of the third cascade module. The input end of the cascade module is connected to the input end of the first global structure perception module, the output end of the first global structure perception module is connected to the input end of the second global structure perception module, the output end of the second global structure perception module is connected to the input end of the first expansion module, the output end of the first expansion module is connected to the input end of the third cascade module, the output end of the second cascade module is connected to the input end of the maximum pooling module, the output end of the maximum pooling module is connected to the input end of the second expansion module, and the output end of the second expansion module is connected to the input end of the third cascade module.
4. The point cloud completion method according to claim 3, characterized in that: The first global structure perception module and the second global structure perception module are both global structure perception modules; The global structure perception module includes a convolution layer, a first nonlinear transformation layer, a second nonlinear transformation layer, a third nonlinear transformation layer, a multi-head interactive attention layer, a first residual normalization layer, a feature adaptive selection layer, and a second residual normalization layer; The input end of the convolutional layer is the input end of the global structure perception module, and the output end of the second residual normalization layer is the output end of the global structure perception module; The output end of the convolutional layer is respectively connected to the input end of the first nonlinear transformation layer, the input end of the second nonlinear transformation layer, the input end of the third nonlinear transformation layer, and the input end of the first residual normalization layer; the input end of the multi-head interactive attention layer is respectively connected to the output end of the first nonlinear transformation layer, the output end of the second nonlinear transformation layer, and the output end of the third nonlinear transformation layer; the output end of the multi-head interactive attention layer is connected to the input end of the first residual normalization layer; the output end of the first residual normalization layer is respectively connected to the input end of the second residual normalization layer and the input end of the feature adaptive selection layer; the output end of the feature adaptive selection layer is connected to the input end of the second residual normalization layer.
5. The point cloud completion method according to claim 1, characterized in that: The multi-layer feature decoding model includes a first mapping module, a second mapping module, a third mapping module, a first addition module, a second addition module, a third addition module, and a multi-layer perception module; The input end of the first mapping module, the input end of the second mapping module, the input end of the third mapping module, the input end of the first adding module, the input end of the second adding module, and the input end of the third adding module are all input ends of the multi-layer feature decoding model, and the output end of the multi-layer perception module is the output end of the multi-layer feature decoding model; The output end of the first mapping module is connected to the input end of the first adding module, the output end of the second mapping module is connected to the input end of the second adding module, and the output end of the third mapping module is connected to the input end of the third adding module.
6. The point cloud completion method according to claim 1, characterized in that: The comprehensive loss function is: in, represents the value of the comprehensive loss function, and represents the hyperparameter, , represents the multi-scale completion loss, Represents the overall loss: in, represents the weighting parameter, Represents the probability of the discriminator's judgment on the generated completed point cloud, represents the completion of the point cloud, Represents the probability of the discriminator's judgment on the real point cloud, Indicates the real point cloud corresponding to the completed point cloud, represents the completion loss at the first density scale, represents the completion loss at the second density scale, Represents the completion loss at the third density scale: in, , Indicates The sampling point set of the completed point cloud at the density scale, express The number of sampling points in Indicates The sampling point set of the real point cloud at the density scale, express The number of sampling points in .
7. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the point cloud completion method based on cross-modality and depth patching as described in any one of claims 1 to 6 is implemented.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the point cloud completion method based on cross-modality and depth patching as described in any one of claims 1 to 6 is implemented.