SwinPoinTr-based point cloud completion method and device

By using the hierarchical geometry-aware Transformer module and shift window method in the point cloud completion method, combining the farthest point sampling and k-nearest neighbor algorithm, the problem of weak local detail completion ability in the existing technology is solved, and higher point cloud completion integrity and accuracy are achieved.

CN120107085APending Publication Date: 2025-06-06QUANZHOU FREEZING POINT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510115656.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the point cloud completion task, it is difficult to accurately capture subtle edges and texture features in the existing technology, and there is a problem of weak local detail completion capabilities.

Method used

The point cloud completion method based on SwinPoinTr is adopted, and the hierarchical geometry-aware Transformer module is used for encoding and decoding. Combining the farthest point sampling algorithm and the k-nearest neighbor algorithm, local and global features are extracted, and communication between different windows is carried out through the shift window method.

Benefits of technology

It improves the integrity and accuracy of point cloud completion, enhances the ability to extract local features, and takes into account the extraction of global features, so as to perform better in tasks with rich details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107085A_ABST
    Figure CN120107085A_ABST
Patent Text Reader

Abstract

The invention provides a point cloud completion method and device based on SwinPoinTr, and the method comprises the steps: obtaining a missing point cloud, obtaining a central point from the missing point cloud through employing a farthest point sampling algorithm, forming a central point cloud, carrying out the extraction of region points based on the central point cloud, obtaining a region point set, and carrying out the extraction of region points based on the region point set, adding the central point cloud and the regional point set to obtain a feature sequence set; the feature sequence set is input into a hierarchical geometric perception Transform module for coding and decoding, then decoding features and a prediction center point of a complete point cloud are obtained, the hierarchical geometric perception Transform module calculates self-attention in a window, and communication of different windows is carried out by using a window shifting method; and according to the decoding features and the prediction center point, respectively recovering the point cloud of the local area where the prediction center point is located, and obtaining a complemented complete point cloud. According to the invention, the completeness and accuracy of point cloud complementation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of point cloud data, and in particular to a point cloud completion method and device based on SwinPoint. Background Art

[0002] Point cloud data is used in many different fields, including autonomous driving, robotics, etc. It has a very uniform structure and can represent very complex models with a small amount of data. However, in practical applications, the collected point cloud data is often incomplete due to the occlusion of objects, differences in the reflectivity of the target surface material, and the limitations of the resolution and viewing angle of the visual sensor. Incomplete point cloud data will seriously affect subsequent processing. Therefore, completing the incomplete point cloud data and restoring the original shape is of great significance to downstream tasks.

[0003] The existing processing method is the Diverse Point Cloud Completion with Geometry-Aware Transformers (PoinTr). It uses a feature extraction module with feature point sampling to abstract the original point cloud into a set of unordered point proxies, so that the model can process point cloud data more efficiently, while avoiding the problems caused by order dependence, making the model more robust to point cloud inputs with different arrangement orders. Then the point cloud is input into the classic encoder-decoder structure, namely the encoder-decoder structure. The encoder is responsible for extracting the features of the input point cloud and converting the point cloud data into an abstract feature representation, which contains key information such as the structure and shape of the point cloud. The decoder gradually restores and generates point proxies of the complete point cloud based on the features output by the encoder to realize the function of point cloud completion. At the same time, in order to solve the problem of lack of inductive bias in the classic encoder-decoder structure, the network introduces geometric information as inductive bias and converts the encoder-decoder structure into an encoder-decoder structure with geometry perception (Geometry-Aware Transformers Encoder-Decoder). Finally, the folding network (FoldingNet) converts the point proxy features of the complete point cloud into the spatial coordinates of the points. However, the visual Transformer architecture used by this network is based on the global attention mechanism, and its ability to model local structures and spatial relationships is relatively weak. It performs poorly when dealing with tasks with rich details and key local features. In edge detection and point cloud completion tasks, it is difficult to accurately capture subtle edge and texture features, and there is a problem of weak local detail completion capabilities. Summary of the invention

[0004] In order to solve the above problems in the prior art, the present invention provides a point cloud completion method and device based on SwinPoint, so as to improve the completeness and accuracy of point cloud completion.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is:

[0006] In a first aspect, the present invention provides a point cloud completion method based on SwinPoint, comprising the steps of:

[0007] S1. Obtain the missing point cloud, use the farthest point sampling algorithm to obtain the center point from the missing point cloud to form a center point cloud, and perform regional point extraction based on the center point cloud to obtain a regional point set, and add the center point cloud and the regional point set to obtain a feature sequence set;

[0008] S2, after inputting the feature sequence set into the hierarchical geometry-aware Transformer module for encoding and decoding, the decoded features and the predicted center point of the complete point cloud are obtained, and the hierarchical geometry-aware Transformer module calculates self-attention within the window and uses the shift window method to communicate between different windows;

[0009] S3. Restore the point cloud of the local area where the predicted center point is located according to the decoded features and the predicted center point to obtain a completed complete point cloud.

[0010] The beneficial effects of the present invention are as follows: the hierarchical geometry-aware Transformer module used in the present invention calculates self-attention within the window, improves the ability to extract local features, and uses a shifting window method to communicate between different windows, taking into account the extraction of global features, thereby improving the completeness and accuracy of point cloud completion.

[0011] Optionally, the step S1 includes the following steps:

[0012] S11, obtaining a missing point cloud, and using a farthest point sampling algorithm to obtain a center point from the missing point cloud to form a center point cloud;

[0013] S12. Use a k-nearest neighbor algorithm to construct an edge relationship for each center point and determine a local area to obtain a regional point set, and use a lightweight dynamic graph convolutional neural network with a hierarchical downsampling structure to extract local features from the regional point set to obtain a regional feature set;

[0014] S13. After performing linear layer dimensionality increase on the global position of the center point coordinates, the global position is added to the local features of the corresponding center point in the regional feature set to obtain a feature sequence set.

[0015] Optionally, the step S1 further includes the steps of:

[0016] S14. Reshape the data shape of the feature sequence set to adapt to the window in the hierarchical geometry-aware Transformer module.

[0017] Optionally, the step S14 is specifically as follows:

[0018] The window size is determined according to the number of center points, and then each feature sequence in the feature sequence set is divided into first feature blocks of equal length. Finally, the first feature blocks in the corresponding number of feature sequences are spliced ​​at corresponding positions in different dimensions according to the window size to obtain a multi-dimensional point cloud feature map.

[0019] Optionally, in step S2, inputting the feature sequence set into a hierarchical geometry-aware Transformer module for encoding comprises the following steps:

[0020] S21, the feature splitting layer splits the multidimensional point cloud feature map obtained by reshaping the feature sequence set into second feature blocks of corresponding sizes, and the second feature blocks split at corresponding positions are spliced ​​and flattened into first feature data;

[0021] S22, the linear embedding layer in the first level projects the first feature data into a fixed dimension, and outputs it to a shifted geometry-aware Transformer block after normalization to obtain second feature data, and the shifted geometry-aware Transformer block performs geometry-aware attention calculation within the window, and then translates the window to perform geometry-aware attention calculation across windows;

[0022] S23, the feature merging layer in the second level uses an image fusion algorithm to increase the number of channels of the feature map, and then outputs it to a shifted geometry-aware Transformer block for self-attention calculation, and then passes through the third and fourth levels with the same structure as the second level to obtain and output the encoded data.

[0023] Optionally, in step S2, inputting the feature sequence set into a hierarchical geometry-aware Transformer module for decoding, and obtaining the decoded features and the predicted center point of the complete point cloud comprises the following steps:

[0024] S24, calculating a query vector according to the encoded data;

[0025] S25, decoding is performed according to the query vector and the multi-dimensional point cloud feature map to obtain a decoding feature;

[0026] S26. Input the decoded features into a linear layer for dimensionality reduction to obtain a predicted center point of the complete point cloud.

[0027] Optionally, the step S24 is specifically:

[0028] First, a linear layer is used to project the encoded data to a higher dimension to obtain high-dimensional data, and then a maximum pooling operation is performed on the high-dimensional data to obtain encoded features, and then an MLP layer is used to reconstruct the encoded features into a query vector.

[0029] In a second aspect, the present invention provides a point cloud completion device based on SwinPoinTr, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, a point cloud completion method based on SwinPoinTr according to the first aspect is implemented.

[0030] Among them, the technical effect corresponding to the point cloud completion device based on SwinPoinTr provided in the second aspect refers to the relevant description of the point cloud completion method based on SwinPoinTr provided in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 A schematic diagram of the main process of a point cloud completion method based on SwinPointer according to an embodiment of the present invention;

[0032] Figure 2 It is a structural diagram of the SwinPoinTr model involved in an embodiment of the present invention;

[0033] Figure 3 This is a flowchart of a hierarchical geometry-aware Transformer module involved in an embodiment of the present invention;

[0034] Figure 4 This is a flowchart of a shifted geometry-aware Transformer block involved in an embodiment of the present invention;

[0035] Figure 5 This is a structural diagram of the geometric perception multi-head self-attention involved in an embodiment of the present invention;

[0036] Figure 6 Visualization of the results of completing different forms of missing point clouds for different models;

[0037] Figure 7 It is a schematic diagram of the framework of a point cloud completion device based on SwinPointer according to an embodiment of the present invention.

[0038] Description of reference numerals:

[0039] 1. A point cloud completion device based on SwinPoinTr;

[0040] 2. Processor;

[0041] 3. Memory. DETAILED DESCRIPTION

[0042] In order to better understand the above technical solution, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0043] Embodiment 1

[0044] Please refer to Figures 1 to 6 , a point cloud completion method based on SwinPoinTr, comprising the steps of:

[0045] S1. Obtain the missing point cloud, use the farthest point sampling algorithm to obtain the center point from the missing point cloud, form a center point cloud, and extract regional points based on the center point cloud to obtain a regional point set. Add the center point cloud and the regional point set to obtain a feature sequence set.

[0046] In this embodiment, step S1 includes the following steps:

[0047] S11. Obtain the missing point cloud, and use the farthest point sampling algorithm to obtain the center point from the missing point cloud to form a center point cloud.

[0048] In one example, an incomplete or sparse point cloud is obtained, and the point cloud coordinates are input into the network, and 1024 center points are obtained from the input data using the farthest point sampling method.

[0049] S12. Use the k-nearest neighbor algorithm to construct the edge relationship of each center point and determine the local area to obtain the regional point set. Use a lightweight dynamic graph convolutional neural network with a hierarchical downsampling structure to extract local features of the regional point set to obtain the regional feature set.

[0050] In one example, the k-nearest neighbor algorithm is used to find 64 neighbors for each center point, and edge relationships are constructed between the point and its neighbors. The feature dimension obtained is (1024, 512), and the point coordinate dimension is (1024, 3).

[0051] S13. After performing linear layer dimensionality increase on the global position of the center point coordinates, the global position is added to the local features of the corresponding center point in the regional feature set to obtain a feature sequence set.

[0052] In one example, the coordinate information is dimensionally upgraded to (1024, 512) through two linear layers and input into the feature extraction module together with the feature information to obtain the feature sequence set {T 1 , T 2 ....T N}, whose dimension is (1024, 392).

[0053] Among them, the individual point cloud coordinates can only describe the position information of each point in space, but cannot directly reflect the feature relationship between points. If the point coordinates are directly used as network input, it is difficult for the network to effectively capture the local feature relationship and global feature relationship in the point cloud data, and if the network directly uses the input data as processing data, the network training burden is too heavy.

[0054] Therefore, refer to Figure 2 It can be seen that the present invention first samples the farthest point of the input point cloud data to obtain N points, and uses the k-nearest neighbor algorithm (k-NN) with these N points as the center to construct its edge relationship and determine the local area. Then, a lightweight dynamic graph convolutional neural network (Dynamic Graph Convolutional Ne-ural Network, DGCNN) with a hierarchical downsampling structure is used to obtain the local features of the point center. At the same time, the global position of the center point coordinates is explicitly encoded through the coordinate dimension upgrade module, and added to the local features of the center point to complete the combination of global features and local features to obtain a feature sequence set.

[0055] S14. Reshape the data of the feature sequence set to fit the window in the hierarchical geometry-aware Transformer module.

[0056] In this embodiment, in order to make the calculated point cloud features suitable for the input of the shift window encoder in the hierarchical geometry-aware Transformer module, it is necessary to reshape the data shape of the feature sequence set. Specifically, step S14 is as follows:

[0057] The window size is determined according to the number of center points, and then each feature sequence in the feature sequence set is divided into first feature blocks of equal length. Finally, the first feature blocks in the corresponding number of feature sequences are spliced ​​at corresponding positions in different dimensions according to the window size to obtain a multi-dimensional point cloud feature map.

[0058] In one example, a feature sequence set is input into a feature reshaping module. The feature reshaping module reshapes each feature sequence T in the feature sequence set. m Cut into 8 7×7 blocks with dimensions of (1024, 8, 7, 7). Then square-join the blocks in the corresponding number of 1024 feature sequences to obtain 8 feature maps with dimensions of (8, 224, 224), thus completing feature reshaping.

[0059] S2. After the feature sequence set is input into the hierarchical geometry-aware Transformer module for encoding and decoding, the decoded features and the predicted center point of the complete point cloud are obtained. The hierarchical geometry-aware Transformer module calculates self-attention within the window and uses the shift window method to communicate between different windows.

[0060] Among them, although the Transformer can capture long-distance dependencies through the self-attention mechanism, it also needs to store a large number of intermediate matrices when calculating self-attention, especially when processing long sequences, the memory consumption will become unacceptable. This feature will lead to an extremely limited number of points input to the network, and it is difficult for the network to learn all the features of the input point cloud, especially in more complex point cloud models, this situation will be more obvious.

[0061] Reference Figure 2 It can be seen that this embodiment proposes a hierarchical geometry-aware Transformer module based on the shift window scheme. In the encoding stage, it is a hierarchical geometry-aware Transformer encoding module. This module can efficiently extract point cloud geometric features while reducing the size of feature maps layer by layer and reducing memory overhead, which is conducive to the processing of complex point cloud models.

[0062] like Figure 3 As shown in , the feature splitting module splits the feature map obtained by feature reshaping into non-overlapping small blocks of equal size, and the small blocks split at corresponding positions are spliced ​​and flattened into a sequence. The linear embedding layer projects the input data into a fixed dimension and normalizes the input data to ensure that the input features have similar scales, which is conducive to the learning of subsequent layers. The feature merging layer splices adjacent small blocks in each window to reduce the spatial size of the feature map, which helps the model capture higher-level abstract features. At the same time, the feature merging layer increases the number of channels of the feature map to accommodate more abstract information. The shifted geometry-aware Transformer block calculates geometry-aware attention in non-overlapping local windows. The process is as follows Figure 4 shown.

[0063] Therefore, in this embodiment, inputting the feature sequence set into the hierarchical geometry-aware Transformer module for encoding in step S2 includes the following steps:

[0064] S21, the feature splitting layer splits the multi-dimensional point cloud feature map obtained by reshaping the feature sequence set into second feature blocks of corresponding sizes, and the second feature blocks split at corresponding positions are spliced ​​and flattened into first feature data.

[0065] S22. The linear embedding layer in the first level projects the first feature data into a fixed dimension, normalizes it, and outputs it to a shifted geometry-aware Transformer block to obtain the second feature data. The shifted geometry-aware Transformer block performs geometry-aware attention calculation within the window, and then translates the window to perform geometry-aware attention calculation across windows.

[0066] S23, the feature merging layer in the second level uses an image fusion algorithm to increase the number of channels of the feature map, and then outputs it to a shifted geometry-aware Transformer block for self-attention calculation, and then passes through the third and fourth levels with the same structure as the second level to obtain and output the encoded data.

[0067] Among them, a complete shifted geometry-aware Transformer block mainly consists of two parts. The first part performs geometry-aware attention calculation within the window, and the second part translates the window to perform geometry-aware attention calculation across windows.

[0068] In the first part, the point cloud data will first be layer normalized (Layer Normalization, LN), and then the geometry-aware multi-head self-attention layer (Geometry Aware-Multi head Self Attention, GA-MSA) will calculate the geometry-aware attention and add it to the input data. The calculation formula is as follows:

[0069] Z i =GAMSA(LN(Z i-1 ))+Z i-1 (1)

[0070] In the formula, GAMSA is computational geometry-aware multi-head self-attention, LN is layer normalization, and Z i-1 is the input data, Z i To obtain the data after calculation.

[0071] The structure of the geometrically-aware multi-head self-attention layer is as follows: Figure 5 As shown. First, the input data is linearly projected to generate three matrices Q, K, and V, and then input into the Multi Head Self Attention (MSA) layer. The Q matrix is ​​the query matrix, the K matrix is ​​the key matrix, and the V matrix is ​​the value matrix. The multi-head self-attention layer first multiplies the Q matrix with the K matrix to calculate the similarity, and then performs a weighted match on the similarity calculation result with the V matrix to complete the extraction of different features. The calculation formula of the multi-head self-attention is as follows:

[0072]

[0073] Where, dk It is a scaling factor used to prevent the gradient from disappearing due to the input value of the softmax function being too large.

[0074] Among them, the single multi-head self-attention mechanism lacks some inductive biases and cannot model the point cloud structure well. In order to facilitate the model to better utilize the inductive bias of the geometric structure of the point cloud, the geometry-aware multi-head self-attention layer uses the k-nearest neighbor algorithm to capture the key features of the point cloud while calculating the self-attention, and combines it with the self-attention calculation results to form a self-attention calculation mechanism with geometry perception ability.

[0075] Getting Z i After layer normalization again, the multi-layer perceptron (MLP) completes the linear transformation of the extracted features to capture the relationship between different features and obtain Z i+1 . And Z i+1 With Z i Add together to complete the feature extraction of the first part. The first part calculates attention within each window. Although it reduces the amount of attention calculation, it also limits the connection and communication between models across windows. Therefore, the shifted window method translates the entire window on the feature map, cuts and fills the protruding part, constructs a new window, and then combines it with the geometry-aware multi-head self-attention layer to form a geometry-aware multi-head self-attention layer with shifted windows (Shifted Windows Geometry Aware-Multihead Self Attention, SGA-MSA). The second part replaces the attention calculation module with a geometry-aware multi-head self-attention layer with shifted windows, and performs calculations again based on the first part to complete the feature encoding.

[0076] in, Figure 5 Connect means connection; Linear means linear; Cat is the abbreviation of Concatenate, which means splicing; Max Pooling means maximum pooling; KNN Query means nearest neighbor query.

[0077] In one example, the feature map is input into a hierarchical geometry-aware Transformer encoding module. First, a 4×4 convolution kernel is used to cut the feature map into patches (small blocks), and its dimension becomes (128, 56, 56). Then the window is cut into 7×7 windows, and the geometry-aware multi-head self-attention is calculated for each window. The window is shifted horizontally and vertically to form a new window layout. The geometry-aware multi-head self-attention is calculated for the new window, and its output dimension is still (128, 56, 56). Then image fusion is used to double the data channel and reduce the length and width to 1 / 2 of the original, with a dimension of (256, 28, 28). Repeat the attention calculation and image fusion steps twice to obtain the encoded data with a dimension of (1024, 7, 7).

[0078] In this embodiment, in step S2, the feature sequence set is input into the hierarchical geometry-aware Transformer module for decoding, and obtaining the decoded features and the predicted center point of the complete point cloud includes the following steps:

[0079] S24. Calculate a query vector based on the encoded data.

[0080] Wherein, step S24 is specifically as follows:

[0081] First, a linear layer is used to project the encoded data to a higher dimension to obtain high-dimensional data. Then, a maximum pooling operation is performed on the high-dimensional data to obtain the encoded features. Finally, an MLP layer is used to reconstruct the encoded features into a query vector.

[0082] In one example, the encoded data is input into the query generation module, and the query vector Q is calculated through three linear layers, whose dimension is (512, 384), and then input into the hierarchical geometry-aware Transformer decoding module.

[0083] S25. Decode according to the query vector and the multi-dimensional point cloud feature map to obtain a decoded feature.

[0084] In one example, the decoding module directly performs geometry-aware multi-head self-attention calculation on the input data to obtain a decoding feature whose dimension is (512, 384).

[0085] S26. Input the decoded features into the linear layer for dimensionality reduction to obtain the predicted center point of the complete point cloud.

[0086] In one example, the decoded data is input into a linear layer for dimensionality reduction to obtain predicted 512 center point coordinates with a dimension of (512, 3).

[0087] S3. Restore the point cloud of the local area where the predicted center point is located according to the decoded features and the predicted center point, and obtain the completed complete point cloud.

[0088] In one example, the predicted center point coordinates are input into the multi-layer reconstruction head together with the decoded data to perform high-resolution reconstruction of the completed point cloud. The multi-layer reconstruction head first extracts the feature information of the decoded data by an MLP layer, with an output dimension of (16384, 6), and then converts the feature information into point cloud coordinate information by a linear layer, with an output dimension of (16384, 3).

[0089] Finally, the high-resolution complete point cloud output by the network is extracted.

[0090] Therefore, in order to illustrate the advantages of the method used in this embodiment, the following evaluation is performed:

[0091] In this evaluation, the dataset used was a self-made incomplete Pleurotus eryngii point cloud dataset. All models kept their optimal hyperparameters unchanged and trained and predicted the same dataset in the same environment.

[0092] (1) Training environment parameters

[0093] The model training was completed under the Windows 10 operating system, using an Intel Core i9-10900K CPU with a main frequency of 3.7 GHz and an NVIDIA GeForce RTX2080 Super GPU. The deep learning environment was python 3.10.0, CUDA 11.8, pytorch 2.2.2, and open3d 0.18.0.

[0094] The number of model training iterations (Epoch) is 150 rounds, the initial learning rate is 0.0001, and it is updated every 21 steps. The learning rate gradually decays with a decay factor of 0.9 until the minimum learning rate is 0.000002.

[0095] (2) Evaluation indicators

[0096] Chamfer Distance (CD), Earth Mover's Distance (EMD) and F1 Score are used as model evaluation indicators.

[0097] Among them, chamfer distance is a method used to measure the similarity between two point clouds. It measures their similarity by calculating the distance from each point in the two point clouds to the nearest point in the other point cloud. The calculation formula is:

[0098]

[0099] Where P and S represent two given point clouds, where P contains the point set {P 1 ,P 2 …P m}, S contains the point set {S 1 ,S 2 …S n}, m and n are the number of points in P and S points respectively, ||p i -s j || means p i Dot and s j Euclidean distance between points.

[0100] The earth moving distance measures the similarity between two point clouds by calculating the minimum distance required to move all points in one point cloud to another point cloud. The calculation formula is:

[0101]

[0102] Where N is the number of points in the point cloud.

[0103] Among them, the F1 score takes into account the precision and recall rate, and calculates the percentage of the reconstructed point cloud within a certain distance from the real point cloud, which represents the accuracy of the reconstruction. Its calculation formula is as follows:

[0104]

[0105] Where ρ(d) and R(d) represent the precision and recall respectively with d as the chamfer distance determination threshold.

[0106] Among them, the smaller the chamfer distance, the closer the output point cloud is to the surface of the Pleurotus eryngii, which can reflect the integrity of the point cloud shape and the smoothness of the edge contour. The smaller the earth movement distance, the smaller the distance, which means that each spatial point output by the completion network can be paired with the point cloud position of the complete point cloud at a small distance, which reflects the similarity between the resolution and density of the point cloud. The higher the F1 score, the more points generated by the completion network are distributed in the correct space, reflecting the improvement of the accuracy of the generated point cloud.

[0107] (3) Comparison of different models on residual cloud completion and reconstruction

[0108] In order to evaluate the completion performance of SwinPoinTr on the incomplete Pleurotus eryngii point cloud, this embodiment is compared with representative models (FoldingNet, GRNet, PCN, PoinTr, SnowFlakeNet, AdaPoinTr, TopNet). Each model maintains the original optimal hyperparameter settings and is trained and verified on the data set produced in this embodiment. The comparative analysis of the completion results of these methods on the Pleurotus eryngii point cloud dataset is shown in Table 1.

[0109] Table 1 Results of different models for completing missing cloud

[0110]

[0111] From the analysis of Table 1, it can be seen that compared with other models, the SwinPoinTr model proposed in this embodiment has improved the three indicators of chamfer distance, earth moving distance and F1 score in completing the incomplete Pleurotus eryngii completion task. This shows that the SwinPoinTr model can better learn the global and local features of the point cloud, not only can it more accurately complete the missing parts of the point cloud and improve the geometric accuracy, but also is superior to other models in terms of the completeness and accuracy of the completion, showing its stronger robustness and generalization ability.

[0112] The effects of different models on the point cloud completion of Pleurotus eryngii with different missing patterns are shown below. Figure 6 As shown in the figure. As can be seen from the figure, AdaPoinTr, SnowFlakeNet, and SwinPoinTr models have achieved excellent results in completing point clouds with different missing methods. Among them, the completion effect of the point cloud with top missing is relatively poor. The reason is that the caps of Pleurotus eryngii vary greatly. When the cap features are completely missing, it is difficult for the model to generate detailed features out of thin air. In comparison, the point cloud generated by SnowFlakeNet is more uniform and the surface is smoother, but many detailed features are also missing. The SwinPoinTr model can not only restore the point cloud better globally, but also improve the ability to restore detailed features.

[0113] Embodiment 2

[0114] Please refer to Figure 7 A point cloud completion device 1 based on SwinPoinTr includes a memory 3, a processor 2, and a computer program stored in the memory 3 and executable on the processor 2. When the processor 2 executes the computer program, the steps in the above-mentioned embodiment 1 are implemented.

[0115] Since the system / device described in the above embodiments of the present invention is a system / device used to implement the method of the above embodiments of the present invention, a person skilled in the art can understand the specific structure and deformation of the system / device based on the method described in the above embodiments of the present invention, and thus will not be described in detail here. All systems / devices used in the method of the above embodiments of the present invention belong to the scope of protection of the present invention.

[0116] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, devices or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0117] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (apparatus) and computer program products according to the embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions.

[0118] It should be noted that in the claims, any reference numerals placed between brackets shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention may be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In the claims enumerating several means, several of these means may be embodied by the same hardware. The use of the words first, second, third, etc., is for convenience of expression only and does not indicate any order. These words may be understood as part of the component name.

[0119] In addition, it should be noted that, in the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.

[0120] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments after knowing the basic creative concept. Therefore, the claims should be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.

[0121] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention should also include these modifications and variations.

Claims

1. A point cloud completion method based on SwinPoint, characterized in that: Includes steps: S1. Obtain the missing point cloud, use the farthest point sampling algorithm to obtain the center point from the missing point cloud to form a center point cloud, and perform regional point extraction based on the center point cloud to obtain a regional point set, and add the center point cloud and the regional point set to obtain a feature sequence set; S2, after inputting the feature sequence set into the hierarchical geometry-aware Transformer module for encoding and decoding, the decoded features and the predicted center point of the complete point cloud are obtained, and the hierarchical geometry-aware Transformer module calculates self-attention within the window and uses the shift window method to communicate between different windows; S3. Restore the point cloud of the local area where the predicted center point is located according to the decoded features and the predicted center point, and obtain a completed complete point cloud.

2. A point cloud completion method based on SwinPointer according to claim 1, characterized in that: The step S1 comprises the following steps: S11, obtaining a missing point cloud, and using a farthest point sampling algorithm to obtain a center point from the missing point cloud to form a center point cloud; S12. Use a k-nearest neighbor algorithm to construct an edge relationship for each center point and determine a local area to obtain a regional point set, and use a lightweight dynamic graph convolutional neural network with a hierarchical downsampling structure to extract local features from the regional point set to obtain a regional feature set; S13. After performing linear layer dimensionality increase on the global position of the center point coordinates, the global position is added to the local features of the corresponding center point in the regional feature set to obtain a feature sequence set.

3. The point cloud completion method based on SwinPointer according to claim 1, characterized in that: The step S1 further comprises the steps of: S14. Reshape the data shape of the feature sequence set to adapt to the window in the hierarchical geometry-aware Transformer module.

4. The point cloud completion method based on SwinPoint according to claim 3, characterized in that: The step S14 is specifically as follows: The window size is determined according to the number of center points, and then each feature sequence in the feature sequence set is divided into first feature blocks of equal length. Finally, the first feature blocks in the corresponding number of feature sequences are spliced ​​at corresponding positions in different dimensions according to the window size to obtain a multi-dimensional point cloud feature map.

5. A point cloud completion method based on SwinPointer according to any one of claims 1 to 4, characterized in that: In step S2, inputting the feature sequence set into the hierarchical geometry-aware Transformer module for encoding comprises the following steps: S21, the feature splitting layer splits the multidimensional point cloud feature map obtained by reshaping the feature sequence set into second feature blocks of corresponding sizes, and the second feature blocks split at corresponding positions are spliced ​​and flattened into first feature data; S22, the linear embedding layer in the first level projects the first feature data into a fixed dimension, and outputs it to a shifted geometry-aware Transformer block after normalization to obtain second feature data, and the shifted geometry-aware Transformer block performs geometry-aware attention calculation within the window, and then translates the window to perform geometry-aware attention calculation across windows; S23, the feature merging layer in the second level uses an image fusion algorithm to increase the number of channels of the feature map, and then outputs it to a shifted geometry-aware Transformer block for self-attention calculation, and then passes through the third and fourth levels with the same structure as the second level to obtain and output the encoded data.

6. The point cloud completion method based on SwinPoint according to claim 5, characterized in that: In step S2, the feature sequence set is input into the hierarchical geometry-aware Transformer module for decoding, and obtaining the decoded features and the predicted center point of the complete point cloud includes the following steps: S24, calculating a query vector according to the encoded data; S25, decoding is performed according to the query vector and the multi-dimensional point cloud feature map to obtain a decoding feature; S26. Input the decoded features into a linear layer for dimensionality reduction to obtain a predicted center point of the complete point cloud.

7. A point cloud completion method based on SwinPointer according to claims 1 to 3, characterized in that: The step S24 is specifically as follows: First, a linear layer is used to project the encoded data to a higher dimension to obtain high-dimensional data, and then a maximum pooling operation is performed on the high-dimensional data to obtain encoded features, and then an MLP layer is used to reconstruct the encoded features into a query vector.

8. A point cloud completion device based on SwinPoint, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the point cloud completion method based on SwinPoinTr described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Integrated point cloud denoising and analysis method based on point-level prompt

    CN121391650A