Three-dimensional point cloud completion method and device based on multi-level feature fusion
Through the multi-level feature fusion method, the farthest point sampling and multi-level feature extraction are used, combined with spatial transformation and attention mechanisms, the problem of poor point cloud completion quality in the existing technology is solved, and efficient and accurate point cloud completion effect is achieved.
Patent Information
- Application Number
- CN202510235755.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-11
AI Technical Summary
The method based on graph convolution in the prior art has insufficient in combining position feature extraction, resulting in poor point cloud completion quality, and the point cloud complemented by deep learning methods is relatively rough, and it is impossible to effectively restore the original point cloud features.
The multi-level feature fusion method is adopted to improve the accuracy and quality of point cloud completion through farthest point sampling, multi-level feature extraction, spatial transformation, continuous edge convolution and attention mechanisms, combined with the coding operation of the graph and the folded decoding operation.
The local feature extraction capability of point clouds is enhanced, local context and spatial connections between point clouds with different resolutions are established, local and global geometric structures of point clouds are captured, and efficient and accurate point cloud completion is achieved.
Smart Images

Figure CN120298635A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of point cloud completion, and in particular, to a three-dimensional point cloud completion method and device with multi-level feature fusion. Background Art
[0002] With the rapid development of computer technologies represented by deep neural networks, sensor devices such as lidar and depth cameras that can acquire three-dimensional information of scenes have been increasingly widely used in fields such as unmanned driving, industrial production, and virtual reality. Compared with 2D images, three-dimensional information can provide a more three-dimensional, comprehensive, and structured description of the real world.
[0003] However, due to the limited resolution of the devices and the problem of being easily occluded, the collected point clouds are sparse or incomplete, which greatly limits their application scope. Point cloud completion is to solve this problem and attempt to find or infer the complete three-dimensional shape from an incomplete three-dimensional object model. In recent years, the point cloud completion task directly starts from the input incomplete and sparse point clouds, usually uses a deep neural network model to learn the feature distribution of the point clouds, and then predicts the missing parts in the point clouds, aiming to construct a point cloud with high credibility, completeness, density, and uniformity.
[0004] However, although the method based on graph convolution in the prior art can capture local structures, it has deficiencies in extracting by combining position features, and the quality of completion by existing deep learning methods is not good, the completed point clouds are relatively rough, and the features of the original point clouds cannot be well restored. Summary of the Invention
[0005] The purpose of the present invention is to provide a three-dimensional point cloud completion method and device with multi-level feature fusion to improve the above problems. To achieve the above purpose, the technical solutions adopted by the present invention are as follows:
[0006] In a first aspect, the present application provides a three-dimensional point cloud completion method with multi-level feature fusion, including:
[0007] Obtain point cloud data;
[0008] Downsample the point cloud data by the farthest point sampling method to obtain point proxy centers;
[0009] Perform multi-level feature extraction on the point cloud data, and obtain point cloud features through spatial transformation, continuous edge convolution, and an attention mechanism. The point cloud features include point cloud local features and point cloud global features;
[0010] Embed the position of the point cloud features by passing the point proxy centers through a multi-layer perceptron to obtain point features with position information;
[0011] Perform feature encoding and feature decoding on the point features in sequence, and based on the graph-based encoding operation and the folding-based decoding operation, obtain the predicted point proxy and the predicted point proxy center;
[0012] Based on the point proxy center, the predicted point proxy center, and the predicted point proxy, perform point cloud reorganization to complete the point cloud data and obtain the predicted point cloud.
[0013] In a second aspect, the present application also provides a three-dimensional point cloud completion device with multi-level feature fusion, including:
[0014] An acquisition unit for acquiring point cloud data;
[0015] A sampling unit for downsampling the point cloud data by the farthest point sampling method to obtain the point proxy center;
[0016] A feature extraction unit for performing multi-level feature extraction on the point cloud data, and obtaining point cloud features through spatial transformation, continuous edge convolution, and an attention mechanism. The point cloud features include point cloud local features and point cloud global features;
[0017] A feature embedding unit for performing position embedding on the point cloud features after passing the point proxy center through a multi-layer perceptron to obtain point features with position information;
[0018] An encoding and decoding unit for performing feature encoding and feature decoding on the point features in sequence, and based on the graph-based encoding operation and the folding-based decoding operation, obtaining the predicted point proxy and the predicted point proxy center;
[0019] A point cloud completion unit for performing point cloud reorganization based on the point proxy center, the predicted point proxy center, and the predicted point proxy to complete the point cloud data and obtain the predicted point cloud.
[0020] The beneficial effects of the present invention are as follows: The present invention enhances the ability to extract local features of points through multi-level feature extraction, thereby improving the completion ability. At the same time, a channel and spatial attention mechanism is introduced to mix information in both the channel and spatial dimensions, thereby adjusting the feature representation, establishing local context and spatial connections between point clouds of different resolutions, and enhancing the semantic correlation between different blocks in the point cloud of the same resolution, so as to achieve the purpose of improving the quality of the completed point cloud. And a graph-based encoder and a folding-based decoder are proposed, which can effectively capture the local and global geometric structures of the point cloud, can flexibly aggregate neighborhood information, and at the same time generate accurate three-dimensional point clouds through folding operations, realizing efficient and accurate point cloud completion operations.
[0021] Other features and advantages of the present invention will be described in the following specification, and in part will be obvious from the specification, or can be understood by implementing the embodiments of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. Brief Description of the Drawings
[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0023] Figure 1 Schematic flow diagram of the three-dimensional point cloud completion method for multi-level feature fusion described in the embodiments of the present invention;
[0024] Figure 2 Schematic structural diagram of the completion network in the embodiments of the present invention;
[0025] Figure 3 Schematic structural diagram of the multi-level feature extraction module in the embodiments of the present invention;
[0026] Figure 4 Schematic structural diagram of the encoder and decoder in the embodiments of the present invention. Detailed Embodiments
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0028] It should be noted that: similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, terms such as "first" and "second" are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0029] Embodiment 1:
[0030] This embodiment provides a three-dimensional point cloud completion method with multi-level feature fusion.
[0031] See Figure 1 , the figure shows that this method includes step S100, step S200, step S300, step S400, step S500, and step S600.
[0032] Step S100: Obtain point cloud data;
[0033] In this embodiment, for the overall structure of the completion network constructed for three-dimensional point cloud completion is as Figure 2 shown, where FPS represents the farthest point sampling method.
[0034] Step S200: Downsample the point cloud data by the farthest point sampling method to obtain point proxy centers;
[0035] In this embodiment, since the number of points in the point cloud of each object is extremely large, selecting features for each point not only has a large computational amount but also is not conducive to the accuracy of point cloud completion. By selecting new points far from the already selected points, it can ensure that the sampled points are more evenly distributed over the entire point cloud rather than being concentrated in a certain area. Therefore, the farthest point sampling method is used to sample the defective point cloud, and the points obtained by downsampling are used as the point proxy centers of the defective point cloud.
[0036] Step S300: Perform multi-level feature extraction on the point cloud data, and obtain point cloud features through spatial transformation, continuous edge convolution, and attention mechanism. The point cloud features include point cloud local features and point cloud global features;
[0037] The step S300 includes:
[0038] Step S301: Construct a multi-level feature extraction module based on spatial transformation and continuous edge convolution, and construct a channel and spatial attention mechanism module based on the attention mechanism;
[0039] Step S302: Extract features from the point cloud data through the multi-level feature extraction module to obtain preliminary point cloud features;
[0040] In this embodiment, the structure of the multi-level feature extraction module is as Figure 3 shown, where n represents the number of points.
[0041] The step S302 includes:
[0042] Step A100: Sequentially pass the point cloud data through four edge convolutions for feature extraction to obtain the first point cloud feature, the second point cloud feature, the third point cloud feature, and the fourth point cloud feature;
[0043] Step A200: Map the fourth point cloud feature to a high dimension through a multi-layer perceptron to obtain a fifth point cloud feature;
[0044] Step A300: Concatenate the first point cloud feature, the second point cloud feature, the third point cloud feature, the fourth point cloud feature, and the fifth point cloud feature to obtain a first concatenated feature;
[0045] Step A400: Perform a pyramid pooling operation on the first concatenated feature to convert the feature dimension and obtain a preliminary point cloud feature.
[0046] In this embodiment, after spatial transformation and continuous edge convolutions, multiple features are obtained. Then, the features obtained by each edge convolution are concatenated, and a pyramid pooling operation is performed to convert the features into the dimensions required for the next stage.
[0047] Step S303: Pass the preliminary point cloud feature through a channel and spatial attention mechanism module to dynamically adjust the channel weight and spatial position weight to obtain a point cloud feature.
[0048] In this embodiment, the channel and spatial attention mechanism module includes a channel attention mechanism and a spatial attention mechanism. Each channel of a feature represents a specific feature information, and the contribution degrees of different channels to the final task are often different. The channel attention mechanism dynamically adjusts the weight of each channel, enabling the completion network to focus on those features that are more critical to the task, thereby improving the overall performance.
[0049] The calculation formula of the channel attention mechanism is:
[0050] M c (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F)))
[0051] In the formula, M c (F) represents the output feature when the channel attention mechanism inputs F, F represents the input feature of the channel attention mechanism, σ(·) represents the Sigmoid activation function, MLP(·) represents the multi-layer perceptron, AvgPool(·) represents the global average pooling operation, and MaxPool(·) represents the max pooling operation.
[0052] The spatial attention mechanism analyzes the spatial layout of the input feature, calculates the attention weights of different spatial positions, and thus adjusts the feature representation to make the completion network pay more attention to the spatial regions helpful for the task.
[0053] The calculation formula of the spatial attention mechanism is:
[0054] M s(F′) = σ(f 7×7 ([AvgPool(F′) + MaxPool(F′)]))
[0055] where M s (F′) represents the output feature when the spatial attention mechanism inputs F′, and F ′ represents the input feature of the spatial attention mechanism. σ(·) represents the Sigmoid activation function, AvgPool(·) represents the global average pooling operation, MaxPool(·) represents the max pooling operation, and f 7×7 (·) represents a convolutional kernel of size 7×7.
[0056] In this embodiment, the channel and spatial attention mechanism module combines the spatial and channel attention mechanisms, realizing the double refinement of the input features, improving the representation ability and decision-making accuracy of the completion network, and being able to dynamically adjust the importance of each channel and spatial position according to the upper and lower features of the input point cloud. This adaptive recalibration mechanism helps the completion network better focus on the features of the completed point cloud and improves the accuracy of point cloud completion.
[0057] The calculation formula of the channel and spatial attention mechanism module is:
[0058]
[0059] where F ′ represents the input feature of the spatial attention mechanism, F represents the input feature of the channel attention mechanism, M c (F) represents the output feature when the input feature of the channel attention mechanism is F, and M s (F′) represents the output feature when the spatial attention mechanism inputs F′, represents the multiplication operation, and F″ represents the output feature of the channel and spatial attention mechanism module.
[0060] Step S400: Pass the point proxy center through a multi-layer perceptron and then perform position embedding on the point cloud features to obtain point features with position information;
[0061] Step S500: Perform feature encoding and feature decoding on the point features in sequence, based on graph-based encoding operations and folding-based decoding operations, to obtain the predicted point proxy and the predicted point proxy center;
[0062] The step S500 includes:
[0063] Step S501: Construct an encoder based on the KNN algorithm, construct a decoder based on three consecutive layers of perceptrons, and connect the encoder and the decoder through a Query generator;
[0064] Step S502: Reconstruct the positions of the point features through the encoder to obtain encoded features;
[0065] In this embodiment, as Figure 4 shown, it is the structural diagram of the encoder and the decoder. Among them, m represents the number of replication times, and n represents the number of points.
[0066] The step S502 includes:
[0067] Step B100: Construct a K-nearest neighbor graph for each point in the point features through the KNN algorithm;
[0068] In this embodiment, when using the KNN algorithm to construct the K-nearest neighbor graph, take K = 16, and K represents the range size of the neighborhood.
[0069] Step B200: Calculate the local covariance matrix according to the three-dimensional positions of the points and their neighbor points in the K-nearest neighbor graph;
[0070] In this embodiment, according to the three-dimensional positions of the points and their neighbor points in the K-nearest neighbor graph, calculate the local covariance matrix of size 3×3 in sequence, and convert it into a local covariance matrix of size 1×9.
[0071] Step B300: Obtain the point position matrix of the points in the point features, and splice the point position matrix and the local covariance matrix to obtain a spliced matrix;
[0072] In this embodiment, use the multi-head self-attention mechanism to calculate the point position matrix of the points in the point features, and then splice the point position matrix and the local covariance matrix to ensure that while mining the long-range semantic correlation, the original local geometric relationship of the point cloud is retained, effectively improving the performance of the completion network.
[0073] Step B400: Perform a point-by-point function calculation on the spliced matrix through a three-layer perceptron to obtain the sixth point cloud feature;
[0074] In this embodiment, the three-layer perceptron acts on each row of the spliced matrix in parallel and performs a point-by-point function calculation on each three-dimensional point.
[0075] Step B500: Pass the sixth point cloud feature through two consecutive layers, and perform a max pooling operation on the neighborhood features of each point to obtain the seventh point cloud feature;
[0076] In this embodiment, the calculation formula of the layer is:
[0077] Y = A max (X)K′
[0078]
[0079] Wherein, Y represents the output feature of the layer, X represents the input feature of the layer, and A max (·) represents the aggregated feature, K' represents the feature mapping matrix, and A max (X) ij represents the element value corresponding to the j-th feature of the i-th point in the aggregated feature, ReLU(·) represents the activation function, represents performing local max pooling operation on x kj where x kj represents the j-th feature value of the k-th point in the input feature of the layer, and Ν(i) represents the set of neighbor nodes of the i-th point.
[0080] Among them, performing local max pooling operation is essentially calculating a local signature based on the graph structure. This signature can represent the aggregated topological information of neighbor nodes. Through the connection of the graph-based max pooling layer, the topological information is propagated to a larger area.
[0081] Step B600: Perform point-by-point function calculation on the seventh point cloud feature through a three-layer perceptron to obtain the encoded feature.
[0082] In this embodiment, feature encoding is performed through an encoder to reconstruct the position of points to obtain the encoded feature, that is, Figure 4 the codeword in, and the number of reconstructed points does not necessarily be the same as the number of points in the point feature. At the same time, a reconstruction error function of the encoder is defined, and the specific formula is:
[0083]
[0084] Wherein, represents the chamfer distance metric, S represents the input point set of the encoder, represents the reconstructed point set, MAX{·} represents taking the maximum value, ∥·∥2 represents the Euclidean distance, |·| represents taking the absolute value, x represents the point in S, represents the point in, and min(·) represents taking the minimum value.
[0085] The chamfer distance metric measures the difference between two point sets by calculating the distance from each point to its nearest neighbor point and taking the maximum value. In this embodiment, the reconstruction error function is used as the loss function of the encoder.
[0086] Step S503: Generate a query vector for the encoded feature through the Query generator;
[0087] Step S504: Input the encoded feature and the query vector into the decoder for feature decoding to obtain the predicted point proxy and the predicted point proxy center.
[0088] The step S504 includes:
[0089] Step C100: Copy the encoded feature a preset number of times to obtain a copied encoded feature;
[0090] Step C200: Concatenate the copied encoded feature and the query vector to obtain a second concatenated feature;
[0091] Step C300: Fold the second concatenated feature through a three-layer perceptron to obtain a first folded feature;
[0092] Step C400: Concatenate the copied encoded feature and the first folded feature to obtain a second concatenated feature;
[0093] Step C500: Fold the second concatenated feature through a three-layer perceptron to obtain a predicted point proxy and a predicted point proxy center.
[0094] In this embodiment, as Figure 4 shown, the folding-based decoder is composed of two consecutive three-layer perceptrons connected together, folding and distorting a two-dimensional grid of a specific size according to the shape of the decoder input feature. Before inputting the codeword into the decoder, it needs to be copied m times to form a copied encoded feature of m×512. In addition, a square surface centered at the origin is selected from the encoded feature by the Query Generator, and the coordinates of m grid points on this square surface can be represented as grid point coordinates of m×2, that is, the query vector. At the same time, the number of copies, that is, the number of grid points selected on the square surface, is determined by the number of points, and its size is the square number closest to the number of points.
[0095] Concatenate the copied encoded feature and the query vector, so that the decoder can accurately reconstruct the point cloud according to the overall structural information of the point cloud. Among them, the query vector is translated into a predicted point proxy after passing through the decoder, and the corresponding predicted point proxy center is obtained through the predicted point proxy.
[0096] Step S600: Reconstruct the point cloud based on the point proxy center, the predicted point proxy center, and the predicted point proxy, and complete the point cloud data to obtain a predicted point cloud.
[0097] In this embodiment, when performing point cloud reconstruction, the FoldingNet network is used to reconstruct the offset coordinates of the predicted point proxy and its corresponding predicted point proxy center, and concatenate them with the original point proxy center to obtain a predicted point cloud. Among them, the formula for point cloud reconstruction is;
[0098] F O = concat(F I , c + FoldingNet(F D ))
[0099] In the formula, F O represents the predicted point cloud, concat(·) represents feature concatenation, and F I represents the point proxy center, c represents the predicted point proxy center, and F D represents the predicted point proxy, and FoldingNet(·) represents the FoldingNet network.
[0100] In summary, the present invention splices the feature information output by each edge convolution through a multi-level feature extraction module, and through pyramid pooling operation, obtains local features and global features that fuse multi-level edge convolution feature information as the output features of the multi-level feature extraction module. The output features of the multi-level feature extraction module contain both low-level features rich in small-range local information and high-level features that can accurately express local overall information, thereby obtaining more detailed point cloud features.
[0101] At the same time, a spatial attention and channel attention mechanism is added to the point cloud feature representation. By using the spatial information and channel image in the features, the completion network can more efficiently calculate the features between each point or region and other points or regions by calculating weights, establish local context and spatial connections between point clouds of different resolutions, and enhance the semantic correlation between different blocks in the point cloud of the same resolution, so that the completion network can better focus on the key point cloud spatial position information, achieving the purpose of improving the quality of the completed point cloud.
[0102] At the same time, an encoder based on K-nearest neighbor geometry is designed to splice the features calculated by the multi-head self-attention mechanism and the local covariance matrix calculated based on the KNN algorithm to obtain feature information with local geometric perception. While ensuring that the completion network mines long-range semantic correlations, the original local geometric relationship of the point cloud is retained, effectively improving the performance of the completion network.
[0103] Embodiment 2:
[0104] This embodiment provides a three-dimensional point cloud completion device with multi-level feature fusion, and the device includes:
[0105] An acquisition unit for acquiring point cloud data;
[0106] A sampling unit for downsampling the point cloud data by the farthest point sampling method to obtain a point proxy center;
[0107] A feature extraction unit for performing multi-level feature extraction on the point cloud data, and obtaining point cloud features through spatial transformation, continuous edge convolution, and an attention mechanism, where the point cloud features include point cloud local features and point cloud global features;
[0108] A feature embedding unit, configured to perform position embedding on the point cloud features after passing the point proxy center through a multi-layer perceptron, so as to obtain point features with position information;
[0109] An encoding and decoding unit, configured to perform feature encoding and feature decoding on the point features in sequence, and obtain a predicted point proxy and a predicted point proxy center based on graph-based encoding operations and folding-based decoding operations;
[0110] A point cloud completion unit, configured to perform point cloud recombination based on the point proxy center, the predicted point proxy center, and the predicted point proxy, and complete the point cloud data to obtain a predicted point cloud.
[0111] The feature extraction unit includes:
[0112] A first construction subunit, configured to construct a multi-level feature extraction module based on spatial transformation and continuous edge convolution, and construct a channel and spatial attention mechanism module based on an attention mechanism;
[0113] A feature extraction subunit, configured to extract features from the point cloud data through the multi-level feature extraction module to obtain preliminary point cloud features;
[0114] An adjustment subunit, configured to pass the preliminary point cloud features through the channel and spatial attention mechanism module to dynamically adjust the channel weight and spatial position weight, so as to obtain point cloud features.
[0115] The encoding and decoding unit includes:
[0116] A second construction subunit, configured to construct an encoder based on the KNN algorithm, construct a decoder based on three consecutive layers of perceptrons, and connect the encoder and the decoder through a Query generator;
[0117] An encoding subunit, configured to perform position reconstruction on the point features through the encoder to obtain encoded features;
[0118] A vector generation subunit, configured to generate a query vector of the encoded features through the Query generator;
[0119] A decoding subunit, configured to input the encoded features and the query vector into the decoder for feature decoding to obtain a predicted point proxy and a predicted point proxy center.
[0120] The decoding subunit includes:
[0121] A replication subunit, configured to replicate the encoded features a preset number of times to obtain replicated encoded features;
[0122] A first splicing subunit, configured to splice the replicated encoded features and the query vector to obtain a second spliced feature;
[0123] The first folding subunit is configured to fold the second splicing feature through a three-layer perceptron to obtain a first folded feature;
[0124] The second splicing subunit is configured to splice the replicated encoded feature and the first folded feature to obtain a second splicing feature;
[0125] The second folding subunit is configured to fold the second splicing feature through a three-layer perceptron to obtain a predicted point proxy and a predicted point proxy center.
[0126] It should be noted that regarding the device in the above embodiments, the specific manners in which each unit performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0127] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0128] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present invention, and all should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A three-dimensional point cloud completion method based on multi-level feature fusion, characterized in that Including: Obtain point cloud data; Downsample the point cloud data by the farthest point sampling method to obtain point proxy centers; Perform multi-level feature extraction on the point cloud data, and obtain point cloud features through spatial transformation, continuous edge convolution, and attention mechanism. The point cloud features include point cloud local features and point cloud global features; Pass the point proxy centers through a multi-layer perceptron and then perform position embedding on the point cloud features to obtain point features with position information; Perform feature encoding and feature decoding on the point features in sequence, based on graph-based encoding operations and folding-based decoding operations, to obtain predicted point proxies and predicted point proxy centers; Perform point cloud reconstruction based on the point proxy centers, the predicted point proxy centers, and the predicted point proxies, and complete the point cloud data to obtain predicted point clouds.
2. The three-dimensional point cloud completion method with multi-level feature fusion according to claim 1, characterized in that ,Perform multi-level feature extraction on the point cloud data, and obtain point cloud features through spatial transformation, continuous edge convolution, and attention mechanism, including: Construct a multi-level feature extraction module based on spatial transformation and continuous edge convolution, and construct a channel and spatial attention mechanism module based on the attention mechanism; Perform feature extraction on the point cloud data through the multi-level feature extraction module to obtain preliminary point cloud features; Pass the preliminary point cloud features through the channel and spatial attention mechanism module to dynamically adjust the channel weight and spatial position weight to obtain point cloud features.
3. The three-dimensional point cloud completion method with multi-level feature fusion according to claim 2, wherein ,The performing feature extraction on the point cloud data through the multi-level feature extraction module to obtain preliminary point cloud features includes: Pass the point cloud data through four edge convolutions in sequence for feature extraction to obtain the first point cloud feature, the second point cloud feature, the third point cloud feature, and the fourth point cloud feature; Map the fourth point cloud feature to a high dimension through a multi-layer perceptron to obtain the fifth point cloud feature; Perform feature splicing on the first point cloud feature, the second point cloud feature, the third point cloud feature, the fourth point cloud feature, and the fifth point cloud feature to obtain the first splicing feature; Perform a pyramid pooling operation on the first splicing feature to convert the feature dimension to obtain preliminary point cloud features.
4. The three-dimensional point cloud completion method with multi-level feature fusion according to claim 1, wherein ,Performing feature encoding and feature decoding on the point features in sequence, based on graph-based encoding operations and folding-based decoding operations, to obtain predicted point proxies and predicted point proxy centers, including: Construct an encoder based on the KNN algorithm, construct a decoder based on three consecutive multi-layer perceptrons, and connect the encoder and the decoder through a Query generator; Perform position reconstruction on the point features through the encoder to obtain encoded features; Generate query vectors of the encoded features through the Query generator; Input the encoded features and the query vectors into the decoder for feature decoding to obtain predicted point proxies and predicted point proxy centers.
5. The three-dimensional point cloud completion method with multi-level feature fusion according to claim 4, wherein ,Performing position reconstruction on the point features through the encoder to obtain encoded features, including: Construct a K-nearest neighbor graph for each point in the point features through the KNN algorithm; Calculate the local covariance matrix according to the three-dimensional positions of the points and neighbor points in the K-nearest neighbor graph; Obtain the point position matrix of the midpoints of the point features, and splice the point position matrix and the local covariance matrix to obtain a spliced matrix; Perform point-wise function calculation on the spliced matrix through a three-layer perceptron to obtain the sixth point cloud feature; Pass the sixth point cloud feature through two consecutive layers, and perform max pooling operations on the neighborhood features of each point to obtain the seventh point cloud feature; Perform point-wise function calculation on the seventh point cloud feature through a three-layer perceptron to obtain the encoded feature.
6. The three-dimensional point cloud completion method with multi-level feature fusion according to claim 4, characterized in that , Input the encoded feature and the query vector into the decoder for feature decoding to obtain the predicted point proxy and the predicted point proxy center, including: Copy the encoded feature a preset number of times to obtain a copied encoded feature; Splice the copied encoded feature and the query vector to obtain a second spliced feature; Fold the second spliced feature through a three-layer perceptron to obtain a first folded feature; Splice the copied encoded feature and the first folded feature to obtain a second spliced feature; Fold the second spliced feature through a three-layer perceptron to obtain the predicted point proxy and the predicted point proxy center.
7. A three-dimensional point cloud completion device with multi-level feature fusion, characterized in that Including: An acquisition unit for acquiring point cloud data; A sampling unit for downsampling the point cloud data by the farthest point sampling method to obtain a point proxy center; A feature extraction unit for performing multi-level feature extraction on the point cloud data, and obtaining point cloud features through spatial transformation, continuous edge convolution, and an attention mechanism, where the point cloud features include point cloud local features and point cloud global features; A feature embedding unit for performing position embedding on the point cloud features after passing the point proxy center through a multi-layer perceptron to obtain point features with position information; An encoding and decoding unit for sequentially performing feature encoding and feature decoding on the point features, and obtaining the predicted point proxy and the predicted point proxy center based on graph-based encoding operations and folding-based decoding operations; A point cloud completion unit for performing point cloud recombination based on the point proxy center, the predicted point proxy center, and the predicted point proxy, and completing the point cloud data to obtain a predicted point cloud.
8. The three-dimensional point cloud completion device with multi-level feature fusion according to claim 7, wherein, The feature extraction unit includes: A first construction subunit for constructing a multi-level feature extraction module based on spatial transformation and continuous edge convolution, and constructing a channel and spatial attention mechanism module based on the attention mechanism; A feature extraction subunit for extracting features from the point cloud data through the multi-level feature extraction module to obtain preliminary point cloud features; An adjustment subunit for passing the preliminary point cloud features through the channel and spatial attention mechanism module to dynamically adjust the channel weights and spatial position weights to obtain point cloud features.
9. The three-dimensional point cloud completion device with multi-level feature fusion according to claim 7, characterized in that, The encoding and decoding unit includes: A second construction subunit for constructing an encoder based on the KNN algorithm, constructing a decoder based on three consecutive layers of perceptrons, and connecting the encoder and the decoder through a Query generator; An encoding subunit for performing position reconstruction on the point features through the encoder to obtain an encoded feature; A vector generation subunit for generating a query vector of the encoded feature through the Query generator; A decoding subunit, configured to input the encoded feature and the query vector into the decoder for feature decoding to obtain a predicted point proxy and a predicted point proxy center.
10. The three-dimensional point cloud completion device with multi-level feature fusion according to claim 9, characterized in that The decoding subunit includes: A replication subunit, configured to replicate the encoded feature a preset number of times to obtain replicated encoded features; A first splicing subunit, configured to splice the replicated encoded features and the query vector to obtain a second spliced feature; A first folding subunit, configured to fold the second spliced feature through a three-layer perceptron to obtain a first folded feature; A second splicing subunit, configured to splice the replicated encoded features and the first folded feature to obtain a second spliced feature; A second folding subunit, configured to fold the second spliced feature through a three-layer perceptron to obtain a predicted point proxy and a predicted point proxy center.