Point cloud completion method based on double-branch feature extraction and attention mechanism
Through the dual-branch feature extractor and attention mechanism, combined with global and local features, the problem of the existing technology being difficult to capture point cloud features at the same time is solved, and a higher-quality point cloud completion effect is achieved.
Patent Information
- Application Number
- CN202510189962.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
AI Technical Summary
The existing point cloud completion technology is difficult to capture the global and local features of point clouds at the same time, resulting in the missing or abnormal points in the detailed parts of the generated point clouds.
A dual-branch feature extractor is adopted, combining global and local features, and introducing channel attention mechanisms and spatial attention mechanisms to establish a more comprehensive and accurate point cloud completion model.
By combining global and local features and using attention mechanisms, the generated point cloud completion results have fewer anomalies and noise, significantly improving the integrity and quality of point cloud data.
Smart Images

Figure CN120125472A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of 3D modeling, and specifically relates to a point cloud completion method based on dual-branch feature extraction and attention mechanism. Background Art
[0002] In recent years, with the rapid development of computer graphics, computer vision, and deep learning technologies, 3D point cloud technology has become increasingly mature. Point cloud is an important concept in the fields of computer graphics and computer vision, and it is not only a key research direction in academic research, but also has important applications in real life, such as unmanned driving, 3D mapping, and architectural design.
[0003] 3D point cloud completion is one of the basic tasks of point cloud. The goal of the point cloud completion task is to restore an incomplete and defective point cloud to a complete point cloud with fine-grained information. Its theoretical significance lies in providing key support for the integrity, accuracy, and robustness of point cloud data.
[0004] By restoring the complete 3D shape from partially missing point cloud data, point cloud completion technology can not only improve the quality of data, but also provide a more reliable input for subsequent point cloud analysis tasks.
[0005] Point cloud completion is an important foundation of 3D point cloud technology, and its practical significance lies in its wide application value, especially in modern science and technology and industrial fields, it plays an important promoting role. With the popularization of 3D sensors such as lidar (LiDAR) and depth cameras, point cloud data has become an important way to describe the 3D structure of objects and scenes. However, due to perspective limitations, sensor errors, resolution limitations, occlusion, and environmental complexity, the obtained point clouds are often incomplete. Point cloud completion technology is the key to solving this problem, which can effectively restore and fill in the missing point cloud data, thereby improving the integrity and quality of point cloud data.
[0006] Existing methods such as PCN mainly focus on global features, which results in missing details in the generated point clouds; while methods like PointAttN focus on local feature extraction, and this way is not sufficient to comprehensively capture all features of the point cloud, which may lead to a lack of global control in the generated point clouds and thus the appearance of abnormal points. Summary of the Invention
[0007] Aiming at the deficiencies of several existing point cloud completion technologies, the present invention focuses on combining global features and local features, and at the same time introduces channel attention mechanism and spatial attention mechanism, and provides a 3D point cloud completion method based on a dual-branch feature extractor and attention mechanism.
[0008] The first aspect of the present invention provides a 3D point cloud completion method based on a dual-branch feature extractor and attention mechanism, and the main steps are as follows:
[0009] Step 1: Prepare a 3D point cloud data set with arbitrary size and quantity, unify the number of points for the defective point cloud, and preprocess all point cloud data.
[0010] Step 2: Design a 3D point cloud completion model. Using the defective point cloud as the input, generate a complete and dense point cloud model while possessing the geometric details of the defective point cloud.
[0011] Step 3: Train the 3D point cloud completion model.
[0012] Step 4: Repeatedly execute Step 3 until a preset number of iterations is reached. Traverse all 3D point clouds in each round. At the end of each round of iteration, update and save the model parameters.
[0013] Step 5: Select the model parameters with the best metrics, load them into the model; input the defective point cloud to generate the corresponding complete point cloud.
[0014] The second aspect of the present invention provides an electronic device for 3D point cloud completion based on dual-branch feature extraction and attention mechanism, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, when the processor executes the program, it implements the above-mentioned 3D point cloud completion method based on dual-branch feature extraction and attention mechanism.
[0015] The third aspect of the present invention provides a computer-readable storage medium, and the storage medium stores a computer program, and the computer program is used to execute the above-mentioned 3D point cloud completion method based on dual-branch feature extraction and attention mechanism.
[0016] The beneficial effects of the present invention:
[0017] The present invention ingeniously combines the local features and global features of the defective point cloud, fully utilizes the information of the feature matrix in the channel dimension and spatial dimension, enables the model to more comprehensively establish the mapping relationship from the defective point cloud to the complete point cloud, and this design makes the completion result have fewer outliers and noises, showing a good completion effect on the 3D point cloud model data set. Description of the Drawings
[0018] Figure 1 It is the structure diagram of the 3D point cloud completion model in the embodiment of the present application;
[0019] Figure 2 It is the structure diagram of the dual-branch feature extraction module in the embodiment of the present application;
[0020] Figure 3 It is the schematic diagram of introducing the channel attention mechanism and spatial attention mechanism in the embodiment of the present application (hereinafter referred to as the CSA module);
[0021] Figure 4 Introduce a sparse point cloud generator with a CSA module for the embodiments of this application;
[0022] Figure 5 Introduce a dense point cloud generator with a CSA module for the embodiments of this application;
[0023] Figure 6 The comparison chart of experimental results for the embodiments of this application;
[0024] Figure 7 The detailed comparison chart for the embodiments of this application. Detailed implementation manners
[0025] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0026] As Figures 1 - 5 shown, the embodiments of this application provide a three-dimensional point cloud completion method based on dual-branch feature extraction and attention mechanism, specifically including:
[0027] Step 1: Prepare a three-dimensional model data set containing any size and quantity, in the format of.h5 or.pcd. Unify the number of points for the defective point cloud (the number of points in the defective point cloud is 2048), and perform necessary preprocessing on all point cloud data.
[0028] Step 2: As Figure 1 shown, design a three-dimensional point cloud completion model, which adopts an encoder-decoder architecture. The encoder takes the defective point cloud as input and obtains the shape encoding of the defective point cloud; the decoder is used to generate a complete point cloud from the shape encoding.
[0029] As Figure 2 shown, the encoder is specifically a dual-branch feature extraction module, which includes two branches: a global feature extraction module and a local feature extraction module. The geometric feature perception module (GDP) and the self-feature enhancement module (SFA) are derived from the baseline model. Add a CSA module to the existing feature extraction module of the baseline model as a branch of feature extraction to obtain local features. At the same time, add a global feature extraction module to obtain global features. The structure of the global feature extraction module is as Figure 2 shown in the global feature extraction branch.
[0030] Furthermore, as Figure 3As shown, the CSA module is composed of a channel attention mechanism and a spatial attention mechanism stacked together. The input of the channel attention module is a feature matrix F with the shape (N, C), that is, F is a feature matrix with N rows and C columns. Max-pooling and average-pooling operations are respectively performed on the channel dimension of F to obtain F p.Max and F p.Avg , that is, F p.Max is a 1-row C-column feature matrix obtained by performing max-pooling on F in the channel dimension, and F p.Avg is a 1-row C-column feature matrix obtained by performing average-pooling on F in the channel dimension.
[0031] In particular, different from the traditional channel attention mechanism, in the embodiments of this application, F p.Max and F p.Avg are respectively fed into a multi-layer perceptron to obtain and with the shape (N, C). Then and are added together, and after being activated by a non-linear function (Sigmoid), the channel attention weight W c is obtained. The feature matrix F and the channel attention weight W c are subjected to Hadamard product operation to obtain the feature matrix based on channel attention
[0032] The decoder includes a sparse point cloud generator and a dense point cloud generator. As Figure 4 and Figure 5 shown, the sparse point cloud generator takes the defective point cloud and the shape encoding as inputs, and generates a sparse but relatively complete point cloud model, denoted as the rough point cloud. To ensure that the dense point cloud has the geometric details of the defective point cloud, the dense point cloud generator takes the defective point cloud shape encoding and the rough point cloud as inputs, and generates a complete and dense point cloud model, which also has the geometric details of the defective point cloud.
[0033] As Figure 4 shown, the sparse point cloud generator is similar to the seed generator framework of the baseline model, and the sampling rate of the SFA module is kept at 1. To utilize the channel information and spatial position information in the feature matrix, a CSA module is added behind the first two multi-layer perceptrons and the last self-feature enhancement module. Then the generated point cloud and the defective point cloud are merged, and finally farthest point sampling (FPS) is performed to retain 512 points as the rough point cloud.
[0034] As Figure 5As shown, the dense point cloud generator is similar to the sparse point cloud generator and is mainly composed of an SFA module and a CSA module. First, the rough point cloud and the shape encoding are respectively fed into a multi-layer perceptron and the CSA module to obtain the feature matrix F 1 and the shape encoding F 2 . The shape encoding is copied to make its shape consistent with that of F 1 . Then, F 1 and F 2 are concatenated to obtain a new feature matrix F 1 with the shape of (n, 256). Next, F 1 is fed into three consecutive SFA modules and upsampled. The shape of the upsampled feature matrix is adjusted to (N, 128), and then F 1 is extended to (N, 128), and they are concatenated to obtain a new feature matrix F with the shape of (N, 256). Finally, F is fed into a multi-layer perceptron to generate the final complete point cloud. The SFA sampling rate in the dense point cloud generator satisfies the following expression:
[0035] N = 2 × n × μ 1 × μ 2 × μ 3
[0036] where μ1, μ2, and μ3 are the upsampling rates of the three SFA modules respectively, N is the number of points in the generated target point cloud, and n is the number of points in the rough point cloud.
[0037] Step 3: Train the three-dimensional point cloud completion model.
[0038] Step 3-1: Assign the corresponding shape encoding to the input defective point cloud.
[0039] Step 3-11: Use the defective point cloud X as the input of the global feature extraction module. After passing through a multi-layer perceptron, the feature matrix M of the defective point cloud is obtained, and its shape is (2048, 512).
[0040] Step 3-12: Perform a pooling operation on this feature matrix to obtain the shape encoding F 1 with the shape of (1, 512), that is, the global feature.
[0041] In particular, in order to retain the significant features of the defective point cloud to the greatest extent and suppress noise, the maximum pooling operation is used in the pooling layer here.
[0042] After obtaining the global feature F 1 , it is also necessary to obtain the local feature F 2 with the shape of (1, 512). The local feature extraction branch is similar to the feature extraction module of the baseline module, and the structure is as shown in the local feature extraction branch in Figure 2 .
[0043] Step 3-13: Obtain the global feature F 1 and the local feature F 2 After that, splice the global feature F 1 and the local feature F 2 to obtain an encoding with a shape of (1, 1024). Then use a multi-layer perceptron to map this shape encoding to a shape encoding F with a shape of (1, 512), and use this encoding F that integrates the global feature and the local feature as the final shape encoding.
[0044] Step 3-2: Use a decoding operation similar to the baseline model and adopt a Coarse-to-Dense strategy to generate a complete point cloud. Input the shape encoding F of the defective point cloud into a sparse point cloud generator similar to the baseline model to generate a rough but complete point cloud p 1 , and the number of points in the rough point cloud is 512.
[0045] Step 3-3: Input the rough point cloud p 1 and F into the first dense point cloud generator to generate the dense point cloud p 2 in the first stage.
[0046] Step 3-4: Input p 2 and F into the second dense point cloud generator to generate the final dense point cloud p 3 . Different from the baseline model, in this embodiment of the present application, a CSA module is added behind most of the multi-layer perceptrons in the sparse point cloud generator and the dense point cloud generator, and the structure of the CSA module is as Figure 3 shown.
[0047] Step 3-5: Calculate the loss using the Chamfer Distance (CD). The expression of the CD distance is as follows:
[0048]
[0049] where S and P are two point clouds to be compared, and s and p are the points in S and P respectively.
[0050] The expression of the complete loss function is as follows:
[0051]
[0052] where λ i is a coefficient of the CD distance calculation result in the i-th stage, and p i , s i are the point cloud predicted and generated in the i-th stage and the point cloud obtained by downsampling the real point cloud respectively.
[0053] Step 4: Repeatedly execute Step 3 until the preset number of iterations (default is 400 rounds) is reached. In each round, traverse all 3D point clouds. At the end of each round of training, update and save the model parameters.
[0054] Step 5: Select the model parameter with the best CD metric and load it into the model; input the defective point cloud to generate the corresponding complete point cloud.
[0055] Specific experimental data and parameters:
[0056] Specific comparison results with the baseline model are as Figure 6 shown. The first row represents the input defective point cloud, the second row is the completion result of the baseline model, the third row is the completion result of this application, and the fourth row is the complete point cloud model. Specific experimental details are as follows:
[0057] (1) The training data includes the Completion3D dataset and the PCN dataset, which are obtained by processing the ShapeNet dataset. Completion3D dataset: The Completion3D dataset is usually used for point cloud completion with 2048 defective point cloud points and 2048 complete point cloud points. The dataset contains 30,958 point cloud models, 8 data categories, and 3 data subsets, and the data format is.h5. The 3 data subsets are the training set (train), the test set (test), and the validation set (val). Each model in the training set and the validation set contains a defective point cloud model and a complete point cloud model, and only defective point cloud models are in the test set. The number of points in the defective model and the complete model is 2048. PCN dataset: The PCN dataset is usually used for point cloud completion with 2048 defective point cloud points and 16,384 complete point cloud points. The dataset contains 30,974 point cloud models, 8 data categories, and 3 data subsets, and the data format is.pcd. The 3 data subsets are the training set (train), the test set (test), and the validation set (val), and each subset contains a defective point cloud model and a complete point cloud model. The number of points in the defective model is 2048, and the number of points in the complete model is 16,384.
[0058] (2) The experimental environment for training and testing the Completion3D dataset is a 12th Gen Intel(R) Core(TM) i9-12900K CPU, 64GB of memory, and a 3080 GPU with 12GB of video memory. The experimental environment for training and testing the PCN dataset is a 13th Gen Intel(R) Core(TM) i7-13700K CPU, 32GB of memory, and a 4090 GPU with 24GB of video memory. The Adam optimizer is used, and a total of 400 rounds of training are performed; the initial learning rate is 0.0001, and every 40 rounds, the learning rate decays to 0.7 times the previous value. When the learning rate decays to 0.000001, it no longer decays. When using the Completion3D dataset, the batch size is set to 32; when using the PCN dataset, the batch size is set to 8. The training time for the Completion3D dataset is approximately 1 day, and the training time for the PCN dataset is approximately 4 days.
[0059] (3) This application uses the training parameters with the minimum loss for the completion test. It can be seen from the visualization results that the completion results of the present invention on the PCN dataset are more complete and reasonable than those of the baseline model.
[0060] In addition, Figure 7 The completion details of this application are shown. The point cloud models predicted by different networks are compared with the complete point cloud model. Among them, the green is the complete point cloud, the orange is the predicted point cloud of the present invention, and the blue is the predicted point cloud of the baseline model. Experiments show that the complete model predicted by this application based on the incomplete point cloud has fewer outliers and noises.
[0061] This embodiment of the application also discloses an electronic device for three-dimensional point cloud completion based on dual-branch feature extraction and attention mechanism, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the three-dimensional point cloud completion method based on dual-branch feature extraction and attention mechanism. Among them, the memory may include internal memory, such as high-speed random access memory, and may also include non-volatile memory, such as at least one disk memory, etc.; the processor, network interface, and memory are interconnected through an internal bus, which can be an Industry Standard Architecture bus, a Peripheral Component Interconnect standard bus, an Extended Industry Standard Architecture bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The memory is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.
[0062] The embodiments of the present application also disclose a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of a three-dimensional point cloud completion method based on dual-branch feature extraction and an attention mechanism. Specifically, the computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. The volatile memory may include random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include read-only memory (ROM), hard disk, flash memory, optical disc, magnetic disk, etc.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A three-dimensional point cloud completion method based on dual-branch feature extraction and attention mechanism, characterized in that: The following steps are involved: Step 1: Prepare a 3D point cloud dataset of any size and quantity, unify the number of points of the residual point cloud, and preprocess all point cloud data; Step 2: Design a 3D point cloud completion model, take the residual point cloud as input, generate a complete and dense point cloud model, and have the geometric details of the residual point cloud; Step 3: training the three-dimensional point cloud completion model; Step 4: Repeat step 3 until the preset number of iterations is reached, traversing all 3D point clouds in each round. At the end of each round of iteration, update and save the model parameters; Step 5: Select the model parameters with the best indicators and load them into the model; input the residual point cloud and generate the corresponding complete point cloud.
2. A three-dimensional point cloud completion method based on dual-branch feature extraction and attention mechanism according to claim 1, characterized in that: The 3D point cloud completion model adopts an encoder-decoder architecture; The encoder takes the residual defect cloud as input and generates a shape code of the residual defect cloud; The decoder is used to enable the shape encoding to generate a complete point cloud.
3. A three-dimensional point cloud completion method based on dual-branch feature extraction and attention mechanism according to claim 2, characterized in that: The encoder is specifically a dual-branch feature extraction module, which includes two branches: a global feature extraction module and a local feature extraction module; The global feature extraction module includes a multi-layer perceptron and a pooling layer, which takes the residual defect cloud as input to obtain the global feature; The local feature extraction module adds a CSA module to the existing feature extraction module of the baseline model to obtain local features.
4. A three-dimensional point cloud completion method based on dual-branch feature extraction and attention mechanism according to claim 3, characterized in that: The CSA module is composed of a stack of a channel attention mechanism and a spatial attention mechanism; The channel attention mechanism performs maximum pooling and average pooling operations on the input feature matrix respectively, and generates a feature matrix obtained by performing maximum pooling and average pooling on the channel dimension of the input feature matrix; The obtained feature matrices are sent to the multi-layer perceptron respectively, and then added. After activation by a nonlinear function, the channel attention weight is obtained. The input feature matrix and the channel attention weight are operated on the Hadamard product to obtain the feature matrix based on channel attention.
5. A three-dimensional point cloud completion method based on dual-branch feature extraction and attention mechanism according to claim 2 or 3, characterized in that: The decoder includes a sparse point cloud generator and a dense point cloud generator; The sparse point cloud generator takes the residual point cloud and shape encoding as input to generate a sparse but relatively complete point cloud model; The dense point cloud generator takes the residual point cloud shape encoding and the rough point cloud as input to generate a complete and dense point cloud model with the geometric details of the residual point cloud.
6. A three-dimensional point cloud completion method based on dual-branch feature extraction and attention mechanism according to claim 1, characterized in that: The step 3 specifically includes: Step 3-1: Assign the corresponding shape code to the input residual defect cloud; Step 3-2: Input the shape code into the sparse point cloud generator to generate a rough but complete point cloud, which is recorded as a rough point cloud; Step 3-3: Input the shape encoding and the rough point cloud into the first dense point cloud generator to generate a dense point cloud of the first stage; Step 3-4: Input the shape encoding and the first stage dense point cloud into the second dense point cloud generator to generate the final dense point cloud; Step 3-5: Calculate the loss function value between the generated point cloud and the real point cloud.
7. A three-dimensional point cloud completion method based on dual-branch feature extraction and attention mechanism according to claim 6, characterized in that: The step 3-1 specifically includes: Step 3-11: Use the residual defect cloud as the input of the global feature extraction module, and obtain the feature matrix of the residual defect cloud after passing through a multi-layer perceptron; Step 3-12: Obtaining global features and local features of the feature matrix; Step 3-13: Concatenate the global features and the local features to obtain a shape code, and use a multi-layer perceptron to map the shape code to obtain the final shape code.
8. A three-dimensional point cloud completion method based on dual-branch feature extraction and attention mechanism according to claim 1, characterized in that: The format of the 3D point cloud dataset prepared in step 1 is .h5 or .pcd.
9. An electronic device for three-dimensional point cloud completion based on dual-branch feature extraction and attention mechanism, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a three-dimensional point cloud completion method based on dual-branch feature extraction and attention mechanism as described in any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, which is used to execute a three-dimensional point cloud completion method based on dual-branch feature extraction and attention mechanism as described in any one of claims 1 to 8.
Citation Information
Cited By
High-fidelity point cloud completion method and system based on double-path attention and fractal structure
CN120655840A
Lightweight three-dimensional point cloud completion method and device and processing equipment
CN120807806A
Attention anti-mask double-branch distillation method for three-dimensional point cloud completion
CN121279396A
Attention unmasking dual-branch distillation method for three-dimensional point cloud completion
CN121279396B