Double-branch indoor scene point cloud completion method and system based on feature fusion

By adopting a dual-branch method of feature fusion in indoor scene point cloud completion, the problem that the existing technology is difficult to complete the missing parts of the main indoor structure is solved, high-quality point cloud completion is achieved, and the accuracy and uniformity of the completion effect are improved.

CN119992128APending Publication Date: 2025-05-13BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510062021.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing cloud completion method for indoor scene points is difficult to effectively complete the missing parts that constitute the main indoor structure, such as walls, floors, doors and windows.

Method used

A dual-branch method based on feature fusion is adopted to generate the overall characteristics of the point cloud through the fusion of global features and local features, and a coarse point cloud is generated through the folding operation in the decoder, and finally the completed point cloud data is obtained through the farthest point sampling.

Benefits of technology

It realizes effective completion of missing structural shapes in indoor scenes, improves the accuracy and distribution uniformity of point cloud repair, reduces the average chamfer distance and average ground movement distance, and improves the F1 score.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992128A_ABST
    Figure CN119992128A_ABST
Patent Text Reader

Abstract

The invention discloses a double-branch indoor scene point cloud completion method based on feature fusion, and the method comprises the steps: inputting a to-be-completed point cloud into a trained point cloud completion model, and obtaining a completed point cloud; the trained point cloud completion model is obtained by the following steps: obtaining a pre-training point cloud data set, wherein each point cloud sample in the data set comprises a complete real point cloud and a missing point cloud corresponding to the complete real point cloud; performing global feature extraction and local feature extraction on the missing point cloud, and performing feature fusion to obtain feature code words; splicing the feature code word with the two-dimensional grid, and carrying out folding decoding to obtain a coarse point cloud; combining the coarse point cloud with the missing point cloud, and performing resampling to obtain a complemented point cloud; training the model by using a loss function LOSS to obtain a trained point cloud completion model; the invention further discloses a system based on the method. According to the method provided by the invention, the missing point cloud of the indoor scene can be complemented, and the method has better robustness for data missing of different degrees.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and more specifically, to a dual-branch indoor scene point cloud completion method and system based on feature fusion. Background Art

[0002] When performing 3D laser scanning in complex indoor scenes, various indoor facilities will inevitably block the indoor space structure, making the point cloud data incomplete, with holes and unevenness, which affects indoor data analysis and visualization, indoor environment modeling, etc. Therefore, it is necessary to complete the missing indoor scenes.

[0003] The existing indoor scene point cloud completion methods mainly include: 1. Indoor point cloud scene repair and completion method based on category-instance segmentation, which uses semantic segmentation to take the missing indoor scene point cloud as the input of the completion network to repair and complete the indoor scene. 2. Point cloud completion of incomplete objects in indoor scenes based on conditional generative adversarial networks. However, the above methods can only complete the specific parts of the indoor home, and do not repair the main structures that constitute the indoor space, such as walls, floors, doors and windows, etc. Summary of the invention

[0004] An object of the present invention is to provide a dual-branch indoor scene point cloud completion method and system based on feature fusion to at least solve the above-mentioned problems.

[0005] In order to achieve the purpose and other advantages of the present invention, a dual-branch indoor scene point cloud completion method based on feature fusion is provided, comprising: inputting the point cloud to be completed into a trained point cloud completion model to obtain the completed point cloud; wherein,

[0006] The trained point cloud completion model is mainly obtained by the following steps: Step 1: Obtain a pre-trained point cloud dataset, where each point cloud sample in the pre-trained point cloud dataset includes a complete real point cloud S gt and the missing point cloud S1 corresponding to it; step 2, extracting global features and local features from the missing point cloud S1 respectively, and performing feature fusion on the extracted global feature vector and local feature vector to obtain a feature codeword; step 3, splicing the feature codeword with the two-dimensional grid, and performing folding decoding to obtain a coarse point cloud S c ; Step 4: transform the coarse point cloud S c Merge with the missing point cloud S1 and resample to obtain the completed point cloud S2; Step 5, train the model with the loss function LOSS to obtain the trained point cloud completion model,

[0007]

[0008] In the formula, S c is the coarse point cloud, S gt is the real point cloud, S2 is the completed point cloud, d cd is the chamfer distance.

[0009] Further,

[0010]

[0011] Where x, y, and z are the coarse point clouds S c , real point cloud S gt And the points in the completed point cloud S2.

[0012] Preferably, in step 2, a multi-layer perceptron and a maximum pooling layer are used to perform global feature extraction on the missing point cloud to obtain the global feature vector.

[0013] Specifically, the missing point cloud is a discrete unordered point cloud matrix S1 composed of N points, and each row represents the three-dimensional coordinates of a point cloud. First, the missing point cloud is input into a multi-layer perceptron (MLP). The multi-layer perceptron in the encoder consists of a two-layer linear structure. Through ReLU activation, a single point cloud data is converted into a feature vector of the point cloud. Therefore, each row of the overall feature of the input point cloud is the feature of each learned point cloud, and its dimension is N×128, where N represents the number of points. Then the feature matrix is ​​passed through the maximum pooling layer to obtain the global feature of the input point cloud, which is a 1×128-dimensional feature codeword. It is expanded to the same dimension as the feature vector and concatenated with the feature vector, and its dimension is N×256. Then it is passed through a multi-layer perceptron to obtain a second N×512-dimensional feature matrix, which is subjected to maximum pooling to obtain the final global feature vector, and its output dimension is 1×512.

[0014] Preferably, in step 2, when extracting local features of the missing point cloud, three layers of EdgeConv are first used to superimpose different feature dimensions, and then the superimposed local features are subjected to maximum pooling and average pooling operations respectively and the results are concatenated to obtain the local feature vector.

[0015] Specifically, the EdgeConv module updates the feature representation of a node by considering the connection relationship between the node and its neighboring nodes. This method enables the network to better capture the local information between point clouds and the relationship between nodes. It maintains the local geometric structure of the data by constructing a dynamically changing local neighborhood graph in the feature space. The core idea of ​​the network is to apply edge convolution to the edges between each node and its neighbors, thereby extracting and propagating local feature information while retaining the geometric relationship of the input data. In order to extract accurate local features, this technical solution adopts a three-layer EdgeConv form, superimposes different feature dimensions (64, 64, 128) to obtain 1×256-dimensional features, and performs maximum pooling and average pooling operations on the superimposed local features and splices them to obtain local features of the point cloud, with an output dimension of 1×512.

[0016] Furthermore, in step 2, the global features and local features extracted by the two branches are merged and fused through a multi-layer perceptron (MLP) to obtain the final feature codeword V, whose output dimension is 1×1024.

[0017] Preferably, in step three, after the feature codeword is copied, it is spliced ​​with the two-dimensional grid and folded and decoded twice to obtain the coarse point cloud.

[0018] Specifically, the feature codeword V of the input point cloud obtained in the encoder of step 2 is copied M times, and the obtained M×1024-dimensional matrix is ​​connected with the two-dimensional matrix containing M grid points, and the two-dimensional matrix contains M grid points on a square centered on the origin. The result of the connection is an M×(1024+2) matrix, which is input into a 3-layer perceptron for processing as the first folding to obtain an M×3-sized matrix. Then, the M×1024 matrix is ​​connected with the output M×3 matrix again to obtain an M×(1024+3) matrix, which is then input into the 3-layer perceptron. After the second folding operation, the fitted coarse point cloud Sc is output.

[0019] Preferably, in step 4, the coarse point cloud and the missing point cloud are merged and then the farthest point sampling is performed to obtain the completed point cloud.

[0020] Specifically, in order to make the completed point cloud have the overall characteristics of the point cloud while retaining the original characteristics of the point cloud, the fitted coarse point cloud Sc is merged with the input missing point cloud S1 and then FPS resampled with 4096 sampling points to obtain the final completed point cloud S2.

[0021] Preferably, in step 5, when training the model, the optimizer used is Adam, the number of training cycles is 300, the number of training batches is 60, and the learning rate is 0.0001.

[0022] The present invention also provides a dual-branch indoor scene point cloud completion system based on feature fusion, comprising: a point cloud completion module, which is used to input the point cloud to be completed into a trained point cloud completion model to obtain the completed point cloud; wherein,

[0023] The trained point cloud completion model is mainly obtained by the following steps: step 1, obtaining a pre-trained point cloud data set, each point cloud sample in the pre-trained point cloud data set includes a complete real point cloud and a missing point cloud corresponding thereto; step 2, performing global feature extraction and local feature extraction on the missing point cloud respectively, and performing feature fusion on the extracted global feature vector and local feature vector to obtain a feature codeword; step 3, splicing the feature codeword with a two-dimensional grid, and performing folding decoding to obtain a coarse point cloud; step 4, merging the coarse point cloud with the missing point cloud, and resampling to obtain a completed point cloud; step 5, training the model with a loss function LOSS to obtain the trained point cloud completion model,

[0024]

[0025] In the formula, S c is the coarse point cloud, S gt is the real point cloud, S2 is the completed point cloud, d cd is the chamfer distance.

[0026] The present invention also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor performs the above-mentioned method.

[0027] The present invention also provides a computer-readable medium on which a computer program is stored. When the program is executed by a processor, the above method is implemented.

[0028] The present invention also provides a computer program product, comprising a computer program / instruction, which implements the above method when executed by a processor.

[0029] The present invention has at least the following beneficial effects:

[0030] Aiming at the problem of missing point cloud data caused by occlusion in indoor scanning, the present invention proposes a dual-branch indoor scene point cloud completion method based on feature fusion. The method uses incomplete missing point cloud data as the input of the encoder, extracts global and local features through two branches respectively, and fuses the two features to form the overall features of the point cloud, and generates a coarse point cloud through two folding operations in the decoder, and then fuses the coarse point cloud with the missing point cloud of the original input, and obtains the complete completed point cloud data through the farthest point sampling technology. Experimental verification shows that the method of the present invention can effectively complete the missing structural shape in the indoor scene and realize the completion of the missing point cloud of the indoor scene. Compared with the existing method, it performs better in terms of point cloud repair error and point cloud distribution uniformity, and its average chamfer distance is reduced by 6.19% to 47.04%, the average ground moving distance is reduced by 12.72% to 40.18%, the average F1 score (F1-Score) is improved by 1.38% to 22.45%, and it has good robustness to different degrees of data missing.

[0031] Other advantages, objectives and features of the present invention will be embodied in part through the following description, and in part will be understood by those skilled in the art through study and practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 This is a structural diagram of a dual-branch indoor scene point cloud completion network based on feature fusion of the present invention;

[0033] Figure 2 is the CD error distribution result of different networks;

[0034] Figure 3 is the EMD error distribution result of different networks;

[0035] Figure 4 This is the visualization of the point cloud completion results of different networks;

[0036] Figure 5 is the point cloud completion result of the point cloud completion network of the present invention at different missing rates;

[0037] Figure 6 It is the completion result of the point cloud completion network of the present invention in the indoor scene;

[0038] Figure 7 It is the completion result of the point cloud completion network of the present invention in an actual indoor occlusion scene. DETAILED DESCRIPTION

[0039] The present invention is further described in detail below in conjunction with embodiments and drawings so that those skilled in the art can implement the invention with reference to the description.

[0040] It should be understood that the terms such as “having”, “including” and “comprising” used in the present invention do not exclude the existence or addition of one or more other elements or combinations thereof.

[0041] It should be noted that the experimental methods described in the following embodiments are conventional methods unless otherwise specified, and the reagents and materials can be obtained from commercial channels unless otherwise specified.

[0042] In order to repair the missing point clouds in indoor scenes, the present invention provides a dual-branch indoor scene point cloud completion model based on feature fusion, which can effectively repair and complete the input missing point clouds. Its network framework is as follows Figure 1 shown.

[0043] The point cloud completion model extracts the global features and local features of the input missing point cloud S1 through two independent branches. The global features are processed by multi-layer perceptron combined with the maximum pooling layer, while the local features are extracted by three layers of EdgeConv. The global features and local features are fused and input into the decoder as feature codewords for operation. The decoder is a decoder structure based on the folding idea. The 2D plane point cloud is folded twice to fit the coarse point cloud S c , and then sample the farthest point with the input missing point cloud S1 to obtain the final completed point cloud data S2.

[0044] The encoder uses a dual-branch encoder, whose main task is to process the information of the input point cloud and output it as the feature vector of the point cloud. The dual-branch encoder extracts global and local features of the input point cloud through two independent branches, and fuses them to obtain the feature codeword of the input point cloud, which contains both the overall information of the point cloud and the local detail features. When extracting global features, the input of the encoder is a discrete unordered point cloud matrix composed of N points, that is, the missing point cloud S1, and each row in the matrix represents the three-dimensional coordinates of a point cloud. The first branch of the encoder uses a multi-layer perceptron (MLP) and a maximum pooling layer to process the extracted feature information to obtain a 512-dimensional global feature vector. Specifically, the point cloud matrix is ​​first input into the multi-layer perceptron. Among them, the multi-layer perceptron in the encoder consists of a two-layer linear structure. Through ReLU activation, the single point cloud data is converted into the feature vector of the point cloud. Therefore, each row of the overall feature of the input point cloud is the feature of each learned point cloud, and its dimension is N×128, where N represents the number of points. Then the feature matrix is ​​passed through the maximum pooling layer to obtain the global features of the input point cloud, which is a 1×128-dimensional feature codeword. It is expanded to the same dimension as the feature vector and concatenated with the feature vector, with a dimension of N×256. Then it is passed through a multi-layer perceptron to obtain a second N×512-dimensional feature matrix, which is subjected to maximum pooling to obtain the final global feature vector, with an output dimension of 1×512. When extracting local features, a second branch is added to the encoder to obtain the local feature vector of the point cloud, with a dimension of 512. Specifically, a three-layer EdgeConv is used to superimpose different feature dimensions (64, 64, 128) to obtain 1×256-dimensional features. The superimposed local features are subjected to maximum pooling and average pooling operations respectively and concatenated to obtain the local features of the point cloud, with an output dimension of 1×512. Finally, the global features and local features extracted by the two branches are merged and fused through a multi-layer perceptron (MLP) to obtain the final feature codeword V, with an output dimension of 1×1024.

[0045] The decoder is based on the folding idea, and is a process of fitting a 2D point cloud plane into a 3D shape through two folding operations. It replicates the feature codeword V obtained from the dual-branch encoder M times, splices it with the 2D point cloud grid, and then obtains a coarse point cloud Sc through two folding operations. This point cloud Sc is merged with the input point cloud S1 in the encoder and then farthest point sampling (FPS) is performed to obtain the completed point cloud S2. Specifically, the feature codeword V of the input point cloud obtained in the encoder in step 2 is replicated M times, and the obtained M×1024-dimensional matrix is ​​connected to a two-dimensional matrix containing M grid points, which contains M grid points on a square centered on the origin. The result of the connection is an M×(1024+2) matrix, which is input into a 3-layer perceptron for processing as the first folding, resulting in an M×3-sized matrix. Then, the M×1024 matrix is ​​connected to the output M×3 matrix again to obtain the M×(1024+3) matrix, which is then input into the 3-layer perceptron. After the second folding operation, the fitted coarse point cloud Sc is output. In order to make the completed point cloud have the overall characteristics of the point cloud while retaining the original characteristics of the point cloud, the fitted coarse point cloud Sc is merged with the input missing point cloud S1 and resampled by FPS. The number of sampling points is 4096, and the final completed point cloud S2 is obtained.

[0046] The loss function is composed of the rough point cloud Sc, the completed point cloud S2 and the real point cloud S gt The mean of the two CD distances is used as the loss function between point clouds to measure the completion quality of the point cloud. The formula is as follows:

[0047]

[0048] In the formula, S c is the coarse point cloud, S gt is the real point cloud, S2 is the completed point cloud, d cd is the chamfer distance. In this formula, Loss takes the average value of the two CD distances, so that the point cloud is evenly distributed without losing its overall characteristics.

[0049] Further,

[0050]

[0051]

[0052] Where x, y, and z are the coarse point clouds S c , real point cloud S gt And the points in the completed point cloud S2.

[0053] In order to verify the application effect of the dual-branch indoor scene point cloud completion network based on feature fusion of the present invention, the following experiments were carried out. The specific experimental process is as follows:

[0054] The point cloud completion network proposed in this invention is implemented under the Ubuntu 20.04 system, with GPU L20, Python version 3.8, and pytorch version 1.11.0.

[0055] 1.1 Dataset and Preprocessing

[0056] Given the structural similarity between ceilings and floors in the indoor model, they are merged into one category in the point cloud completion process. In the experiment, the following point cloud data were collected: 4206 wall samples, 3480 ceiling and floor samples, 3528 column samples, 3048 beam samples, 3630 door samples, and 3488 window samples. They were randomly sampled to 4096 points.

[0057] Since the range of occlusion of the point cloud of the real scanned indoor scene is different, when making incomplete point cloud data samples, for each complete point cloud model, a cubic bounding box is constructed, and four sampling points are randomly selected on the surface of the bounding box. Then, with these four sampling points as the center, according to two different missing rates of 25% and 40%, the corresponding proportion of point cloud data closest to each sampling point is removed respectively, and incomplete point cloud data with different missing ratios at 8 different positions is obtained.

[0058] The present invention selects 60% of all data set samples as training sets, 20% as validation sets, and the rest as test sets. Each point cloud sample set includes a complete point cloud and at least one missing point cloud corresponding to it, and the ratio of complete point cloud to missing point cloud in the training set is 1:8, and the ratio of validation set to test set is 1:1. The method uses missing point cloud data as input and uses the corresponding real point cloud as verification for training.

[0059] The point cloud completion network of the present invention sets the number of training cycles to 300, the training batches to 60, the optimizer to Adam, and the learning rate to 0.0001.

[0060] 1.2 Comparison of results with different point cloud completion methods

[0061] In order to analyze the performance of the point cloud completion network proposed in this invention, the network proposed in this invention is qualitatively and quantitatively evaluated with FoldingNet, PCN, non-resampling point cloud completion network (without resampling) and point cloud completion network without local feature extraction (without EdgeConv). In this experiment, the same training set and test set are used for training, the number of sampling points of all missing point cloud models is unified to 3072, and the number of sampling points of the point cloud models generated after completion by different networks and the real point cloud models are all 4096.

[0062] The completion effect is quantitatively analyzed using three statistical indicators: CD distance (chamfer distance), EMD distance (ground moving distance) and F1-score. The smaller the CD spacing, the more complete the outline of the point cloud and the higher the boundary smoothness; the smaller the EMD spacing, the higher the spatial resolution of the point cloud and the closer the density; the higher the F1-score value, the more points in the point cloud generated by the completion network are located in the correct spatial position, reflecting the accuracy of the generated point cloud.

[0063] The point cloud completion effects of these methods on the test data set are compared and analyzed. The results are shown in Table 1, Table 2, and Table 3, where the bold part is the best completion effect. The error distribution results of different networks in CD and EMD distance are shown in Figure 2 , Figure 3 .

[0064] Table 1 Comparison of point cloud completion results of different networks (CD / 10 3 )

[0065]

[0066] Table 2 Comparison of point cloud completion results of different networks (EMD)

[0067]

[0068] Table 3 Comparison of point cloud completion results of different networks (F1-score / %)

[0069]

[0070] from Figure 2 and Figure 3It can be seen that the network of the present invention can achieve a lower CD value compared with the above networks. In addition, in most categories, the EMD value of the network of the present invention is also lower than that of other networks. As shown in Table 1, the CD values ​​of various point cloud models are better than those of other comparison networks, and their average CD value is 6.19% lower than that of the optimal network. As shown in Table 2, the PCN network is slightly better than the network of the present invention in wall completion, the EMD value of the column class of the network (without EdgeConv) is slightly lower than that of the network of the present invention, and the EMD value of the beam class of the network (without Resampling) is the lowest. The EMD values ​​of other types of point cloud models show better results compared with other networks, and the average EMD value is 12.72% lower than that of the optimal network. As shown in Table 3, the F1-score values ​​of various point clouds are better than those of other comparison networks, and their average F1-score values ​​are 1.38% higher than that of the optimal network. This is because the dual-branch network of the present invention uses EdgeConv to extract local features of the point cloud on the basis of extracting global features, so that the generated point cloud is more evenly distributed while preserving local information, thereby obtaining lower CD and EMD values. At the same time, the input point cloud is fused with the coarse point cloud generated by the network decoder and resampled, which retains the input point cloud information, improves the accuracy of the point cloud completion network, and optimizes the point cloud completion effect.

[0071] The visualization of the completed data is as follows Figure 4 As shown in the figure, observing the completion results of the input point cloud, it can be concluded that: FoldingNet can roughly restore the overall structure in point cloud completion, but because its decoder only uses the 2D network superposition method, the output point cloud is insufficient in detail performance, especially the edges of beams and columns are distorted, and the frames of doors and windows are not accurately restored. The PCN network focuses on the global features of the point cloud in the encoding stage and generates high-density point clouds through global features, but this approach leads to uneven distribution of point clouds and inaccurate detail restoration. The network (without Resampling) integrates the global features and local features of the input point cloud without performing point cloud resampling operations, so that the generated point cloud model can better retain the features of the input point cloud, such as the frame part of doors and windows, but there is still the problem of uneven point cloud density, and there is a certain gap with the actual point cloud model. The network (without EdgeConv) does not extract local features, resulting in the edges of the completed point cloud not being smooth enough. For example, at the boundary between the ceiling and the beam, the features of the door and window frames are not obvious.

[0072] Comprehensive comparison shows that the network proposed in the present invention performs best in point cloud completion. The generated point cloud model has clear edges, high fusion with the original point cloud, significant geometric features, and relatively uniform point cloud density distribution. Experimental results confirm that the method of the present invention can effectively complete the missing point cloud structure indoors, and the completion result has high accuracy.

[0073] 1.3 Analysis of completion results for different missing situations

[0074] In order to evaluate the completion effect of the point cloud completion network proposed in this invention on point cloud models with different degrees of missingness, we tested its robustness. Figure 5 The network is shown to complete point cloud models with different missing proportions. Figure 5 It can be seen that even when the degree of missing point cloud is different, the network can still maintain the shape and structural information of the original point cloud well. However, as the missing ratio increases, the features of the point cloud become less significant, which leads to a decrease in the completion effect. The experimental results confirm that when facing point cloud models with different degrees of missing, the network proposed in the present invention can still effectively complete the missing parts, showing good robustness.

[0075] 1.4 Indoor scene completion effect

[0076] In order to verify the effectiveness of the point cloud completion network proposed in this invention in completing the missing point cloud of indoor scenes, the network is used to complete the missing point cloud data, and the completed point cloud model is fused to obtain the completed complete point cloud of indoor scenes. In the experiment, the complete indoor point cloud converted from the obj model is used, and some point cloud data is randomly removed to simulate the missing situation that may occur in the actual scanning process.

[0077] Figure 6 The point cloud data comparison of three different indoor scenes is shown. First, the first row shows the missing point cloud of the indoor scene, the second row shows the point cloud processed by the point cloud completion network, and the third row shows the complete point cloud of the indoor scene as a reference. The experimental results show that the network proposed in this invention effectively completes the missing point cloud data in indoor scenes. Specifically, Figure 6 The overall completion effect of the classrooms and conference rooms is very good, and the missing indoor columns are well completed, and the completion effect is uniform and natural. Figure 6 In the lobby, although the completion effect of the indoor grid-shaped windows failed to accurately restore its characteristic details, the completion effect is still considerable overall. This is mainly because when processing point cloud data with special structures, the network of the present invention cannot fully and effectively learn and predict its precise structure. Nevertheless, the network can still capture and reconstruct the general outline and overall shape of the window, thus presenting a relatively satisfactory completion effect overall.

[0078] 1.5 Analysis of indoor completion results under actual occlusion conditions

[0079] When performing point cloud scanning of indoor scenes, the point clouds obtained are often incomplete and unevenly distributed due to the limitation of scanning angle and the occlusion of indoor objects. In order to test the completion ability of the point cloud completion network proposed in this invention when the actual indoor scene point cloud is missing, the large indoor scene dataset S3DIS (Stanford 3D Indoor Semantics Dataset) is used to complete the point cloud obtained under actual occlusion conditions as test samples. The missing point cloud data is put into the point cloud completion network for completion, and then the completed model is fused to obtain the completed indoor scene point cloud model under actual conditions.

[0080] Figure 7 This is the completion effect of the network of the present invention on the actual indoor scene point cloud. Figure 7 (a) shows the point cloud data of the indoor scene obtained by the original scan. Figure 7 (b) shows the missing point cloud data formed after removing the indoor occlusions, while Figure 7 (c) is the complete indoor scene point cloud after completion. It can be clearly seen from the figure that the point cloud completion network proposed in this study can effectively complete the missing point clouds in the actual scene, proving the practicality and applicability of the network model in practical applications, and can meet the needs of indoor scene point cloud completion.

[0081] In summary, this paper proposes a dual-branch indoor scene point cloud completion network based on feature fusion, that is, in the encoder part, the global and local features of the point cloud are extracted by two branches respectively for fusion, and then decoding and iterative optimization are performed, and finally a high-quality and evenly distributed completed point cloud model is generated, which realizes the completion of the missing point cloud of the indoor scene. The results show that:

[0082] 1) Under the same test set, compared with other networks, the average CD error of the network of the present invention is reduced by 6.19% to 47.04%, the average EMD error is reduced by 12.72% to 40.18%, the average F1-score is improved by 1.38% to 22.45%, and the generated point cloud is more uniform.

[0083] 2) Effective completion can be achieved for point cloud data with different degrees of missingness. However, as the missing proportion increases, the completion effect will also decrease.

[0084] 3) By fusing indoor point cloud data of different categories with multiple details, the missing indoor scenes can be effectively improved to meet the needs of indoor scene point cloud completion.

[0085] In addition, the present invention also provides a dual-branch indoor scene point cloud completion system based on feature fusion, which is based on the same inventive concept as the above-mentioned dual-branch indoor scene point cloud completion method based on feature fusion, and will not be repeated here.

[0086] The present invention also provides an electronic device, which is a device including a processor (CPU / MCU / SOC) and a memory (ROM / RAM), such as a portable computer, a smart phone, etc. In particular, a computer program is stored in the memory, and when the processor loads and executes the computer program, all or part of the steps of the above-mentioned dual-branch indoor scene point cloud completion method based on feature fusion are implemented.

[0087] The present invention also provides a computer-readable medium, which includes: various media that can store program codes, such as ROM, RAM, magnetic disk or optical disk, wherein a computer program is stored. When the computer program is loaded and executed by a processor, all or part of the steps of the above-mentioned feature fusion-based dual-branch indoor scene point cloud completion method are implemented.

[0088] The present invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements all or part of the steps of the above-mentioned dual-branch indoor scene point cloud completion method based on feature fusion.

[0089] The number of devices and processing scale described here are used to simplify the description of the present invention. The application, modification and variation of the dual-branch indoor scene point cloud completion method based on feature fusion of the present invention are obvious to those skilled in the art.

[0090] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and the implementation modes, and they can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and the illustrations shown and described herein.

Claims

1. A dual-branch indoor scene point cloud completion method based on feature fusion, characterized in that: include: Input the point cloud to be completed into the trained point cloud completion model to obtain the completed point cloud; The trained point cloud completion model is mainly obtained by the following steps: Step 1: obtaining a pre-training point cloud data set, wherein each point cloud sample in the pre-training point cloud data set includes a complete real point cloud and a missing point cloud corresponding thereto; Step 2: extracting global features and local features from the missing point cloud respectively, and fusing the extracted global feature vectors and local feature vectors to obtain feature codewords; Step 3: splice the feature codeword with a square two-dimensional grid generated with the origin as the center, and perform folding decoding to obtain a coarse point cloud; Step 4: Merge the coarse point cloud with the missing point cloud, and resample them to obtain a completed point cloud; Step 5: Train the model with the loss function LOSS to obtain the trained point cloud completion model. In the formula, S c is the coarse point cloud, S gt is the real point cloud, S2 is the completed point cloud, d cd is the chamfer distance.

2. The dual-branch indoor scene point cloud completion method based on feature fusion as claimed in claim 1, characterized in that: In step 2, a multi-layer perceptron and a maximum pooling layer are used to extract global features of the missing point cloud to obtain the global feature vector.

3. The dual-branch indoor scene point cloud completion method based on feature fusion as claimed in claim 1, characterized in that: In step 2, when extracting local features of the missing point cloud, three layers of EdgeConv are first used to superimpose different feature dimensions, and then the superimposed local features are subjected to maximum pooling and average pooling operations respectively, and the results are concatenated to obtain the local feature vector.

4. The dual-branch indoor scene point cloud completion method based on feature fusion as claimed in claim 1, characterized in that: In step three, the feature codeword is copied, spliced ​​with the two-dimensional grid, and folded and decoded twice to obtain the coarse point cloud.

5. The dual-branch indoor scene point cloud completion method based on feature fusion as claimed in claim 1, characterized in that: In step 4, the coarse point cloud and the missing point cloud are merged and the farthest point sampling is performed to obtain the completed point cloud.

6. The dual-branch indoor scene point cloud completion method based on feature fusion as claimed in claim 1, characterized in that: In step 5, when training the model, the optimizer used is Adam, the number of training cycles is 300, the training batch is 60, and the learning rate is 0.0001.

7. A dual-branch indoor scene point cloud completion system based on feature fusion, characterized in that: include: The point cloud completion module is used to input the point cloud to be completed into the trained point cloud completion model to obtain the completed point cloud; wherein, The trained point cloud completion model is mainly obtained by the following steps: Step 1: obtaining a pre-training point cloud data set, wherein each point cloud sample in the pre-training point cloud data set includes a complete real point cloud and a missing point cloud corresponding thereto; Step 2: extracting global features and local features from the missing point cloud respectively, and fusing the extracted global feature vectors and local feature vectors to obtain feature codewords; Step 3: splice the feature codeword with a square two-dimensional grid generated with the origin as the center, and perform folding decoding to obtain a coarse point cloud; Step 4: Merge the coarse point cloud with the missing point cloud, and resample them to obtain a completed point cloud; Step 5: Train the model with the loss function LOSS to obtain the trained point cloud completion model. In the formula, S c is the coarse point cloud, S gt is the real point cloud, S2 is the completed point cloud, d cd is the chamfer distance.

8. An electronic device, characterized in that include: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor executes the method according to any one of claims 1 to 6.

9. A computer readable medium having a computer program stored thereon, characterized in that When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. Computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Point cloud completion method, system and device based on self-supervised learning and medium

    CN120182144A