Model training, point cloud missing completion method, device, equipment and medium
By adjusting the parameters of the initial model, using the target to reconstruct the network and initially generate the network, the problem of low data accuracy in point cloud missing completion is solved, and a higher point cloud data accuracy is achieved.
Patent Information
- Application Number
- CN202111129999.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-09-26
AI Technical Summary
The data accuracy of the existing technology is low in point cloud defect completion, and data repair cannot be effectively carried out from the perspective of existing data.
By obtaining the training missing point cloud data, enter the initial model to obtain the training repair point cloud data, and adjust the parameters of the initial model based on the training repair point cloud data and the original point cloud data until the training completion conditions are met, and the initial model is determined to be the point cloud completion model. The model includes the target reconstruction network and the initial generation network, and uses comparative learning and reconstruction loss values for parameter adjustment.
The accuracy of point cloud data after completion processing is improved, the problem of low data accuracy is solved, and the missing point cloud is accurately predicted through a global structure containing information from different local areas.
Smart Images

Figure CN113850916B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to model training, point cloud missing completion methods, devices, electronic devices and computer-readable storage media. Background Art
[0002] 3D reconstruction technology reconstructs 3D objects in the virtual world and is the basis for the realization of 3D visual technologies such as VR / AR (Virtual Reality). In recent years, with the development of sensors and deep learning, 3D point clouds have become the mainstream representation of 3D reconstruction results. However, due to mutual occlusion between objects and technical limitations of hardware equipment, 3D reconstruction results based on point clouds have holes or missing shape structures. At present, some research works have proposed point-based completion methods, which can directly process point cloud data and obtain point cloud features, and predict complete 3D point clouds or missing point clouds through decoders based on full connection or folding, and repair and complete 3D reconstruction results. Compared with voxel model representation, directly inputting point clouds reduces the amount of input data and the parameter scale of neural networks, which greatly improves the network training speed. However, the related technology extracts features from missing input point clouds to obtain feature representations of input point clouds, which can only repair data from the perspective of existing data, thereby reducing the accuracy of the completed data generated by the model.
[0003] Therefore, the problem of low data accuracy in the related technology is a technical problem that needs to be solved by those skilled in the art. Summary of the invention
[0004] In view of this, the purpose of the present application is to provide a model training, point cloud missing completion method, device, electronic device and computer-readable storage medium, so as to improve the accuracy of the processed point cloud data after completion.
[0005] To solve the above technical problems, the present application provides a model training method, including:
[0006] Get missing point cloud data for training;
[0007] Inputting the training missing point cloud data into an initial model to obtain training repair point cloud data, and adjusting the parameters of the initial model based on the training repair point cloud data and the original point cloud data corresponding to the training missing point cloud data;
[0008] If it is detected that the training completion condition is met, determining that the initial model is a point cloud completion model;
[0009] Among them, the initial model includes a target reconstruction network and an initial generation network, the target reconstruction network includes a target encoding network, the target encoding network uses the training missing point cloud data for comparative learning, the training missing point cloud data is input into the target encoding network to obtain input features, the input features are input into the initial generation network to obtain missing point cloud data, and the missing point cloud data is used to generate the training repair point cloud data.
[0010] Optionally, the generation process of the initial model includes:
[0011] Using the training missing point cloud data to perform learning and training on the initial reconstruction network to obtain the target reconstruction network;
[0012] The initial model is obtained by combining the target reconstruction network with the initial generation network.
[0013] Optionally, the using the training missing point cloud data to perform learning and training on the initial reconstruction network to obtain the target reconstruction network includes:
[0014] Determine an anchor point cloud from the training missing point cloud data;
[0015] Based on the anchor point cloud, the training missing point cloud data is input into the initial reconstruction network to obtain target data; wherein the target data includes the input features and the reconstructed point cloud data;
[0016] Using the input features to obtain a contrastive learning loss value, using the reconstructed point cloud data to obtain a reconstruction loss value, and using the contrastive learning loss value and the reconstruction loss value to adjust parameters of the initial reconstruction network;
[0017] If it is detected that the pre-training completion condition is met, the initial reconstructed network is determined to be the target reconstructed network.
[0018] Optionally, inputting the training missing point cloud data into the initial reconstruction network to obtain target data includes:
[0019] Inputting the training missing point cloud data into the initial encoding network in the initial reconstruction network to obtain the input features;
[0020] Inputting the input features into an initial decoding network in the initial reconstruction network to obtain the reconstructed point cloud data;
[0021] Correspondingly, the use of the contrastive learning loss value and the reconstruction loss value to adjust the parameters of the initial reconstruction network includes:
[0022] Generate a first loss value using the contrastive learning loss value and the reconstruction loss value;
[0023] The first loss value is used to adjust parameters of the initial reconstruction network.
[0024] Optionally, the initial encoding network includes several feature extraction blocks, each of which includes a multi-layer perceptron and a downsampling layer based on farthest point sampling; the initial decoding network includes multiple multi-layer perceptrons and multiple upsampling layers.
[0025] Optionally, obtaining missing point cloud data for training includes:
[0026] Acquire a number of original missing point clouds as the original point cloud data;
[0027] Different degrees of missing point processing are performed on each of the original missing point clouds to obtain the training missing point cloud data; the missing point processing is cropping processing.
[0028] Optionally, the initial generation network includes a missing point cloud generation network and a correction network, and the generation process of the training repair point cloud data includes:
[0029] Inputting the input features into the missing point cloud generation network to obtain the missing point cloud data;
[0030] Inputting the missing point cloud data and the output data of the target reconstruction network into the correction network to obtain the training repaired point cloud data;
[0031] The missing point cloud generation network includes a missing point cloud modulation module and a folding decoding module, and the inputting the input features into the missing point cloud generation network to obtain the missing point cloud data includes:
[0032] Inputting the input features into the missing point cloud modulation module to obtain missing point cloud features;
[0033] The missing point cloud features and the input features are input into the folded decoding module to obtain the missing point cloud data.
[0034] Optionally, adjusting the parameters of the initial model based on the original point cloud data corresponding to the training repaired point cloud data and the training missing point cloud data includes:
[0035] Using the training repaired point cloud data and the original point cloud data to obtain a corrected reconstruction loss value;
[0036] Obtaining a missing reconstruction loss value using the missing point cloud data and the missing point cloud true value data;
[0037] generating a second loss value using the modified reconstruction loss value and the missing reconstruction loss value;
[0038] Using the second loss value to adjust parameters of the initial model;
[0039] The missing point cloud true value data is the difference data between the training missing point cloud data and the corresponding original point cloud data.
[0040] This application also provides a point cloud missing completion method, including:
[0041] Obtain the point cloud data to be completed;
[0042] The point cloud data to be completed is input into the above-mentioned point cloud completion model to obtain processed point cloud data.
[0043] The present application also provides a model training device, comprising:
[0044] The first acquisition module is used to acquire missing point cloud data for training;
[0045] A training module, used for inputting the training missing point cloud data into an initial model to obtain training repair point cloud data, and adjusting the parameters of the initial model based on the training repair point cloud data and the original point cloud data corresponding to the training missing point cloud data;
[0046] A determination module, configured to determine that the initial model is a point cloud completion model if it is detected that a training completion condition is met;
[0047] Among them, the initial model includes a target reconstruction network and an initial generation network, the target reconstruction network includes a target encoding network, the target encoding network uses the training missing point cloud data for comparative learning, the training missing point cloud data is input into the target encoding network to obtain input features, the input features are input into the initial generation network to obtain missing point cloud data, and the missing point cloud data is used to generate the training repair point cloud data.
[0048] The present application also provides a point cloud missing completion device, comprising:
[0049] The second acquisition module is used to acquire the point cloud data to be completed;
[0050] The completion processing module is used to input the point cloud data to be completed into the above-mentioned point cloud completion model to obtain processed point cloud data.
[0051] The present application also provides an electronic device, including a memory and a processor, wherein:
[0052] The memory is used to store the computer program;
[0053] The processor is used to execute the computer program to implement the above-mentioned model training method and / or the above-mentioned point cloud missing completion method.
[0054] The present application also provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the above-mentioned model training method and / or the above-mentioned point cloud missing completion method.
[0055] The model training method provided in the present application obtains training missing point cloud data; inputs the training missing point cloud data into the initial model to obtain training repair point cloud data, and adjusts the parameters of the initial model based on the training repair point cloud data and the original point cloud data corresponding to the training missing point cloud data; if it is detected that the training completion conditions are met, the initial model is determined to be a point cloud completion model; wherein the initial model includes a target reconstruction network and an initial generation network, the target reconstruction network includes a target encoding network, the target encoding network uses the training missing point cloud data for comparative learning, the training missing point cloud data is input into the target encoding network to obtain input features, the input features are input into the initial generation network to obtain missing point cloud data, and the missing point cloud data is used to generate training repair point cloud data.
[0056] It can be seen that in this method, the initial model includes a target reconstruction network and an initial generation network, wherein the target reconstruction network can take a certain training missing point cloud data as an anchor point and learn the global structure from the perspective of other training missing point cloud data with different missing conditions. That is, several training missing point cloud data corresponding to the same original point cloud data have the same global structure, but due to the different missing parts, they have limited and different receptive fields. Based on the training method of comparison learning, the global structure of the point cloud learned by the network can contain information from different local areas, thereby enabling more accurate feature extraction. The initial generation network is used to generate missing point cloud data, which infers the missing point cloud part lost by the training missing point cloud data based on the input features corresponding to the training missing point cloud data. During training, the missing point cloud features are extracted according to the input feature learning. When the initial model meets the training completion conditions, it is determined as a point cloud completion model. The point cloud completion model can obtain a global structure with local area information, and accurately predict the missing point cloud according to the input data, thereby improving the accuracy of the processed point cloud data after completion processing, and solving the problem of low data accuracy in related technologies.
[0057] In addition, the present application also provides a point cloud missing completion method, a model training device, a point cloud missing completion device, an electronic device and a computer-readable storage medium, which also have the above-mentioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0059] Figure 1 A flow chart of a model training method provided in an embodiment of the present application;
[0060] Figure 2 A specific point cloud completion model structure diagram provided in an embodiment of the present application;
[0061] Figure 3 A schematic diagram of the structure of a model training device provided in an embodiment of the present application;
[0062] Figure 4 A schematic diagram of the structure of a point cloud missing completion device provided in an embodiment of the present application;
[0063] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0065] Please refer to Figure 1 , Figure 1 A flow chart of a model training method provided in an embodiment of the present application. The method includes:
[0066] S101: Obtain missing point cloud data for training.
[0067] Training missing point cloud data refers to incomplete three-dimensional point cloud data used for model training. Each training missing point cloud data corresponds to an original point cloud data. The original point cloud data can be used as label data for the training process, and the loss value can be calculated using it so that the trained model can recognize the difference between the two, and then learn the ability to predict the missing parts of the incomplete three-dimensional point cloud.
[0068] Regarding the method of obtaining missing training point cloud data, in one embodiment, the missing training point cloud data and its corresponding original point cloud data can be obtained from an existing data set. The original point cloud data is usually point cloud truth data, that is, the complete point cloud data of a certain object. In practical applications, point cloud truth data is difficult to obtain, the quantity is small, and it is usually not accurate enough. Therefore, after using point cloud truth data as the original point cloud data to train the model, the result obtained by the model for point cloud completion has a certain difference from the actual result. In order to solve the above problem, in another embodiment, three-dimensional point cloud data with missing data can be used as the original point cloud data, and further missing processing can be performed on it to obtain training missing point cloud data. Specifically, this method can include the following steps:
[0069] Step 11: Get several original missing point clouds.
[0070] Step 12: Perform different degrees of missing point processing on each original missing point cloud to obtain training missing point cloud data.
[0071] The original missing point cloud refers to incomplete three-dimensional point cloud data as the original point cloud data, and this embodiment does not limit its specific number. Missing processing refers to processing that causes incomplete three-dimensional point cloud data, which can be specifically a cropping process. Cropping processing is to select and delete part of the content on the original missing point cloud. It is understandable that the cropping process causes the loss of part of the content in the original missing point cloud, resulting in training missing point cloud data with a higher degree of damage than the original missing point cloud. It is understandable that after an original missing point cloud has undergone different degrees of missing processing, a corresponding plurality of training missing point cloud data can be obtained, and the specific form of each training missing point cloud data is related to the degree of missing processing.
[0072] It is understandable that while the missing processing generates the training missing point cloud data, the missing content of the training missing point cloud data and the original missing point cloud can be clarified, and the missing content can be referred to as missing true value data. In this application, the initial model can not only extract features from the incomplete three-dimensional point cloud and learn its global structure, but also predict the missing point cloud (i.e., the missing part) to obtain the predicted missing point cloud. Therefore, in one embodiment, the missing true value data can also be used as label data for a certain part of the initial model during training so that the model can make accurate predictions.
[0073] S102: Inputting the training missing point cloud data into an initial model to obtain training repair point cloud data, and adjusting the parameters of the initial model based on the training repair point cloud data and the original point cloud data corresponding to the training missing point cloud data.
[0074] S103: If it is detected that the training completion condition is met, the initial model is determined to be a point cloud completion model.
[0075] The specific content and form of the training completion condition are not limited, and may be, for example, a training round condition, a training time condition, a model accuracy condition, or any other optional condition.
[0076] It should be noted that the initial model includes two parts: a target reconstruction network and an initial generation network. The target reconstruction network refers to a network that is at least used to extract features from the training missing point cloud data. In addition, under normal circumstances, the target reconstruction network can also reconstruct data based on the extracted features, and remove noise from the training missing point cloud data by data reconstruction. Among them, the target reconstruction network includes a target encoding network, and the target encoding network refers to a network used for feature extraction. It can be understood that if the target reconstruction network does not perform the data reconstruction step, the target reconstruction network is the target encoding network. The initial generation network refers to a network that generates missing point cloud data and uses it to generate training repair point cloud data.
[0077] In the related art, the data part and the label part of the training data used in model training correspond one to one. In this case, the model can only learn the global structure of the training data from the perspective of the overall situation of the data part, and extract features based on the global structure. The acquisition of this global structure depends on the degree of missingness of the data part compared to the label part, so it is usually not accurate enough. In order to obtain a better global structure, and then obtain input features that can more accurately reflect the missing training point cloud data.
[0078] Specifically, in the present application, the target encoding network uses certain training missing point cloud data as anchor point clouds for comparative learning. Anchor point cloud refers to a point cloud used as a learning benchmark for comparative learning. The training missing point cloud data corresponding to the same original point cloud data as the anchor point cloud is a positive sample, and the training missing point cloud data corresponding to other original point cloud data is a negative sample. During the training process, after the training missing point cloud data is input into the target encoding network, the corresponding input features can be obtained. After the input features are input into the initial generation network, the missing point cloud data is obtained, and the missing point cloud data is used to generate training repair point cloud data.
[0079] In one implementation, in order to improve the convergence speed of the initial model, the initial model can be constructed using a pre-trained target reconstruction network. Pre-training will basically determine the parameters of the target reconstruction network. When training the initial model, it is only necessary to fine-tune it on the existing basis and adjust the parameters of the initial generation network at the same time. Specifically, the generation process of the initial model includes:
[0080] Step 21: Use the training missing point cloud data to perform comparative learning and training on the initial reconstruction network to obtain the target reconstruction network.
[0081] Step 22: Use the target reconstruction network and the initial generation network to combine to obtain the initial model.
[0082] The initial reconstruction network refers to an untrained reconstruction network. It can be pre-trained using the missing point cloud data. The pre-training is also comparative learning training to obtain the target reconstruction network. The initial model can be obtained by combining the target reconstruction network with the initial generation network. When the initial model is subsequently trained, since the target reconstruction network has been pre-trained and has basically reached convergence, compared with the solution of using an untrained initial reconstruction network as the target reconstruction network and forming an initial model, the pre-trained target initial model can reach convergence faster.
[0083] Specifically, the process of using the training missing point cloud data to compare and learn the initial reconstruction network to obtain the target reconstruction network may include the following steps:
[0084] Step 31: Determine the anchor point cloud from the training missing point cloud data.
[0085] Step 32: Based on the anchor point cloud, the training missing point cloud data is input into the initial reconstruction network to obtain the target data.
[0086] Step 33: Use the input features to obtain the contrastive learning loss value, use the reconstructed point cloud data to obtain the reconstruction loss value, and use the contrastive learning loss value and the reconstruction loss value to adjust the parameters of the initial reconstruction network.
[0087] Step 34: If it is detected that the pre-training completion condition is met, the initial reconstructed network is determined as the target reconstructed network.
[0088] In this embodiment, the initial reconstruction network not only extracts features from the input training missing point cloud data, but also reconstructs data based on the extracted features, so as to remove noise from the training missing point cloud data. Therefore, the target data includes input features and reconstructed point cloud data. Input features refer to features obtained after feature extraction of the input training missing point cloud data; reconstructed point cloud data refers to reconstructed data obtained after data reconstruction using the input features.
[0089] In this embodiment, P in Represents the original point cloud data, using S in Represents a collection of original point cloud data, using S S Represents a set of training missing point cloud data. When training the initial reconstruction network, any training missing point cloud data can be selected as the anchor point cloud P S , and S S Zhong and P SThe corresponding training missing point cloud data is used as positive samples, and other training missing point cloud data is used as negative samples. For example, if the original point cloud data corresponding to the anchor point cloud is the airplane point cloud, then S S The training missing point cloud data corresponding to the airplane (i.e., each airplane point cloud with missing points obtained from the airplane point cloud) is a positive sample, and the training missing point cloud data not corresponding to the airplane (e.g., a chair point cloud with missing points) is a negative sample. It can be understood that when a positive sample or a negative sample is input into the initial reconstruction network, the corresponding sample type (i.e., a positive sample or a negative sample) needs to be declared.
[0090] After obtaining the target data, the corresponding loss values are calculated using the input features and the reconstructed point cloud data respectively. Specifically, the contrastive learning loss value is obtained using the input features. The contrastive learning loss value refers to the loss value used to adjust the parameters of the feature extraction part; the reconstruction loss value is calculated using the reconstructed point cloud data. The reconstruction loss value refers to the loss value used to adjust the parameters of the data reconstruction part. After obtaining the above two loss values, use them to adjust the parameters of the initial reconstruction network, and determine the initial reconstruction network as the target reconstruction network when the pre-training completion conditions are met. Among them, the specific content and form of the pre-training completion conditions are not limited, for example, it can be a training round condition, or it can be a training time condition, or it can be any other optional condition.
[0091] Specifically, the process of inputting the training missing point cloud data into the initial reconstruction network to obtain the target data may include the following steps:
[0092] Step 41: Input the training missing point cloud data into the initial encoding network in the initial reconstruction network to obtain input features.
[0093] Step 42: Input the input features into the initial decoding network in the initial reconstruction network to obtain reconstructed point cloud data.
[0094] Accordingly, the process of adjusting the parameters of the initial reconstruction network using the contrastive learning loss value and the reconstruction loss value may include the following steps:
[0095] Step 43: Generate a first loss value using the contrastive learning loss value and the reconstruction loss value.
[0096] Step 44: Use the first loss value to adjust parameters of the initial reconstruction network.
[0097] In this embodiment, the initial reconstruction network includes an initial encoding network and an initial decoding network. The target encoding network is used to extract features from the training missing point cloud data to obtain input features. The initial decoding network is used to decode the input features to complete data reconstruction and obtain reconstructed point cloud data. After integrating the contrastive learning loss value and the reconstruction loss value, a first loss value can be obtained, and then the first loss value is used to adjust the parameters of the entire initial reconstruction network. This embodiment does not limit the specific generation method of the first loss value. For example, in one implementation, the two can be added to obtain the first loss value.
[0098] In one embodiment, the initial encoding network can use PointNet++ (a network structure for processing point clouds) as the basic framework, which includes several feature extraction blocks, each of which contains an MLP (Multi-layer Perceptron) and a downsampling layer. MLP is used to optimize the extracted point cloud features, and downsamples the point cloud using a downsampling layer based on FPS (Farthest Point Sampling), and obtains point clouds of multiple resolutions from fine to coarse, thereby learning local features of point clouds at multiple scales, and finally uses the pooling layer in the initial encoding network for pooling processing to obtain the global features of the point cloud. The initial decoding network includes multiple MLPs for feature dimension transformation and uses multiple upsampling layers for upsampling, which can iteratively reconstruct the shape of the input point cloud. The initial encoding network and the initial decoding network of this structure cooperate to better remove noise from the input point cloud and optimize the shape of the input point cloud.
[0099] It is understandable that for one original point cloud data, there are multiple training missing point cloud data with different missing situations, and each training missing point cloud data has the same global structure. However, different training missing point cloud data, as different local parts of the same original point cloud data, have a limited receptive field. Using contrastive learning to train the initial encoding network can make the global structure of the point cloud learned by the initial encoding network contain information from different local areas.
[0100] Let’s take an example to illustrate the above process: input the missing training point cloud data of the category “aircraft” into the initial reconstruction network. The initial encoding network can obtain the local detail features representing each part of the aircraft, as well as the global structural features representing the whole, that is, the global structure. In the same way, the global structure of the positive and negative samples of contrastive learning is obtained as the input of the initial decoding network to obtain the reconstructed point cloud data. Then, the contrastive learning loss and reconstruction loss are minimized, and the network parameters are updated to continuously optimize the local and global features extracted from the input point cloud. Specifically, it can be used Represents the contrastive learning loss value, using L inRepresents the reconstruction loss value, and uses InfoNCE loss as the loss function of the contrastive learning loss value. The specific calculation formula is:
[0101]
[0102] Among them, v represents the characteristics of the anchor point cloud, v+ represents the input characteristics of the positive sample, and v- represents the input characteristics of the negative sample. Represents the set of input features of all positive samples, Represents the set of input features of all negative samples, and τ is a constant.
[0103] At the same time, you can use:
[0104]
[0105] Calculate the reconstruction loss, where S 1 To reconstruct point cloud data, S 2 To train the original point cloud data corresponding to the missing point cloud data, x and y represent the points therein.
[0106] Based on the above embodiment, in a feasible implementation, the initial generation network can directly splice the generated missing point cloud data and the training missing point cloud data (or the reconstructed point cloud data obtained through reconstruction) to obtain the training repair point cloud data. In another implementation, the data obtained by direct splicing can be rough three-dimensional point cloud data, and the initial generation network can further optimize the rough three-dimensional point cloud data to obtain the training repair point cloud data. Specifically, the initial generation network includes a missing point cloud generation network and a correction network, and the generation process of the training repair point cloud data can include the following steps:
[0107] Step 51: Input the input features into the missing point cloud generation network to obtain the missing point cloud data.
[0108] Step 52: Input the missing point cloud data and the output data of the target reconstruction network into the correction network to obtain the training repair point cloud data.
[0109] Among them, the missing point cloud generation network refers to a network used to generate corresponding missing point cloud data according to input features. The correction network refers to a network that corrects the shape of the output data (which can be unreconstructed training missing point cloud data or reconstructed reconstructed point cloud data). The specific structures of the missing point cloud generation network and the correction network are not limited and can be set as needed.
[0110] For example, in one embodiment, the missing point cloud generation network includes a missing point cloud modulation module and a folding decoding module, and the process of inputting the input features into the missing point cloud generation network to obtain the missing point cloud data may include the following steps:
[0111] Step 53: Input the input features into the missing point cloud modulation module to obtain the missing point cloud features.
[0112] Step 54: Input the missing point cloud features and the input features into the folded decoding module to obtain the missing point cloud data.
[0113] Specifically, the missing point cloud generation network includes multiple decoding modules, each of which includes a missing point cloud modulation module and a folding-based decoding layer (i.e., a folded decoding module). The missing point cloud modulation module transforms the input features through an MLP as the learned missing point cloud features. Based on the folded decoding layer, the randomly sampled two-dimensional grid, the learned missing point cloud features and the input features are processed to obtain the missing point cloud data. By increasing the density of the two-dimensional grid layer by layer, a higher resolution missing point cloud can be predicted.
[0114] This embodiment does not limit the specific process of obtaining the training repair point cloud data by the correction network. In one embodiment, after the correction network fuses the reconstructed point cloud data and the missing point cloud data, the rough three-dimensional point cloud is obtained by FPS sampling. The correction network includes multiple MLPs and a correction layer based on folding. For the input rough three-dimensional point cloud, after multiple MLP processing, point cloud features can be obtained, and then a two-dimensional grid is randomly sampled from a two-dimensional plane of a fixed size. The sampled two-dimensional grid, point cloud features, and three-dimensional coordinates of the point cloud are input into the correction layer based on folding, and the rough three-dimensional point cloud is optimized to obtain the training repair point cloud data.
[0115] It can be understood that in the presence of a correction network, the process of adjusting the parameters of the initial model based on the training repair point cloud data and the training missing point cloud data may include the following steps:
[0116] Step 61: Use the training repaired point cloud data and the original point cloud data to obtain a corrected reconstruction loss value.
[0117] Step 62: Obtain the missing reconstruction loss value using the missing point cloud data and the missing point cloud true value data.
[0118] Step 63: Generate a second loss value using the corrected reconstruction loss value and the missing reconstruction loss value.
[0119] Step 64: Use the second loss value to adjust the parameters of the initial model, where the missing true value data is the difference data between the training missing point cloud data and the corresponding original point cloud data. r Represents the corrected reconstruction loss value, using L c Represents the missing reconstruction loss value. L r and L c The calculation method is the same as L in same.
[0120] Please refer to Figure 2 , Figure 2 A specific point cloud completion model structure diagram provided for an embodiment of the present application. Among them, the incomplete three-dimensional point cloud is the training missing point cloud data or the input point cloud data to be completed when the model is trained and used, the input point cloud reconstruction network based on contrastive learning is the target reconstruction network, the missing point cloud decoding and modulation network is the missing point cloud generation network, and the rough point cloud prediction and correction network is the correction network. Among them, module 1 is the target encoding network (or initial encoding network), which is used for feature encoding, module 2 is the initial decoding network, which is used for fully connected decoding, module 3 is the folded decoding module, which is used for folded decoding, module 4 is the correction network, which is used for rough point cloud correction, and module 5 is the missing point cloud modulation module, which is used to modulate the missing point cloud and generate missing point cloud features. Figure 2 The calculation methods of each loss value in can refer to the above process and will not be repeated here.
[0121] It is understandable that after the model training is completed, it can be used to complete the point cloud data. Therefore, the present application also provides a method for completing missing point cloud data. The method may include the following steps:
[0122] Step 71: Obtain the point cloud data to be completed.
[0123] Step 72: Input the point cloud data to be completed into the point cloud completion model as described above to obtain processed point cloud data.
[0124] The model training method provided by the embodiment of the present application is applied, and the initial model includes a target reconstruction network and an initial generation network, wherein the target reconstruction network can use the original point cloud data as an anchor point to learn the global structure from the perspective of the training missing point cloud data with different missing conditions. That is, a number of training missing point cloud data corresponding to the same original point cloud data have the same global structure, but due to the different missing parts, they have limited and different receptive fields. Based on the training method of comparison learning, the global structure of the point cloud learned by the network can contain information from different local areas, thereby enabling more accurate feature extraction. The initial generation network is used to generate missing point cloud data, which infers the missing point cloud part lost by the training missing point cloud data based on the input features corresponding to the training missing point cloud data. During training, the missing point cloud features are extracted according to the input feature learning. When the initial model meets the training completion conditions, it is determined as a point cloud completion model. The point cloud completion model can obtain a global structure with local area information, and accurately predict the missing point cloud according to the input data, thereby improving the accuracy of the processed point cloud data after completion processing, and solving the problem of low data accuracy in the related technology.
[0125] The model training device provided in an embodiment of the present application is introduced below. The model training device described below and the model training method described above can be referenced to each other.
[0126] Please refer to Figure 3 , Figure 3 A schematic diagram of the structure of a model training device provided in an embodiment of the present application includes:
[0127] A first acquisition module 110 is used to acquire missing point cloud data for training;
[0128] The training module 120 is used to input the training missing point cloud data into the initial model to obtain the training repair point cloud data, and adjust the parameters of the initial model based on the training repair point cloud data and the original point cloud data corresponding to the training missing point cloud data;
[0129] A determination module 130, configured to determine that the initial model is a point cloud completion model if it is detected that the training completion condition is met;
[0130] Among them, the initial model includes a target reconstruction network and an initial generation network. The target reconstruction network includes a target encoding network. The target encoding network uses the training missing point cloud data for comparative learning. The training missing point cloud data is input into the target encoding network to obtain input features. The input features are input into the initial generation network to obtain missing point cloud data. The missing point cloud data is used to generate training repair point cloud data.
[0131] Optionally include:
[0132] The pre-training module is used to train the initial reconstruction network using the training missing point cloud data to obtain the target reconstruction network;
[0133] The combination module is used to obtain an initial model by combining the target reconstruction network with the initial generation network.
[0134] Optionally, pre-training modules include:
[0135] An anchor point determination unit, used to determine the anchor point cloud from the training missing point cloud data;
[0136] An input unit is used to input the training missing point cloud data into the initial reconstruction network based on the anchor point cloud to obtain target data; wherein the target data includes input features and reconstructed point cloud data;
[0137] A parameter adjustment unit, used to obtain a contrastive learning loss value using input features, obtain a reconstruction loss value using reconstructed point cloud data, and adjust parameters of an initial reconstruction network using the contrastive learning loss value and the reconstruction loss value;
[0138] The target reconstruction network determination unit is used to determine the initial reconstruction network as the target reconstruction network if it is detected that the pre-training completion condition is met.
[0139] Optionally, the input unit includes:
[0140] A feature acquisition subunit is used to input the training missing point cloud data into the initial encoding network in the initial reconstruction network to obtain input features;
[0141] A reconstruction subunit, used to input the input features into the initial decoding network in the initial reconstruction network to obtain reconstructed point cloud data;
[0142] Correspondingly, the parameter adjustment unit includes:
[0143] A first loss generating subunit, used for generating a first loss value by using a contrastive learning loss value and a reconstruction loss value;
[0144] The initial reconstruction network adjustment subunit is used to adjust the parameters of the initial reconstruction network using the first loss value.
[0145] Optionally, the initial encoding network includes several feature extraction blocks, each feature extraction block includes a multi-layer perceptron and a downsampling layer based on farthest point sampling; the initial decoding network includes multiple multi-layer perceptrons and multiple upsampling layers.
[0146] Optionally, the first acquisition module 110 includes:
[0147] An original missing point acquisition unit, used for acquiring a number of original missing point clouds as original point cloud data;
[0148] The missing point processing unit is used to perform missing point processing to different degrees on each original missing point cloud to obtain training missing point cloud data; the missing point processing is cropping processing.
[0149] Optionally, the training module 120 includes:
[0150] A missing point cloud generation unit inputs the input features into a missing point cloud generation network to obtain missing point cloud data;
[0151] A correction unit, used for inputting the missing point cloud data and the output data of the target reconstruction network into the correction network to obtain the training repair point cloud data;
[0152] The missing point cloud generation network includes a missing point cloud modulation module and a folding decoding module, and the missing point cloud generation unit includes:
[0153] A missing feature acquisition subunit is used to input the input feature into the missing point cloud modulation module to obtain the missing point cloud feature;
[0154] The folding decoding subunit is used to input the missing point cloud features and the input features into the folding decoding module to obtain the missing point cloud data.
[0155] Optionally, the training module 120 includes:
[0156] A corrected reconstruction loss generating unit, used for obtaining a corrected reconstruction loss value by using the training repaired point cloud data and the original point cloud data;
[0157] A missing reconstruction loss generating unit, used for obtaining a missing reconstruction loss value by using the missing point cloud data and the missing point cloud true value data;
[0158] A second loss generating unit, configured to generate a second loss value by using the corrected reconstruction loss value and the missing reconstruction loss value;
[0159] The initial model adjustment unit is used to adjust the parameters of the initial model using the second loss value.
[0160] The point cloud missing completion device provided in an embodiment of the present application is introduced below. The point cloud missing completion device described below and the point cloud missing completion method described above can be referenced to each other.
[0161] Please refer to Figure 4 , Figure 4 A schematic diagram of the structure of a point cloud missing completion device provided in an embodiment of the present application includes:
[0162] The second acquisition module 210 is used to acquire the point cloud data to be completed;
[0163] The completion processing module 220 is used to input the point cloud data to be completed into the above-mentioned point cloud completion model to obtain processed point cloud data.
[0164] The electronic device provided in the embodiment of the present application is introduced below. The electronic device described below and the model training method described above can be referenced to each other.
[0165] Please refer to Figure 5 , Figure 5 The electronic device 100 may include a processor 101 and a memory 102 , and may further include one or more of a multimedia component 103 , an information input / information output (I / O) interface 104 , and a communication component 105 .
[0166] Among them, the processor 101 is used to control the overall operation of the electronic device 100 to complete all or part of the steps in the above-mentioned model training method; the memory 102 is used to store various types of data to support the operation of the electronic device 100, and these data may include, for example, instructions for any application or method used to operate on the electronic device 100, and application-related data. The memory 102 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), erasable programmable read-only memory (Erasable Programmable Read-Only Memory, EPROM), programmable read-only memory (Programmable Read-Only Memory, PROM), read-only memory (Read-Only Memory, ROM), magnetic memory, flash memory, magnetic disk or optical disk. One or more.
[0167] The multimedia component 103 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone, which is used to receive external audio signals. The received audio signal may be further stored in the memory 102 or sent through the communication component 105. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 104 provides an interface between the processor 101 and other interface modules, and the above-mentioned other interface modules may be keyboards, mice, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 105 is used for wired or wireless communication between the electronic device 100 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G or 4G, or a combination of one or more of them, so the corresponding communication component 105 may include: Wi-Fi components, Bluetooth components, NFC components.
[0168] The electronic device 100 can be implemented by one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), controllers, microcontrollers, microprocessors or other electronic components to execute the model training method given in the above embodiment.
[0169] The computer-readable storage medium provided in the embodiments of the present application is introduced below. The computer-readable storage medium described below and the model training method described above can be referenced to each other.
[0170] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned model training method are implemented.
[0171] The computer-readable storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0172] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0173] Those skilled in the art may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0174] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0175] Finally, it should be noted that, in this article, relationships such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms include, include or any other variations are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0176] Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A model training method, It is characterized in that include: Get missing point cloud data for training; Inputting the training missing point cloud data into an initial model to obtain training repair point cloud data, and adjusting the parameters of the initial model based on the training repair point cloud data and the original point cloud data corresponding to the training missing point cloud data; If it is detected that the training completion condition is met, determining that the initial model is a point cloud completion model; The initial model includes a target reconstruction network and an initial generation network, the target reconstruction network includes a target encoding network that performs comparative learning using training missing point cloud data; the training missing point cloud data is input into the target encoding network to obtain input features, the target reconstruction network also outputs reconstructed point cloud data, the input features are input into the missing point cloud generation network of the initial generation network to obtain missing point cloud data, and the missing point cloud data and the reconstructed point cloud data are input into the correction network of the initial generation network to obtain the training repair point cloud data; The generation process of the initial model includes: Select any missing point cloud data for training as the anchor point cloud; The training missing point cloud data of the same original point cloud data corresponding to the anchor point cloud in the set of training missing point cloud data is taken as a positive sample, and the other training missing point cloud data in the set of training missing point cloud data is taken as a negative sample; Input the sample into the initial encoding network of the initial reconstruction network to obtain input features; input the input features into the initial decoding network in the initial reconstruction network to obtain reconstructed point cloud data; The contrastive learning loss value is obtained by using the input features, the reconstruction loss value is obtained by using the reconstructed point cloud data, and the contrastive learning loss value and the reconstruction loss value are used to adjust the parameters of the initial reconstruction network; If it is detected that the pre-training completion condition is met, the initial reconstructed network is determined as the target reconstructed network; The target reconstruction network is combined with the initial generation network to obtain the initial model.
2. The model training method according to claim 1, It is characterized in that The step of adjusting parameters of the initial reconstruction network by using the contrastive learning loss value and the reconstruction loss value includes: Generate a first loss value using the contrastive learning loss value and the reconstruction loss value; The first loss value is used to adjust parameters of the initial reconstruction network.
3. The model training method according to claim 2, It is characterized in that The initial encoding network includes several feature extraction blocks, each of which includes a multi-layer perceptron and a down-sampling layer based on farthest point sampling; the initial decoding network includes multiple multi-layer perceptrons and multiple up-sampling layers.
4. The model training method according to claim 1, It is characterized in that The step of obtaining missing point cloud data for training includes: Acquire a number of original missing point clouds as the original point cloud data; Different degrees of missing point processing are performed on each of the original missing point clouds to obtain the training missing point cloud data; the missing point processing is cropping processing.
5. The model training method according to claim 1, It is characterized in that The missing point cloud generation network includes a missing point cloud modulation module and a folding decoding module, and the inputting the input features into the missing point cloud generation network to obtain the missing point cloud data includes: Inputting the input features into the missing point cloud modulation module to obtain missing point cloud features; The missing point cloud features and the input features are input into the folded decoding module to obtain the missing point cloud data.
6. The model training method according to claim 5, It is characterized in that The adjusting the parameters of the initial model based on the original point cloud data corresponding to the training repaired point cloud data and the training missing point cloud data includes: Using the training repaired point cloud data and the original point cloud data to obtain a corrected reconstruction loss value; Obtaining a missing reconstruction loss value using the missing point cloud data and the missing point cloud true value data; generating a second loss value using the modified reconstruction loss value and the missing reconstruction loss value; Using the second loss value to adjust parameters of the initial model; The missing point cloud true value data is the difference data between the training missing point cloud data and the corresponding original point cloud data.
7. A method for completing missing point clouds. It is characterized in that include: Obtain the point cloud data to be completed; The point cloud data to be completed is input into the point cloud completion model according to any one of claims 1 to 6 to obtain processed point cloud data.
8. A model training device, It is characterized in that include: The first acquisition module is used to acquire missing point cloud data for training; A training module, used for inputting the training missing point cloud data into an initial model to obtain training repair point cloud data, and adjusting the parameters of the initial model based on the training repair point cloud data and the original point cloud data corresponding to the training missing point cloud data; A determination module, configured to determine that the initial model is a point cloud completion model if it is detected that a training completion condition is met; The initial model includes a target reconstruction network and an initial generation network, the target reconstruction network includes a target encoding network that performs comparative learning using training missing point cloud data; the training missing point cloud data is input into the target encoding network to obtain input features, the target reconstruction network also outputs reconstructed point cloud data, the input features are input into the missing point cloud generation network of the initial generation network to obtain missing point cloud data, and the missing point cloud data and the reconstructed point cloud data are input into the correction network of the initial generation network to obtain training repair point cloud data; A pre-training module is used to select any training missing point cloud data as an anchor point cloud; use the training missing point cloud data of the same original point cloud data corresponding to the anchor point cloud in the set of training missing point cloud data as positive samples, and use other training missing point cloud data in the set of training missing point cloud data as negative samples; input the samples into the initial encoding network of the initial reconstruction network to obtain input features; input the input features into the initial decoding network in the initial reconstruction network to obtain reconstructed point cloud data; use the input features to obtain a contrastive learning loss value, use the reconstructed point cloud data to obtain a reconstruction loss value, and use the contrastive learning loss value and the reconstruction loss value to adjust the parameters of the initial reconstruction network; if it is detected that the pre-training completion condition is met, the initial reconstruction network is determined to be the target reconstruction network; The combination module is used to obtain an initial model by combining the target reconstruction network with the initial generation network.
9. A point cloud missing completion device, It is characterized in that include: The second acquisition module is used to acquire the point cloud data to be completed; The completion processing module is used to input the point cloud data to be completed into the point cloud completion model as described in any one of claims 1 to 6 to obtain processed point cloud data.
10. An electronic device, It is characterized in that comprising a memory and a processor, wherein: The memory is used to store the computer program; The processor is used to execute the computer program to implement the model training method as described in any one of claims 1 to 6, and / or the point cloud missing completion method as described in claim 7.
11. A computer-readable storage medium, It is characterized in that Used to save a computer program, wherein when the computer program is executed by a processor, the model training method according to any one of claims 1 to 6 and / or the point cloud missing completion method according to claim 7 are implemented.
Citation Information
Patent Citations
Deep learning-based point cloud completion method
CN113205104A