Local perception point cloud representation learning pre-training method and device based on attention mechanism

By adopting the locally perceived point cloud representation learning pre-training method based on attention mechanism in point cloud representation learning, the problem of insufficient local feature perception ability of Transformer model is solved, and more efficient point cloud feature extraction and large-scale point cloud data processing is achieved, improving the accuracy of point cloud downstream tasks.

CN119418143BActive Publication Date: 2025-05-09INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD +1

Patent Information

Application Number
CN202510033344.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-09
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

The existing Transformer model has weak local feature perception capabilities in point cloud representation learning and is difficult to effectively process large-scale point cloud data.

Method used

The locally perceived point cloud representation learning pre-training method is adopted based on the attention mechanism. By sampling the point cloud data on the surface of the three-dimensional model, the pre-trained data set is determined, and the converter model structure and the local perceived self-attention mechanism are used to group data, embed and random masks, extract local and global features, and determine the training loss through density-aware chamfer distance loss computer system.

Benefits of technology

The model's ability to extract locally perceived point cloud data features is improved, and large-scale point cloud data can be effectively processed, thereby improving the quality of point cloud feature extraction and the accuracy and accuracy of point cloud downstream classification tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119418143B_ABST
    Figure CN119418143B_ABST
Patent Text Reader

Abstract

The present application discloses a pre-training method and device for learning local perception point cloud representation based on an attention mechanism, and relates to the field of point cloud representation learning technology, including: determining a pre-training data set by sampling point cloud data on the surface of a three-dimensional model; the pre-training data set is a data set of point cloud shape; determining a model to be trained based on a converter model structure and a preset local perception self-attention mechanism, and processing the pre-training data set using a pre-training strategy of the model to be trained and mask modeling; extracting local and global features based on the data processing results and the model to be trained, and determining the training loss using the feature extraction results and the density-aware chamfer distance loss calculation mechanism in the model to be trained to obtain a pre-trained model; collecting point cloud data of a target product to determine a fine-tuning data set, and using the pre-trained model and the fine-tuning data set to perform feature extraction and point cloud classification, completing fine-tuning, and obtaining a target model. The present application improves the quality of point cloud feature extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of point cloud representation learning, and in particular to a local perception point cloud representation learning pre-training method and device based on an attention mechanism. Background Art

[0002] With the rapid development of computing devices and deep learning, point cloud data has gradually become an indispensable part of the perception field, and point cloud representation learning has also received more and more attention. Limited by the difficulty of data annotation and the disorder of point clouds, the existing solution uses the transformer model structure for point cloud perception and point cloud understanding tasks.

[0003] However, due to the weak perception of local features in the pre-trained point cloud model of the Transformer model structure, the performance of local feature extraction, which is an important part of point cloud representation learning, is poor in existing solutions. Although there have been many studies on local self-attention in vision, there are few related studies in point clouds, and due to its spatial continuity, related methods in vision are not directly applicable. At the same time, the quadratic complexity of the Transformer with respect to sequence length limits the ability to apply the Transformer to larger-scale point cloud data. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a local perception point cloud representation learning pre-training method and device based on the attention mechanism, which can effectively improve the model's local perception point cloud data features and the ability to process large-scale point cloud data, thereby improving the situation where the pre-trained point cloud model based on the converter model structure has weak local feature perception, improving the quality of point cloud feature extraction, and further improving the precision and accuracy of processing point cloud downstream classification tasks. The specific scheme is as follows:

[0005] In a first aspect, the present application provides a local perception point cloud representation learning pre-training method based on an attention mechanism, comprising:

[0006] By sampling point cloud data on the surface of the three-dimensional model, a corresponding pre-training data set is determined; the pre-training data set is a data set in the shape of a point cloud;

[0007] Based on the converter model structure and the preset local perception self-attention mechanism, the model to be trained is determined, and the pre-training strategy of the model to be trained and mask modeling is used to group and embed the pre-training data set in sequence, and the embedded point cloud data is randomly masked to obtain the point cloud data processing result;

[0008] Performing local and global feature extraction based on the model to be trained, the pre-training strategy and the point cloud data processing result, and determining the training loss using the obtained feature extraction result and the density-aware chamfer distance loss calculation mechanism in the model to be trained to obtain a pre-trained model;

[0009] The fine-tuning data set is determined by collecting point cloud data of the target product, and the pre-trained model and the fine-tuning data set are used to perform feature extraction and point cloud classification, so as to complete the model fine-tuning operation based on the obtained classification results and obtain the target model.

[0010] Optionally, the step of sampling point cloud data on the surface of the three-dimensional model to determine a corresponding pre-training data set includes:

[0011] Get the preset number of samples;

[0012] Point cloud data sampling is performed on the surface of the three-dimensional model using the preset sampling quantity to determine a corresponding pre-training data set; and no annotation information exists in the pre-training data set.

[0013] Optionally, the pre-training strategy using the model to be trained and mask modeling sequentially groups and embeds the pre-training data set, and randomly masks the embedded point cloud data, including:

[0014] Determine a corresponding number of sampling points from the pre-training data set based on the model to be trained, the farthest point sampling strategy and the preset number of groups, and use the K nearest neighbor algorithm to determine the neighboring points of each sampling point to complete the corresponding point cloud data grouping operation and obtain the grouping result;

[0015] For any point cloud group in the grouping result, the corresponding point cloud coordinates are mapped to the feature space through a combination of a multi-layer linear mapping layer, a splicing layer, and a maximum pooling layer in the model to be trained to complete the point cloud embedding operation and obtain the embedded feature vector of each point cloud group;

[0016] Performing random masking on each of the point cloud groups in the grouping result based on a preset ratio to obtain a corresponding random masking result;

[0017] A point cloud data processing result is determined based on the random mask result and the embedded feature vector of each point cloud group.

[0018] Optionally, the extracting local and global features based on the model to be trained, the pre-training strategy and the point cloud data processing result includes:

[0019] Determine an unmasked point cloud group and a masked point cloud group based on the model to be trained and the random mask result;

[0020] For any data point in the unmasked point cloud group, a neighborhood point corresponding to the current data point is determined based on a K-nearest neighbor algorithm, and a local feature corresponding to the current data point is determined for the current data point and the neighborhood point using a self-attention mechanism in a corresponding local neighborhood area;

[0021] Downsampling the unmasked point cloud group based on a downsampling operator, and determining a global feature corresponding to the current data point and the obtained downsampling result using a self-attention mechanism;

[0022] The local features and the global features corresponding to the current data point are spliced, and the obtained spliced ​​features are weighted using a normalized exponential function to complete a feature fusion operation and obtain a fused feature corresponding to the current data point.

[0023] Optionally, the determining the training loss by using the obtained feature extraction result and the density-aware chamfer distance loss calculation mechanism in the model to be trained includes:

[0024] Obtaining learnable parameters corresponding to feature information of the masked point cloud group;

[0025] Determining point cloud group sequence information according to the grouping result, and using the point cloud group sequence information to splice the learnable parameters and the feature extraction results corresponding to the unmasked point cloud group to obtain a corresponding splicing result;

[0026] Predicting the coordinates of each of the point cloud groups based on the decoder in the model to be trained and the splicing result to obtain corresponding coordinate prediction results;

[0027] Reconstructing the masked point cloud group based on the coordinate prediction result to obtain a corresponding reconstruction result;

[0028] The density-aware chamfer distance loss calculation mechanism in the model to be trained is used to trigger the corresponding point cloud group density information acquisition operation, and the training loss is determined based on the obtained point cloud group density information acquisition result, the reconstructed point coordinates in the reconstruction result, and the real point coordinates corresponding to the masked point cloud group to obtain the pre-trained model.

[0029] Optionally, determining the fine-tuning dataset by collecting point cloud data of the target product includes:

[0030] Collect the point cloud data corresponding to the target product to determine the corresponding collection results;

[0031] Obtain the scaling factor and translation factor corresponding to each coordinate axis in the three-dimensional coordinate axis;

[0032] Processing the point cloud data in the acquisition result based on the scaling factor and the translation factor to obtain corresponding processed data;

[0033] A fine-tuning dataset is constructed based on the processed data.

[0034] Optionally, the using the pre-trained model and the fine-tuning dataset to perform feature extraction and point cloud classification includes:

[0035] Removing the decoder in the pre-trained model, and adding a classification head constructed based on a multi-layer linear mapping layer to the pre-trained model to obtain a corresponding processed model;

[0036] Using the processed model and the fine-tuning data set to perform local and global feature extraction to obtain corresponding target extraction results;

[0037] Based on the processed model, point cloud category mapping is performed on the target extraction result to complete the corresponding point cloud classification operation and obtain the classification result.

[0038] In a second aspect, the present application provides a local perception point cloud representation learning pre-training device based on an attention mechanism, comprising:

[0039] A data set determination module is used to determine a corresponding pre-training data set by sampling point cloud data on the surface of the three-dimensional model; the pre-training data set is a data set in the shape of a point cloud;

[0040] A data processing module is used to determine the model to be trained based on the converter model structure and the preset local perception self-attention mechanism, and to group and embed the pre-trained data set in sequence using the pre-training strategy of the model to be trained and mask modeling, and to randomly mask the embedded point cloud data to obtain a point cloud data processing result;

[0041] A pre-training module, used for performing local and global feature extraction based on the model to be trained, the pre-training strategy and the point cloud data processing result, and determining the training loss using the obtained feature extraction result and the density-aware chamfer distance loss calculation mechanism in the model to be trained, so as to obtain a pre-trained model;

[0042] The model fine-tuning module is used to determine the fine-tuning data set by collecting point cloud data of the target product, and use the pre-trained model and the fine-tuning data set to perform feature extraction and point cloud classification, so as to complete the model fine-tuning operation based on the obtained classification results and obtain the target model.

[0043] It can be seen that in the present application, point cloud data sampling is performed on the surface of the three-dimensional model to determine the corresponding pre-training data set; the pre-training data set is a data set of point cloud shape; the model to be trained is determined based on the converter model structure and the preset local perception self-attention mechanism, and the pre-training strategy of the model to be trained and mask modeling is used to group and embed the data of the pre-training data set in turn, and the embedded point cloud data is randomly masked to obtain the point cloud data processing result; local and global feature extraction is performed based on the model to be trained, the pre-training strategy and the point cloud data processing result, and the training loss is determined using the obtained feature extraction results and the density-aware chamfer distance loss calculation mechanism in the model to be trained to obtain the pre-trained model; the fine-tuning data set is determined by collecting the point cloud data of the target product, and the pre-trained model and the fine-tuning data set are used to perform feature extraction and point cloud classification, so as to complete the model fine-tuning operation based on the obtained classification result and obtain the target model. That is, in this application, firstly, a pre-training data set is obtained, and then the model to be trained is determined by using the preset local perception self-attention mechanism and the converter model structure, and then the pre-training strategy of mask modeling of the model to be trained is used to process the pre-training data set, extract local and global features, and determine the training loss to complete the pre-training, and then the point cloud data of the target product is collected to extract features and classify the point cloud of the obtained pre-trained model to complete fine-tuning and determine the target model. In this way, the model's local perception of point cloud data features and its ability to process large-scale point cloud data can be effectively improved, thereby improving the situation where the pre-trained point cloud model based on the converter model structure has weak perception of local features, improving the quality of point cloud feature extraction, and thereby improving the precision and accuracy of processing point cloud downstream classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0045] Figure 1 A flowchart of a local perception point cloud representation learning pre-training method based on an attention mechanism provided for this application;

[0046] Figure 2 A schematic diagram of a specific point cloud pre-training process provided for this application;

[0047] Figure 3 A specific local and global feature extraction process diagram provided for this application;

[0048] Figure 4 A schematic diagram of the structure of a local perception point cloud representation learning pre-training device based on an attention mechanism provided in this application;

[0049] Figure 5 A structural diagram of an electronic device provided for this application. DETAILED DESCRIPTION

[0050] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0051] However, due to the weak perception of local features in the pre-trained point cloud model of the Transformer model structure, the performance of local feature extraction, which is an important part of point cloud representation learning, is poor in existing schemes. Although there have been many works in vision that study local self-attention, there are few related studies in point clouds, and due to the nature of its spatial continuity, related methods in vision are not directly applicable. At the same time, the quadratic complexity of the Transformer with respect to sequence length limits the ability to apply the Transformer to larger-scale point cloud data. To this end, the present application provides a pre-training scheme for locally perceived point cloud representation learning based on an attention mechanism, which can effectively improve the model's local perception of point cloud data features and its ability to process large-scale point cloud data, thereby improving the situation where the pre-trained point cloud model based on the transformer model structure has weak perception of local features, improving the quality of point cloud feature extraction, and thereby improving the precision and accuracy of processing point cloud downstream classification tasks.

[0052] See also Figure 1 As shown, the embodiment of the present invention discloses a local perception point cloud representation learning pre-training method based on an attention mechanism, comprising:

[0053] Step S11, sampling point cloud data on the surface of the three-dimensional model to determine a corresponding pre-training data set; the pre-training data set is a data set in the shape of a point cloud.

[0054] Specifically, in this embodiment, the pre-training data set is constructed by point cloud data obtained by sampling a certain number of points on the surface of the three-dimensional model, that is, first, a preset sampling number is obtained, and then the point cloud data of the surface of the three-dimensional model is sampled by the preset sampling number to determine the corresponding pre-training data set; the pre-training data set is a point cloud shape data set and there is no annotation information therein.

[0055] Step S12: determine the model to be trained based on the converter model structure and the preset local perception self-attention mechanism, use the model to be trained and the pre-training strategy of mask modeling to group and embed the pre-training data set in sequence, and randomly mask the embedded point cloud data to obtain the point cloud data processing result.

[0056] In this embodiment, after obtaining the pre-training data set, the model to be trained is determined based on the converter model structure and the preset local perception self-attention mechanism, so that the original self-attention mechanism in the converter model structure is replaced by the local perception self-attention mechanism, which is used to extract the global and local features of the point cloud at the same time, so as to enhance the local perception ability of the model. Then, it should be pointed out that the pre-training strategy of mask modeling is a rule related to model pre-training. In the process of pre-training the model to be trained by using this rule and the point cloud data in the pre-training data set, it is necessary not only to complete the grouping and embedding of the data, but also to perform a preset proportion of random masking on the point cloud data after the embedding process is completed, so as to divide the grouped point cloud groups into visible parts and invisible parts, and take corresponding processing measures for different point cloud groups according to these two types.

[0057] Specific, combined Figure 2 As shown in FIG. 1 , the specific steps of grouping, embedding and randomly masking the pre-training data set are as follows: first, based on the model to be trained, the farthest point sampling strategy and the preset grouping number, a corresponding number of points are sampled from the pre-training data set and determined as sampling points (i.e. Figure 2 The center point in the sample point is then determined by using the K nearest neighbor algorithm to complete the corresponding point cloud data grouping operation and obtain the grouping result; after the grouping is completed, for any point cloud group in the grouping result (i.e. Figure 2 The point cloud blocks in the training model are combined with the multi-layer linear mapping layer, the concatenation layer and the maximum pooling layer (i.e. Figure 2 The point cloud embedding module in the method maps the corresponding point cloud coordinates to the feature space to complete the point cloud embedding operation and obtain the embedding feature vector of each point cloud group; randomly masks each point cloud group in the grouping result based on a preset ratio to obtain a corresponding random masking result; and determines the point cloud data processing result based on the random masking result and the embedding feature vector of each point cloud group.

[0058] Step S13: perform local and global feature extraction based on the model to be trained, the pre-training strategy and the point cloud data processing results, and determine the training loss using the obtained feature extraction results and the density-aware chamfer distance loss calculation mechanism in the model to be trained to obtain the pre-trained model.

[0059] In this embodiment, the Transformer model structure is used to extract features from point cloud data. The embedded feature vector of the point cloud group obtained after grouping and embedding by the point cloud encoding module is used to extract local and global features using the attention mechanism of the Transformer in the model to be trained. That is, the encoder is used to extract features from the visible part of the point cloud group. Specifically, combined with Figure 3 As shown, based on the model to be trained and the random mask result, an unmasked (visible) point cloud group and a masked (invisible) point cloud group are determined; in the local feature extraction branch, for any data point in the unmasked point cloud group (i.e. Figure 3 ), based on the K nearest neighbor algorithm, determine the neighborhood points corresponding to the current data point as key-value pairs, and use the self-attention mechanism in the model to determine the local feature similarity matrix corresponding to the current data point for the current data point and the neighborhood points in the corresponding local neighborhood area, and then add a mask to the matrix to eliminate the influence of the features of the invisible data points on the matrix, so as to determine the local features corresponding to the current data point in the end; in the global feature extraction branch, downsample all data points in the unmasked point cloud group based on the downsampling operator to form key-value pairs, and use the self-attention mechanism to determine the global feature similarity matrix corresponding to the current data point and the obtained downsampling result, and then add a mask to the matrix to eliminate the influence of the features of the invisible data points on the matrix, so as to determine the global features corresponding to the current data point in the end.

[0060] Furthermore, in this embodiment, after obtaining the features of the two branches, feature fusion needs to be performed. Taking the current data point as an example, specifically, the local features and the global features corresponding to the current data point are spliced, and the normalized exponential function (i.e. Figure 3 The Softmax in the above example is used to perform weighted processing on the concatenated features to complete the feature fusion operation and obtain the fused features corresponding to the current data point. Among them, regarding the weighted processing, a Softmax is used to normalize the concatenated features to obtain the affinity matrix of the current data point; then the affinity matrix is ​​split into a local affinity matrix and a global affinity matrix, respectively. The values ​​in the affinity matrix are the weights of the query points to the corresponding values, and the feature values ​​are weighted using the weights to obtain new query point features. It should be pointed out that the weights here are used for the point features in the feature map and will not be used in the similarity matrix. For example, the value in the local affinity matrix is ​​the weighted value of the query point to its neighborhood point features, that is, all domain point features are multiplied by their corresponding weights and all weighted neighborhood point features are summed to obtain the new query point feature value. The same is true for the global affinity matrix.

[0061] It should be understood that the application of the downsampling operator in the global feature extraction branch can be specifically to group the features of all data points in the unmasked point cloud group using the farthest point sampling and K nearest neighbor grouping methods to obtain the grouped features of the downsampled number, map the grouped features using a linear layer, then take the average, and then regularize the features to obtain the downsampled point cloud group features, that is, the downsampling result.

[0062] In this embodiment, in addition to extracting local and global features from the visible point cloud group, a set of learnable parameters will be used to represent the features of the invisible point cloud group. Afterwards, the features of the invisible point cloud group are spliced ​​back to the features of the visible point cloud group in the order of the original point cloud group, and the point cloud is reconstructed based on the spliced ​​complete point cloud features, and then the training loss is calculated to determine whether the pre-training is completed. Specifically, it may include: obtaining learnable parameters corresponding to the feature information of the masked point cloud group; determining the point cloud group order information according to the grouping result, and using the point cloud group order information to splice the learnable parameters and the feature extraction results corresponding to the unmasked point cloud group to obtain a corresponding splicing result; predicting the coordinates of each point cloud group based on the decoder in the model to be trained and the splicing result to obtain a corresponding coordinate prediction result; reconstructing the masked point cloud group based on the coordinate prediction result to obtain a corresponding reconstruction result; using the density-aware chamfer distance loss calculation mechanism in the model to be trained to trigger the corresponding point cloud group density information acquisition operation, and determining the training loss based on the obtained point cloud group density information acquisition result, the reconstructed point coordinates in the reconstruction result, and the real point coordinates corresponding to the masked point cloud group to obtain a pre-trained model. That is to say, Figure 2 As shown in the figure, after the splicing is completed, the obtained complete point cloud features are input into the decoder in the model, and then the decoder uses a single-layer prediction head to predict the coordinates of the data points in each point cloud group, including the point cloud group of the invisible part, to reconstruct the point cloud of the invisible part according to the coordinates in the prediction results. Then the training loss is determined based on the density-aware chamfer distance loss of the reconstructed point coordinates of the invisible part and the corresponding real point coordinates in the real point cloud.

[0063] It is further understood that in a specific embodiment, regarding the application of density-aware chamfer distance loss: on the basis of the original chamfer distance loss that was replaced in the model, that is, on the basis of calculating the average shortest distance from each point in two point sets to the nearest neighbor point in another point set, a point cloud density metric is added, that is, for high point cloud density areas, its weight will be lower, and for low point cloud density areas, its weight will be higher. Through the perception of density, the model can obtain local information of the point cloud, thereby having stronger local perception capabilities.

[0064] Step S14: determine a fine-tuning data set by collecting point cloud data of the target product, and use the pre-trained model and the fine-tuning data set to perform feature extraction and point cloud classification, so as to complete the model fine-tuning operation based on the obtained classification results and obtain the target model.

[0065] Combination Figure 2 As shown, in this embodiment, after completing the pre-training, in order to accurately apply it to the processing of 3D (Three-Dimensional) tasks (point cloud classification, point cloud segmentation, three-dimensional target detection and other tasks) downstream of the point cloud, it is also necessary to fine-tune the pre-trained model according to the task information.

[0066] In a specific embodiment, taking the point cloud classification task as an example, first, it is necessary to construct a fine-tuning data set, collect the point cloud data corresponding to the target product, and determine the corresponding collection results; then obtain the scaling coefficients and translation coefficients corresponding to each coordinate axis in the three-dimensional coordinate axis, and process the point cloud data in the collection result based on the scaling coefficients and the translation coefficients to obtain the corresponding processed data; then construct a fine-tuning data set based on the processed data. Among them, regarding scaling and translation, three scaling coefficients and three translation coefficients can be randomly sampled on the three coordinate axes x, y, and z respectively; for each point cloud data, the three coordinates of the point cloud are multiplied by the scaling coefficient and the translation coefficient is added. In this way, the diversity of point cloud data is increased through the data enhancement method to prevent overfitting during the model fine-tuning process. Then, it is necessary to perform some processing on the pre-trained model, specifically, it is necessary to remove the decoder in the pre-trained model, and add a classification head constructed based on a multi-layer linear mapping layer to the pre-trained model to obtain the corresponding processed model. After that, the processed model and the local and global feature extraction of the fine-tuned data set can be used as described in the previous steps to obtain the corresponding target extraction result, and the target extraction result is mapped to a point cloud category based on the processed model to map the features to point cloud categories, thereby achieving the function of classifying the point cloud and obtaining the classification result. In this way, by fine-tuning the pre-trained model using the three-dimensional point cloud data of the target product, a point cloud classification inference model for the target product scene can be obtained for online inference.

[0067] In summary, the training method described in this embodiment has the following beneficial effects:

[0068] 1) It can improve the local perception ability of the model. This embodiment uses the design of the self-attention mechanism to explicitly model the local geometry, which can improve the model's perception ability of local geometry while maintaining the original global feature perception.

[0069] 2) Only a small amount of calculation and parameters are added. This embodiment uses a downsampling operator in the global feature extraction branch to limit the increase in the amount of calculation of the model during the calculation process, while only adding a small amount of parameters. This ensures that the overall amount of calculation and parameters of the model will not increase significantly, making it easier to apply the model to larger-scale point cloud data.

[0070] 3) Improve the accuracy of 3D classification task results. This embodiment improves the model's ability to extract features of local point cloud geometry through the local-aware self-attention mechanism and density-aware chamfer distance loss design, allowing the model to extract better features for point cloud downstream tasks in actual scenarios, thereby improving the processing accuracy of task results.

[0071] It can be seen that in the present application, point cloud data sampling is performed on the surface of the three-dimensional model to determine the corresponding pre-training data set; the pre-training data set is a data set of point cloud shape; the model to be trained is determined based on the converter model structure and the preset local perception self-attention mechanism, and the pre-training strategy of the model to be trained and mask modeling is used to group and embed the data of the pre-training data set in turn, and the embedded point cloud data is randomly masked to obtain the point cloud data processing result; local and global feature extraction is performed based on the model to be trained, the pre-training strategy and the point cloud data processing result, and the training loss is determined using the obtained feature extraction results and the density-aware chamfer distance loss calculation mechanism in the model to be trained to obtain the pre-trained model; the fine-tuning data set is determined by collecting the point cloud data of the target product, and the pre-trained model and the fine-tuning data set are used to perform feature extraction and point cloud classification, so as to complete the model fine-tuning operation based on the obtained classification result and obtain the target model. That is, in this application, firstly, a pre-training data set is obtained, and then the model to be trained is determined by using the preset local perception self-attention mechanism and the converter model structure, and then the pre-training strategy of mask modeling of the model to be trained is used to process the pre-training data set, extract local and global features, and determine the training loss to complete the pre-training, and then the point cloud data of the target product is collected to extract features and classify the point cloud of the obtained pre-trained model to complete fine-tuning and determine the target model. In this way, the model's local perception of point cloud data features and its ability to process large-scale point cloud data can be effectively improved, thereby improving the situation where the pre-trained point cloud model based on the converter model structure has weak perception of local features, improving the quality of point cloud feature extraction, and thereby improving the precision and accuracy of processing point cloud downstream classification tasks.

[0072] See also Figure 4 As shown, the embodiment of the present application also discloses a local perception point cloud representation learning pre-training device based on an attention mechanism, including:

[0073] The data set determination module 11 is used to determine a corresponding pre-training data set by sampling point cloud data on the surface of the three-dimensional model; the pre-training data set is a data set in the shape of a point cloud;

[0074] The data processing module 12 is used to determine the model to be trained based on the converter model structure and the preset local perception self-attention mechanism, and use the model to be trained and the pre-training strategy of mask modeling to sequentially group and embed the pre-training data set, and randomly mask the embedded point cloud data to obtain the point cloud data processing result;

[0075] A pre-training module 13 is used to perform local and global feature extraction based on the model to be trained, the pre-training strategy and the point cloud data processing result, and determine the training loss using the obtained feature extraction result and the density-aware chamfer distance loss calculation mechanism in the model to be trained to obtain a pre-trained model;

[0076] The model fine-tuning module 14 is used to determine the fine-tuning data set by collecting point cloud data of the target product, and use the pre-trained model and the fine-tuning data set to perform feature extraction and point cloud classification, so as to complete the model fine-tuning operation based on the obtained classification results and obtain the target model.

[0077] Among them, for more specific working processes of the above-mentioned modules, please refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.

[0078] It can be seen that in this application, firstly, a pre-training data set is obtained, and then the model to be trained is determined by using the preset local perception self-attention mechanism and the converter model structure, and then the pre-training strategy of mask modeling of the model to be trained is used to process the pre-training data set, extract local and global features, and determine the training loss to complete the pre-training, and then the point cloud data of the target product is collected to extract features and classify the point cloud of the obtained pre-trained model to complete fine-tuning and determine the target model. In this way, the model's local perception of point cloud data features and its ability to process large-scale point cloud data can be effectively improved, thereby improving the situation where the pre-trained point cloud model based on the converter model structure has weak perception of local features, improving the quality of point cloud feature extraction, and thereby improving the precision and accuracy of processing point cloud downstream classification tasks.

[0079] In some specific embodiments, the data set determination module 11 may specifically include:

[0080] A sampling quantity acquisition unit, used to acquire a preset sampling quantity;

[0081] The pre-training data set determination unit is used to perform point cloud data sampling on the surface of the three-dimensional model by the preset sampling quantity to determine the corresponding pre-training data set; there is no annotation information in the pre-training data set.

[0082] In some specific embodiments, the data processing module 12 may specifically include:

[0083] A data grouping unit, used to determine a corresponding number of sampling points from the pre-training data set based on the model to be trained, the farthest point sampling strategy and the preset number of groups, and to determine the neighboring points of each sampling point using a K-nearest neighbor algorithm to complete the corresponding point cloud data grouping operation and obtain a grouping result;

[0084] A data embedding unit, for mapping the corresponding point cloud coordinates to the feature space through a combination of a multi-layer linear mapping layer, a splicing layer, and a maximum pooling layer in the model to be trained for any point cloud group in the grouping result, so as to complete the point cloud embedding operation and obtain the embedded feature vector of each point cloud group;

[0085] A random masking unit, configured to perform random masking on each of the point cloud groups in the grouping result based on a preset ratio to obtain a corresponding random masking result;

[0086] A processing result determination unit is used to determine a point cloud data processing result based on the random mask result and the embedded feature vector of each point cloud group.

[0087] In some specific embodiments, the pre-training module 13 may specifically include:

[0088] A grouping determination unit, configured to determine an unmasked point cloud group and a masked point cloud group based on the model to be trained and the random masking result;

[0089] A local feature extraction unit is used to determine, for any data point in the unmasked point cloud group, a neighborhood point corresponding to the current data point based on a K-nearest neighbor algorithm, and determine a local feature corresponding to the current data point for the current data point and the neighborhood point using a self-attention mechanism in a corresponding local neighborhood area;

[0090] A global feature extraction unit, configured to downsample the unmasked point cloud group based on a downsampling operator, and determine a global feature corresponding to the current data point using a self-attention mechanism for the current data point and the obtained downsampling result;

[0091] The feature fusion unit is used to splice the local features corresponding to the current data point and the global features, and use a normalized exponential function to weight the obtained spliced ​​features to complete the feature fusion operation and obtain the fused features corresponding to the current data point.

[0092] In some specific embodiments, the pre-training module 13 may specifically include:

[0093] A parameter acquisition unit, used to acquire learnable parameters corresponding to feature information of the masked point cloud group;

[0094] a result splicing unit, configured to determine point cloud group sequence information according to the grouping result, and splice the learnable parameters and the feature extraction results corresponding to the unmasked point cloud group using the point cloud group sequence information to obtain a corresponding splicing result;

[0095] A coordinate prediction unit, used for predicting the coordinates of each of the point cloud groups based on the decoder in the model to be trained and the splicing result, so as to obtain a corresponding coordinate prediction result;

[0096] A point cloud group reconstruction unit, configured to reconstruct the masked point cloud group based on the coordinate prediction result to obtain a corresponding reconstruction result;

[0097] A loss determination unit is used to trigger the corresponding point cloud group density information acquisition operation by using the density-aware chamfer distance loss calculation mechanism in the model to be trained, and determine the training loss based on the obtained point cloud group density information acquisition result, the reconstructed point coordinates in the reconstruction result, and the real point coordinates corresponding to the masked point cloud group to obtain the pre-trained model.

[0098] In some specific embodiments, the model fine-tuning module 14 may specifically include:

[0099] A data acquisition unit is used to collect point cloud data corresponding to the target product to determine the corresponding acquisition results;

[0100] A coefficient acquisition unit, used to acquire a scaling coefficient and a translation coefficient corresponding to each coordinate axis in the three-dimensional coordinate axis;

[0101] A data processing unit, used for processing the point cloud data in the acquisition result based on the scaling factor and the translation factor to obtain corresponding processed data;

[0102] The fine-tuning data set determining unit is used to construct a fine-tuning data set based on the processed data.

[0103] In some specific embodiments, the model fine-tuning module 14 may specifically include:

[0104] A model processing unit, used to remove the decoder in the pre-trained model and add a classification head constructed based on a multi-layer linear mapping layer to the pre-trained model to obtain a corresponding processed model;

[0105] A feature extraction unit, used to perform local and global feature extraction using the processed model and the fine-tuning data set to obtain a corresponding target extraction result;

[0106] The point cloud classification unit is used to perform point cloud category mapping on the target extraction result based on the processed model to complete the corresponding point cloud classification operation and obtain the classification result.

[0107] Furthermore, the present application also discloses an electronic device. Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.

[0108] Figure 5 A schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the local perception point cloud representation learning pre-training method based on the attention mechanism disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0109] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0110] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0111] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the local perception point cloud representation learning pre-training method based on the attention mechanism performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.

[0112] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned disclosed local perception point cloud representation learning pre-training method based on attention mechanism is implemented. The specific steps of the method can refer to the corresponding contents disclosed in the aforementioned embodiments, and will not be repeated here.

[0113] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0114] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0115] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0116] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0117] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of ​​the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A local perception point cloud representation learning pre-training method based on attention mechanism, characterized in that: include: By sampling point cloud data on the surface of the three-dimensional model, a corresponding pre-training data set is determined; The pre-training data set is a data set in the shape of a point cloud; Based on the converter model structure and the preset local perception self-attention mechanism, the model to be trained is determined, and the pre-training strategy of the model to be trained and mask modeling is used to group and embed the pre-training data set in sequence, and the embedded point cloud data is randomly masked to obtain the point cloud data processing result; Performing local and global feature extraction based on the model to be trained, the pre-training strategy and the point cloud data processing result, and determining the training loss using the obtained feature extraction result and the density-aware chamfer distance loss calculation mechanism in the model to be trained to obtain a pre-trained model; Determine a fine-tuning data set by collecting point cloud data of a target product, and use the pre-trained model and the fine-tuning data set to perform feature extraction and point cloud classification, so as to complete a model fine-tuning operation based on the obtained classification result and obtain a target model; The pre-training strategy using the model to be trained and mask modeling sequentially groups and embeds the pre-training data set, and randomly masks the embedded point cloud data, including: Determine a corresponding number of sampling points from the pre-training data set based on the model to be trained, the farthest point sampling strategy and the preset number of groups, and use the K nearest neighbor algorithm to determine the neighboring points of each sampling point to complete the corresponding point cloud data grouping operation and obtain the grouping result; For any point cloud group in the grouping result, the corresponding point cloud coordinates are mapped to the feature space through a combination of a multi-layer linear mapping layer, a splicing layer, and a maximum pooling layer in the model to be trained to complete the point cloud embedding operation and obtain the embedded feature vector of each point cloud group; Performing random masking on each of the point cloud groups in the grouping result based on a preset ratio to obtain a corresponding random masking result; Determining a point cloud data processing result based on the random mask result and the embedded feature vector of each point cloud group; The extracting local and global features based on the model to be trained, the pre-training strategy and the point cloud data processing result includes: Determine an unmasked point cloud group and a masked point cloud group based on the model to be trained and the random mask result; For any data point in the unmasked point cloud group, a neighborhood point corresponding to the current data point is determined based on a K-nearest neighbor algorithm, and a local feature corresponding to the current data point is determined for the current data point and the neighborhood point using a self-attention mechanism in a corresponding local neighborhood area; Downsampling the unmasked point cloud group based on a downsampling operator, and determining a global feature corresponding to the current data point and the obtained downsampling result using a self-attention mechanism; The local features and the global features corresponding to the current data point are spliced, and the obtained spliced ​​features are weighted using a normalized exponential function to complete a feature fusion operation and obtain a fused feature corresponding to the current data point.

2. The local perception point cloud representation learning pre-training method based on the attention mechanism according to claim 1 is characterized in that: The method of sampling point cloud data on the surface of the three-dimensional model to determine the corresponding pre-training data set includes: Get the preset number of samples; Point cloud data sampling is performed on the surface of the three-dimensional model using the preset sampling quantity to determine a corresponding pre-training data set; and no annotation information exists in the pre-training data set.

3. The local perception point cloud representation learning pre-training method based on the attention mechanism according to claim 1 is characterized in that: The method of determining the training loss by using the obtained feature extraction result and the density-aware chamfer distance loss calculation mechanism in the model to be trained includes: Obtaining learnable parameters corresponding to feature information of the masked point cloud group; Determining point cloud group sequence information according to the grouping result, and using the point cloud group sequence information to splice the learnable parameters and the feature extraction results corresponding to the unmasked point cloud group to obtain a corresponding splicing result; Predicting the coordinates of each of the point cloud groups based on the decoder in the model to be trained and the splicing result to obtain corresponding coordinate prediction results; Reconstructing the masked point cloud group based on the coordinate prediction result to obtain a corresponding reconstruction result; The density-aware chamfer distance loss calculation mechanism in the model to be trained is used to trigger the corresponding point cloud group density information acquisition operation, and the training loss is determined based on the obtained point cloud group density information acquisition result, the reconstructed point coordinates in the reconstruction result, and the real point coordinates corresponding to the masked point cloud group to obtain the pre-trained model.

4. The local perception point cloud representation learning pre-training method based on the attention mechanism according to claim 1 is characterized in that: The step of determining the fine-tuning data set by collecting point cloud data of the target product includes: Collect the point cloud data corresponding to the target product to determine the corresponding collection results; Obtain the scaling factor and translation factor corresponding to each coordinate axis in the three-dimensional coordinate axis; Processing the point cloud data in the acquisition result based on the scaling factor and the translation factor to obtain corresponding processed data; A fine-tuning dataset is constructed based on the processed data.

5. The local perception point cloud representation learning pre-training method based on the attention mechanism according to any one of claims 1 to 4, characterized in that: The method of using the pre-trained model and the fine-tuning data set to perform feature extraction and point cloud classification includes: Removing the decoder in the pre-trained model, and adding a classification head constructed based on a multi-layer linear mapping layer to the pre-trained model to obtain a corresponding processed model; Using the processed model and the fine-tuning data set to perform local and global feature extraction to obtain corresponding target extraction results; Based on the processed model, point cloud category mapping is performed on the target extraction result to complete the corresponding point cloud classification operation and obtain the classification result.

6. A local perception point cloud representation learning pre-training device based on attention mechanism, characterized in that: include: A data set determination module, used to determine a corresponding pre-training data set by sampling point cloud data on the surface of the three-dimensional model; The pre-training data set is a data set in the shape of a point cloud; A data processing module is used to determine the model to be trained based on the converter model structure and the preset local perception self-attention mechanism, and to group and embed the pre-trained data set in sequence using the pre-training strategy of the model to be trained and mask modeling, and to randomly mask the embedded point cloud data to obtain a point cloud data processing result; A pre-training module, used for performing local and global feature extraction based on the model to be trained, the pre-training strategy and the point cloud data processing result, and determining the training loss using the obtained feature extraction result and the density-aware chamfer distance loss calculation mechanism in the model to be trained, so as to obtain a pre-trained model; A model fine-tuning module, used to determine a fine-tuning data set by collecting point cloud data of a target product, and perform feature extraction and point cloud classification using the pre-trained model and the fine-tuning data set, so as to complete a model fine-tuning operation based on the obtained classification result and obtain a target model; Wherein, the data processing module includes: A data grouping unit, used to determine a corresponding number of sampling points from the pre-training data set based on the model to be trained, the farthest point sampling strategy and the preset number of groups, and to determine the neighboring points of each sampling point using a K-nearest neighbor algorithm to complete the corresponding point cloud data grouping operation and obtain a grouping result; A data embedding unit, for mapping the corresponding point cloud coordinates to the feature space through a combination of a multi-layer linear mapping layer, a splicing layer, and a maximum pooling layer in the model to be trained for any point cloud group in the grouping result, so as to complete the point cloud embedding operation and obtain the embedded feature vector of each point cloud group; A random masking unit, configured to perform random masking on each of the point cloud groups in the grouping result based on a preset ratio to obtain a corresponding random masking result; a processing result determining unit, configured to determine a point cloud data processing result based on the random mask result and the embedded feature vector of each point cloud group; The pre-training module comprises: A grouping determination unit, configured to determine an unmasked point cloud group and a masked point cloud group based on the model to be trained and the random masking result; A local feature extraction unit is used to determine, for any data point in the unmasked point cloud group, a neighborhood point corresponding to the current data point based on a K-nearest neighbor algorithm, and determine a local feature corresponding to the current data point for the current data point and the neighborhood point using a self-attention mechanism in a corresponding local neighborhood area; A global feature extraction unit, configured to downsample the unmasked point cloud group based on a downsampling operator, and determine a global feature corresponding to the current data point using a self-attention mechanism for the current data point and the obtained downsampling result; The feature fusion unit is used to splice the local features corresponding to the current data point and the global features, and use a normalized exponential function to weight the obtained spliced ​​features to complete the feature fusion operation and obtain the fused features corresponding to the current data point.

Citation Information

Patent Citations

  • Morphological feature extraction method and morphological structure generation method of neurons

    CN118918258A

Cited By

  • Universal efficient low-parameter fine tuning method for point cloud analysis

    CN121937834A