Multi-scale point cloud pre-training method for sampling from top to bottom, equipment and medium

Through the multi-scale point cloud pre-training method of top-down sampling, combined with mask automatic encoder and U-Net architecture, the problem of insufficient local features and global information extraction in point cloud data processing is solved, and higher accuracy and robustness is achieved, and it is suitable for tasks such as three-dimensional shape classification, object detection and partial segmentation.

CN120431441APending Publication Date: 2025-08-05SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510376133.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing point cloud data processing methods have shortcomings in local features and global information extraction, multi-scale information is easily lost, and the problem of location information leakage affects the performance of the model in downstream tasks.

Method used

The multi-scale point cloud pre-training method with top-down sampling is adopted, combining mask autoencoder and U-Net architecture, through self-supervised learning, upsampling and integration of high-level abstract features and underlying detailed features, and a shared and trainable mask token mechanism is used to avoid location information leakage.

Benefits of technology

It significantly improves the multi-scale feature extraction and reconstruction capabilities of point cloud data, and improves the accuracy and generalization capabilities of tasks such as three-dimensional shape classification, object detection and partial segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431441A_ABST
    Figure CN120431441A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale point cloud pre-training method for sampling from top to bottom, equipment and a medium. According to the method, through a U-Net type asymmetric encoder-decoder architecture and in combination with a multi-scale mask, the extraction capability of local features and global information of point cloud data is enhanced. According to the method, firstly, input point cloud data is subjected to multi-scale division and mask processing, then multi-scale features are extracted through an encoder, a masked point cloud area is reconstructed through a decoder, and a masked local point cloud block is recovered through a linear projection layer. According to the invention, the problem of multi-scale information leakage in the prior art is effectively avoided, and the point cloud data processing precision is improved. After self-supervised pre-training, the model shows excellent performance in a plurality of downstream tasks, especially in applications such as 3D shape classification, object detection and partial segmentation. The method has the advantages of high robustness and wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and point cloud processing, and specifically relates to a top-down sampling multi-scale point cloud pre-training method, equipment and medium, which can be used for tasks such as 3D shape classification, object detection and part segmentation. Background Art

[0002] As an unstructured data form that expresses the shape of three-dimensional objects, 3D point clouds have a wide range of applications in autonomous driving, robotic navigation, 3D reconstruction, and other fields. However, point cloud data presents challenges such as sparsity, irregularity, and complex geometric features, making it difficult for traditional convolutional neural networks (CNNs) to effectively process both local and global information from point cloud data.

[0003] Existing masked autoencoders (MAEs) have made significant progress in natural language processing and two-dimensional vision tasks and are gradually being applied to point cloud data processing. However, existing MAE models, when processing point clouds, are primarily limited to extracting global information while neglecting the capture of local features. This deficiency leads to the loss of multi-scale information in point clouds, making it impossible to balance local details with global semantic expression. Furthermore, existing point cloud pre-training models exhibit positional information leakage issues in their information masking strategies, which in turn affects the fusion and transfer of multi-scale features, limiting performance in downstream tasks.

[0004] Point cloud pre-training methods based on masked autoencoders (e.g., “Point-M2AE: multi-scale masked autoencoders for hierarchical point cloud pre-training” https: / / proceedings.neurips.cc / paper_files / paper / 2022 / hash / ad1d7a4df30a9c0c46b387815a774a84-Abstract-Conferenc e.html”) suffer from shortcomings such as inconsistent multi-scale information, leakage of low-scale position information, imbalance between local details and global semantic extraction, and lack of effective U-Net-style multi-scale feature fusion. Existing technologies usually only randomly mask local point cloud blocks at the lowest scale, and fail to ensure the consistency of the masked area when downsampling at higher scales, resulting in partial leakage of low-scale position information, making it easy for the model to "peek" at the unmasked parts during self-supervised training, making it difficult to fully capture fine geometric structures; in addition, some methods over-rely on multi-layer Transformer stacking, ignoring the advantages of using jump connections to fuse low-level details with high-level abstract semantics, and fail to adopt a shared and trainable mask token mechanism for the masked areas, further limiting the reconstruction accuracy and the migration ability of downstream tasks. Summary of the Invention

[0005] In order to solve at least one of the problems existing in the prior art, the present invention provides a self-supervised learning method for 3D point cloud data. By combining a masked autoencoder (MAE) with a Unet-like architecture, the method solves the problem of insufficient extraction of local details and global semantic information of point cloud data in the prior art, improves the multi-scale feature extraction and reconstruction capabilities, and is applied to tasks such as 3D shape classification, few-shot learning, partial segmentation, and 3D object detection. In order to achieve the above-mentioned purpose of the invention, the present invention proposes a multi-scale point cloud pre-training method for top-down sampling, comprising the following steps:

[0006] Select a representative center point from the standardized point cloud data;

[0007] Select neighboring points for each center point to construct a local point cloud block;

[0008] According to the predetermined mask ratio, the local point cloud blocks are randomly masked at the lowest scale and gradually continue to mask at higher scales;

[0009] The masked local point cloud block serves as the input to the encoder-decoder model. The encoder is used to extract point cloud features at different scales, interact with contextual information, and use a skip connection mechanism to upsample and fuse high-level abstract features with low-level detailed features. It outputs a visible token and a shared and trainable mask token. The decoder is used to reconstruct the features of the visible token and the shared and trainable mask token. Then, a linear projection layer is used to map the visible token and the shared and trainable mask token back to the three-dimensional coordinate space to recover the masked local point cloud block.

[0010] The encoder-decoder model is trained, and after training is completed, the encoder is used as a pre-trained model for downstream tasks.

[0011] Furthermore, the farthest point sampling algorithm is used to select representative center points.

[0012] Furthermore, a top-down masking strategy is adopted, which masks the point cloud at the lowest scale and then ensures the consistency of the mask at higher scales step by step. The first step of the masking operation is to mask the local point cloud block at the lowest scale s1 with a preset ratio α;

[0013] The mask ratio is:

[0014] m i =m i-1 ×(1-α)

[0015] Through this top-down multi-scale masking strategy, the mask areas of local point cloud blocks at different scales remain consistent, ensuring that position information will not leak at lower scales, significantly improving the retention of local features and the ability to process multi-scale information.

[0016] Furthermore, the k-nearest neighbor algorithm is used to select neighboring points for each center point to construct a local point cloud block, which is defined as follows:

[0017] P=KNN(FPS(X),X),

[0018] Among them, X is the input point cloud data, m1 is the number of center points sampled at the lowest scale s1, and k1 is the number of neighborhood points of each center point.

[0019] The contraction path in the encoder is used to extract multi-scale local and global features, and the expansion path is used to fuse high-level abstract features with low-level detail features layer by layer. The multiple scale encoders of the encoder are used for context information interaction, and the formula is as follows:

[0020]

[0021] Among them, T i is the input feature of the i-th scale, is the output feature after the transformer block (scale encoder), d i is the feature dimension.

[0022] The fusion process of the jump connection is described as:

[0023]

[0024] Among them, θ proj It is a linear projection layer used to transform high-level features With low-level decoding features Fusion. Through this process, high-level semantic features are transferred layer by layer, allowing features to be propagated from the highest scale to the lowest scale, ensuring the multi-scale representation capability of the point cloud. This design combines the advantages of a Unet-like architecture, extracting local features at a lower level while obtaining global features at a higher level through a skip connection mechanism. This effectively solves the problem of fusing local details with global semantics in point cloud data.

[0025] The input point cloud first generates a center point through the farthest point sampling (FPS), and then the k-nearest neighbor algorithm (KNN) is used to select neighboring points for each center point to construct a local point cloud block. The Mini-PointNet structure is used to extract local geometric features. The formula is as follows:

[0026] P=KNN(FPS(X),X),

[0027] For each local point cloud block, by calculating its relative position with the center point, its relative coordinate representation is obtained, which helps to capture the local geometric structure of the point cloud. This relative position representation is calculated as:

[0028] TC1=Mini-PointNet(P vis ),

[0029] In the pre-training stage of the model, Chamfer distance is used as the loss function to calculate the difference between the reconstructed local point cloud block and the original local point cloud block, which is defined as:

[0030]

[0031] Among them, P is the reconstructed local point cloud block, For the original local point cloud block, by minimizing the Chamfer distance, the reconstruction accuracy of the point cloud can be effectively improved. This loss function effectively reduces the geometric errors generated during the reconstruction process, allowing the masked area to be accurately reconstructed.

[0032] For each masked local point cloud block, the output of the decoder is projected through a linear projection layer and the coordinates of the neighboring points are restored as follows:

[0033] P rec =θ rec (T D ),

[0034] Among them, P rec is the reconstructed point cloud block, T D is the output of the decoder, θ rec It is a linear projection function used to map the token back to the point cloud coordinates.

[0035] The L2 Chamfer distance is used as the reconstruction loss function to reconstruct the masked region, guiding the model to capture local and global feature representations of the point cloud in a self-supervised learning process.

[0036] Furthermore, a prediction head of the corresponding downstream task is added to the obtained pre-trained model to output the final result for the specific task; when used for three-dimensional shape classification tasks, a classification layer is attached to the pre-trained model, and the classification result is obtained after the data to be classified is input; when used for segmentation tasks, a segmentation head is attached to the pre-trained model, and the segmentation result is generated after the result to be segmented is input; when used for three-dimensional object detection, a detection head is attached to the pre-trained model, and the object position and category in the point cloud are generated after the point cloud data to be detected is input.

[0037] The present invention also provides a computer device.

[0038] The present invention also provides a computer-readable storage medium.

[0039] The present invention implements random masking at the lowest scale and maintains mask consistency at each level. Combining a U-Net-style asymmetric encoder-decoder structure with a shared, trainable mask token mechanism, it can not only effectively fuse multi-scale local and global features, but also significantly improve the reconstruction accuracy of point clouds and the performance of downstream tasks (such as 3D shape classification, object detection, and part segmentation), thus making up for the shortcomings of the existing technology.

[0040] Compared with the prior art, the present invention has at least the following beneficial effects:

[0041] This invention effectively avoids the multi-scale information leakage problem encountered in existing technologies and improves the accuracy of point cloud data processing. Through self-supervised pre-training, the model more fully characterizes the local and global features of 3D point clouds. This enables higher accuracy, stronger generalization, and more refined segmentation or detection in downstream tasks such as 3D shape classification, few-shot learning, partial segmentation, and 3D object detection. This significantly improves the final results of these tasks and possesses great practical application value.

[0042] The present invention has strong robustness, is suitable for a variety of 3D point cloud data processing tasks, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a flow chart of the present invention.

[0044] Figure 2 The detailed process of the proposed method is described, in which the encoder is responsible for multi-scale feature extraction and interaction with context information. Subsequently, a simple decoder and reconstruction head are used to reconstruct the masked point cloud.

[0045] Figure 3 Schematic diagram of the visualization results after training the ModelNet40 dataset. DETAILED DESCRIPTION

[0046] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. In addition, it should be understood that the specific embodiments described herein are merely for explanation of the present invention and are not intended to limit the present invention. Therefore, the scope of protection of the present invention is not limited by the specific embodiments disclosed below.

[0047] like Figure 1 As shown, an embodiment of the present invention provides a multi-scale point cloud pre-training method for top-down sampling, and the specific implementation steps are as follows:

[0048] Step 1: Standardize the input raw point cloud data.

[0049] In some embodiments of the present invention, this step includes the following sub-steps:

[0050] Step 1: Dataset collection and preparation

[0051] Step 1-1: Determination of original point cloud data

[0052] The raw point cloud data uses the publicly available ModelNet40 dataset. This dataset contains 40 common object categories, such as chairs, tables, and sofas. Each category contains multiple sets of 3D object models, providing rich 3D spatial geometry information. It is understood that the raw point cloud data is not limited to the ModelNet40 dataset; other datasets may also be used in other embodiments.

[0053] Steps 1-2: Data format and conversion

[0054] The 3D object models in the ModelNet40 dataset are usually in mesh form (such as .obj files). To use the point cloud format, each mesh-based 3D object model is sampled and converted into a point cloud representation.

[0055] For each 3D object model, a point cloud set containing N points is generated using uniform sampling or farthest point sampling (FPS) algorithm, where each point Represents the three-dimensional coordinates (x, y, z) of the corresponding three-dimensional object model.

[0056] Steps 1-3: Standardization

[0057] To unify the data scale, the point cloud of each 3D object model is normalized. The 3D coordinates of each point cloud are first processed with zero mean, and then scaled according to the maximum boundary range to normalize them to the unit cube range of [-1, 1].

[0058] The normalization formula is as follows:

[0059]

[0060] Here, x represents the input 3D point cloud data matrix, x norm Represents the normalized coordinates.

[0061] The standardized point cloud data has a uniform scale and is suitable for subsequent model input and training.

[0062] Step 2: Select a representative center point from the standardized point cloud data.

[0063] like Figure 2 As shown in the upper part of , in order to achieve multi-scale feature extraction and prevent over-reliance on local area information, multi-scale sampling is performed on each standardized point cloud.

[0064] In some embodiments of the present invention, a representative center point set P is selected from the normalized point cloud using the farthest point sampling (FPS) method. centroid This method can ensure that the sampling points are evenly distributed and improve the ability of the point cloud block to represent the global structure.

[0065] P centroid =FPS(X),

[0066] Among them, FPS(X) represents the operation of sampling the farthest point of point cloud X, m is the number of center points after sampling, Indicates that the coordinates of the selected center point are three-dimensional.

[0067] Step 3: Select neighboring points for each representative center point and construct a local point cloud block.

[0068] Use the k-nearest neighbor algorithm (KNN) to select k neighborhood points from the standardized point cloud for each center point to form a local point cloud block P, that is:

[0069] P = KNN(P centroid ,X),

[0070] but:

[0071] P=KNN(FPS(X),X),

[0072] Among them, P represents the selected local point cloud block, X is the input point cloud data, m1 is the number of center points after sampling, k1 is the number of neighborhood points of each center point, and k is the number of neighborhood points.

[0073] The farthest point sampling (FPS) algorithm is used to generate local point cloud blocks at each scale to ensure the sparsity of the point cloud.

[0074] For each local point cloud block, by calculating the relative position of the local point cloud block and the center point, the relative position representation of the local point cloud block is obtained, which helps to capture the local geometric structure of the point cloud. This relative position representation is calculated as:

[0075] TC1=Mini-PointNet(P),

[0076] Among them, TC1 represents the local feature representation obtained after Mini-PointNet processing, Mini-PointNet(P) represents the use of a lightweight PointNet network for feature extraction of the local point cloud block P, and d1 represents the dimension of the extracted feature vector.

[0077] By constructing local point cloud blocks, the geometric details of the point cloud can be effectively captured, laying the foundation for subsequent feature extraction.

[0078] Step 4: According to the predetermined mask ratio, the local point cloud blocks are randomly masked at the lowest scale, and the masking is gradually continued at higher scales to ensure the consistency of the mask across scales and avoid low-scale information leakage.

[0079] When masking, the lowest-scale local point cloud block is masked according to the preset mask ratio α, and the masked local point cloud block is used to calculate the reconstruction loss; at higher scales s i On the top, use the farthest point sampling algorithm FPS to the previous scale s i-1 The visible local point cloud blocks are downsampled; the KNN algorithm is used to select k for each downsampled center point i neighborhood points to form a higher-scale feature merging.

[0080] Through the top-down masking mechanism, the consistency of masked areas at different scales can be ensured.

[0081] On the local point cloud block, a preset proportion of local point cloud blocks are randomly masked at the lowest scale s1, and the mask ratio is set to α. This masking operation can prevent the model from directly relying on complete position information, thereby improving the self-supervised learning effect of the model.

[0082] The number of center points after masking is calculated by decreasing the following formula to ensure the consistency of the mask at higher scales:

[0083] m i =m i-1 ×(1-α)

[0084] Among them, m i Indicates the number of center points at the current scale.

[0085] This step ensures that the positions of the mask points remain consistent at different scales, preventing the leakage of position information at low scales.

[0086] Step 5: Construct an encoder-decoder model. The encoder extracts point cloud features of different scales in the contraction path, interacts with contextual information through the Transformer module, and uses the skip connection mechanism to upsample and fuse high-level abstract features with low-level detail features.

[0087] Step 5-1. Encoder design

[0088] In this paper, a U-Net-style asymmetric encoder-decoder architecture is used, and the encoder and decoder are asymmetric: the encoder adopts a U-Net-style architecture, including deep contraction paths and expansion paths, and realizes multi-scale feature extraction and fusion through multi-layer Transformer modules to fully capture the local details and global semantics of the point cloud; the decoder is a lightweight structure, including a small number of modules, which focuses on mapping the features output by the encoder back to the masked local point cloud blocks to reconstruct the masked areas.

[0089] The contraction path includes multiple scale encoders (visual transformer blocks). Each scale encoder is used to first extract local features of the masked local point cloud block. These local features are downsampled and integrated layer by layer to form representations of different scales. Then, the Transformer block is used to perform self-attention calculations to interact and fuse the local features and global context information at each scale, thereby completing the extraction of multi-scale features.

[0090]

[0091] Among them, T i is the input feature of the i-th scale, is the output feature after the scale encoder (visual transformer block), d i is the dimension of the feature representation of the i-th layer.

[0092] In some embodiments of the present invention, the multi-scale encoder captures local features of local point cloud blocks through the local feature extraction module of Mini-PointNet (including MLP and pooling operations).

[0093] like Figure 2As shown in the figure, the encoder-decoder model adopts a U-Net-style encoder-decoder structure. It takes local point cloud patches as input and uses the Mini-PointNet module to extract local geometric features. These features are converted into visible tokens and fed into the Transformer block for global context interaction. As the encoder-decoder model advances from scale i to scale i+1, it further downsamples and merges visible tokens from the previous scale to obtain higher-level, more abstract token representations. This process is repeated multiple times, gradually forming a feature pyramid from fine-grained to coarse-grained. During this process, the contraction path is primarily responsible for continuously downsampling and aggregating features. In the expansion path, the encoder-decoder model preserves and fuses feature representations at different scales through skip connections. Interpolation or upsampling mechanisms are used to gradually project high-level semantic features back to higher resolutions, fully integrating local details with global information. Ultimately, the multi-scale fused features output by the expansion path provide the decoder with rich and complete contextual support, facilitating the subsequent reconstruction of masked regions and better recovering the geometric structure of the point cloud.

[0094] The contraction path is responsible for extracting multi-scale features of the point cloud, and each layer achieves the fusion of local and global information through the Transformer block (Visual Transformer (ViT) block). For each scale s i The features of the local point cloud blocks in are aggregated using a multi-layer perceptron (MLP) and a maximum pooling layer to aggregate neighborhood information and pass it to the next scale. The formula for feature aggregation is as follows:

[0095]

[0096] in, represents the feature representation of the i-th scale, θ max is the maximum pooling operator, which is used to select the maximum value in each feature dimension to aggregate neighborhood information, d i is the dimension of the feature representation of the i-th layer, is the feature of the i-1th scale after being processed by the Transformer module, Represents a size m i ×d i The characteristic matrix of .

[0097] The expansion path restores low-scale features layer by layer to capture high-scale global information, ensuring that features at each scale contain rich contextual information. The expansion path uses a skip connection mechanism to propagate high-level features back to lower layers to enhance the encoder's multi-scale feature fusion capabilities. This skip connection mechanism is described by the following formula:

[0098]

[0099] Among them, θ proj It is a linear projection layer used to transform high-level features With low-level decoding features Fusion. The feature fusion formula of the extended path is as follows:

[0100]

[0101] in represents the output feature of the i-th scale on the contraction path, It indicates that the extended path samples or interpolates features from a higher scale (such as scale i+1) when restoring the features of scale i. Indicates that and The final features are obtained after fusion through the linear projection layer.

[0102] Skip connections are used in the encoder to preserve and pass multi-scale features to the decoder. The skip connection mechanism can preserve low-level geometric details and avoid losing local details in multi-layer feature extraction.

[0103] Step 5-2, decoder design

[0104] The lightweight decoder is used to process the masked area, using a shared and trainable mask token. The decoder processes the masked point cloud blocks through a shared and trainable mask token mechanism, attaches the mask token at the initial scale and passes it to the Transformer for reconstruction. First, the decoder input is the visible token output by the encoder and the shared and trainable mask token. Then, the decoder uses a lightweight Transformer module to process the input features, and maps high-level features to a lower scale through upsampling technology and weighted interpolation to restore more refined local geometric information. Figure 2 As can be seen, the shared and trainable mask token is concatenated with the visible token and input into the decoder. The decoder passes the concatenated token into the lightweight Transformer module. The output features are then mapped back to three-dimensional coordinates through the linear projection layer to complete the reconstruction of the masked area.

[0105] The decoder recovers the masked local point cloud blocks layer by layer through multiple layers of lightweight Transformer modules. The decoder receives the feature information transmitted by the encoder through skip connections and recovers the geometric features of the point cloud blocks by gradually fusing low-level and high-level information.

[0106] At the initial scale s i-1 On the other hand, the shared and trainable mask token T maskAppended after the visible token and passed to the decoder for processing, it is defined as:

[0107]

[0108] T D Indicates that it will be visible With mask tokenT mask The output features are input into the decoder and processed by several layers of lightweight Transformer or other modules. It means that the size of the matrix (or tensor) is m1×d1, that is, there are m1 tokens and the feature dimension of each token is d1.

[0109] Step 5-3: Linear Projection and Reconstruction

[0110] The output of the decoder passes through the linear projection layer, which maps the feature vector output by the decoder back to three-dimensional coordinates to complete the reconstruction of the mask area. The linear projection layer reconstructs the nearest neighbor points of the initial scale using the k-nearest neighbor algorithm. The projection formula is:

[0111] P rec =θ rec (T D ),

[0112] Among them, P rec is the reconstructed local point cloud block, θ rec is the linear projection function, T D is the feature representation output by the decoder, m1 is the number of center points after sampling, k1 is the number of neighborhood points for each center point, Represents the point cloud data reconstructed by the decoder or the coordinate set of the masked area that needs to be processed.

[0113] Step 6: Train and optimize the encoder-decoder model, and use the trained encoder as a pre-trained model for downstream tasks.

[0114] Step 6-1. Loss function design

[0115] like Figure 2 As shown, after the decoder output is completed, the reconstruction head generates a reconstructed point cloud block, and then the L2Chamfer distance is used to calculate the reconstructed point cloud block P rec With the original local point cloud block The error between them. L2Chamfer distance is used as the reconstruction loss function to accurately measure the geometric difference between two sets of point cloud blocks, ensuring that the encoder-decoder model can accurately reconstruct the masked area. The loss function formula is:

[0116]

[0117] in, Represents the reconstructed point cloud block P rec With the original local point cloud block The error between .

[0118] Step 6-2, Optimizer and Training Settings

[0119] The Adam optimizer is selected, the learning rate is set to 0.001, and the batch size is 64. In each iteration, the Chamfer distance loss is calculated, and the parameters of the encoder and decoder are updated based on the loss to improve the reconstruction ability and feature learning ability of the encoder-decoder model.

[0120] Step 6-3: Model validation and performance evaluation

[0121] Validation set and test set division:

[0122] The ModelNet40 dataset is divided into training, validation, and test sets. The validation set is used to optimize hyperparameters and check for overfitting of the encoder-decoder model, while the test set is used to evaluate the final performance of the encoder-decoder model.

[0123] Step 6-4. Performance evaluation indicators

[0124] The reconstructed L2 Chamfer distance is used as the main evaluation metric for the model in the self-supervised task. At the same time, the trained model features are used for the 3D shape classification task, and the classification accuracy is calculated using a linear classifier.

[0125] In one embodiment of the present invention, Figure 3 In the figure, t-SNE is used to visualize the point feature distribution obtained using different training strategies on the ModelNet40 dataset. Specifically, this example shows four different settings: (a) using only the feature distribution at the simplest initialization, (b) training directly from scratch on the ModelNet40 dataset without any external data or pre-training, (c) pre-training on the ShapeNet dataset, and (d) migrating the pre-trained model parameters learned on the ShapeNet dataset to the ModelNet40 dataset and performing targeted fine-tuning. On the test set, the pre-trained model of the present invention achieved a classification accuracy of 93.1%, and after fine-tuning it reached 94.2%. A classification accuracy of 86.95% was achieved on the ScanObjectNN dataset, significantly exceeding existing point cloud self-supervised learning methods.

[0126] In one embodiment of the present invention, a computer device is further provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the aforementioned embodiment when executing the computer program.

[0127] In one embodiment of the present invention, a computer-readable storage medium is further provided, wherein the computer-readable storage medium stores a computer program, and wherein the computer program implements the method described in the above embodiment when executed by a processor.

[0128] The pre-trained model obtained through training is applied to downstream tasks. In specific applications, after adding a prediction head dedicated to the corresponding task to the obtained pre-trained model (or fine-tuning it), the final result can be output for the specific task. For example, in the 3D shape classification task, a classification layer is added to the pre-trained model, and the classification result is obtained after the data to be classified is input; in the segmentation task, a segmentation head is added to the pre-trained model, and the segmentation result can be generated after the result to be segmented is input; in 3D object detection, a detection head is introduced based on the pre-trained model to predict the position and category of objects in the point cloud. By improving the feature extraction and representation capabilities of point cloud data, the accuracy and robustness of downstream tasks can be improved.

[0129] The embodiments of the present invention can improve the representation of local and global features of three-dimensional point clouds, improve multi-scale feature extraction and reconstruction capabilities, and improve the final results of downstream tasks through multi-scale masking of local point cloud blocks, design optimization of encoder-decoders, and the use of reconstruction loss functions.

[0130] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein, but is intended to be embodied in the widest possible manner consistent with the principles and novel features disclosed herein.

Claims

1. A multi-scale point cloud pre-training method, characterized in that: The following steps are involved: Select a representative center point from the standardized point cloud data; Select neighboring points for each center point to construct a local point cloud block; According to the predetermined mask ratio, the local point cloud blocks are randomly masked at the lowest scale and gradually continue to mask at higher scales; The masked local point cloud block serves as the input to the encoder-decoder model. The encoder is used to extract point cloud features at different scales, interact with contextual information, and use a skip connection mechanism to upsample and fuse high-level abstract features with low-level detailed features. The encoder outputs a visible token and a shared and trainable mask token. The decoder is used to reconstruct the features of the visible token and the shared and trainable mask token, and then map the visible token and the shared and trainable mask token back to the three-dimensional coordinate space through the linear projection layer to recover the masked local point cloud block; The encoder-decoder model is trained, and after training is completed, the encoder is used as a pre-trained model for downstream tasks.

2. A multi-scale point cloud pre-training method according to claim 1, characterized in that: The farthest point sampling algorithm is used to select representative center points.

3. The multi-scale point cloud pre-training method according to claim 1, characterized in that: Use the k-nearest neighbor algorithm to select neighboring points for each center point and construct a local point cloud block. The formula is: Among them, P represents the selected local point cloud block, X is the input point cloud data, m1 is the number of center points after sampling, and k1 is the number of neighborhood points of each center point.

4. The multi-scale point cloud pre-training method according to claim 1, characterized in that: Use L2Chamfer distance as the reconstruction loss function to reconstruct the mask area. The expression of the reconstruction loss function is: in, Represents the reconstructed point cloud block P rec With the original local point cloud block The error between .

5. The multi-scale point cloud pre-training method according to claim 1, characterized in that: The steps for masking include: The lowest-scale local point cloud block is masked according to the ratio α; At higher scales i On the scale s, use the farthest point sampling algorithm to i-1 The visible local point cloud blocks are downsampled, and the number of center points after downsampling is: m i =m i-1 ×(1-a) Use the KNN algorithm to select k for each downsampled center point i neighborhood points to form a higher-scale feature merging.

6. The point cloud processing method based on self-supervised learning according to claim 1, characterized in that: An asymmetric encoder-decoder architecture is used. The encoder includes multiple scale encoders to extract local and global features. The formula is as follows: Among them, T i is the input feature of the i-th scale, is the output feature after the scale encoder, d i is the dimension of the feature representation of the i-th layer.

7. The point cloud processing method based on self-supervised learning according to claim 1, characterized in that: At the initial scale s i-1 A shared and trainable mask token is appended to the visible token and passed to the decoder for processing, which is defined as: Where, T D represents the output features of the decoder, Indicates that the matrix size is m1×d1, that is, there are m1 tokens, and the feature dimension of each token is d1. T is a visible token. mask A shared and trainable mask token; The decoder is used to reconstruct the masked local point cloud blocks, and the output of the decoder is passed through a linear projection layer using the k-nearest neighbor algorithm to reconstruct the nearest neighbor points of the initial scale.

8. A point cloud processing method based on self-supervised learning according to any one of claims 1 to 7, characterized in that: A prediction head of the corresponding downstream task is added to the obtained pre-trained model to output the final result for the specific task; when used for three-dimensional shape classification tasks, a classification layer is attached to the pre-trained model, and the classification result is obtained after the data to be classified is input; when used for segmentation tasks, a segmentation head is attached to the pre-trained model, and the segmentation result is generated after the result to be segmented is input; when used for three-dimensional object detection, a detection head is attached to the pre-trained model, and the object position and category in the point cloud are generated after the point cloud data to be detected is input.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Aspheric mirror surface shape measuring method and system based on variational auto-encoder

    CN121616761A

  • Method and system for measuring aspheric mirror surface based on variational autoencoder

    CN121616761B

  • Adaptive neighborhood selection-based point cloud self-supervised classification and segmentation method and device, and medium

    CN122024222A

  • Point cloud model training method and electronic equipment

    CN122313196A