Self-supervised learning method of point cloud combining partial-overall perception

Through the joint part-holistic perception point cloud self-supervised learning method, the overall representation is enriched by component information, which solves the problem of ignoring local information in the existing methods and improves the performance of point cloud model in three-dimensional visual tasks.

CN120278221APending Publication Date: 2025-07-08ZHONGBEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510342690.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing self-supervised learning method ignores the impact of some information on object perception in point cloud data, resulting in a degradation of the generalization ability and performance of the model in three-dimensional visual tasks.

Method used

A joint part-holistic perception point cloud self-supervised learning method is adopted to enrich the overall representation through data augmentation and feature extraction through point cloud component information, and combine the point cloud feature extractor network and implicit field learning local features to optimize the network to improve feature capture capabilities.

Benefits of technology

Improves the performance of unsupervised point cloud representations on downstream classification and segmentation tasks, especially on ModelNet40 datasets, and improves the detection accuracy of feature extractors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278221A_ABST
    Figure CN120278221A_ABST
Patent Text Reader

Abstract

The invention discloses a self-supervised learning method of a point cloud combining partial-overall perception, and belongs to the technical field of artificial intelligence. Aiming at the problem that the invariant perception of point cloud features is generally enhanced by expanding point cloud data and the influence of part of information on object perception is ignored in the existing method, the method comprises the following steps: preprocessing the point cloud data, extracting enhanced point cloud and part of point cloud feature vectors, and then enhancing the similarity condition of the point cloud using comparison loss evaluation feature vectors; a part of feature vectors of the point cloud are obtained through a point cloud feature extractor network and a point cloud feature projection head network, the part of feature vectors are put into an implicit field to learn implicit representation of the part, and the feature vectors of the enhanced point cloud are integrated to obtain an integrated overall feature vector and an integrated partial feature vector; and comparing the difference between the two feature vectors, and performing back propagation to optimize the network according to the difference condition. According to the method, the relation between local and overall point cloud is fully combined, and richer point cloud information representation is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a self-supervised learning method for point clouds that combines part-whole perception. Background Art

[0002] Point clouds are usually represented by a set of sparse three-dimensional points and are a basic geometric data structure. It is a commonly used three-dimensional representation in visual recognition and scene understanding and is widely used in various three-dimensional vision tasks such as three-dimensional shape analysis, object classification, object detection, and object segmentation. Point cloud analysis usually involves extracting and distinguishing its geometric features. Ideally, with sufficient three-dimensional annotation data, the process of understanding and differentiating point cloud geometric features becomes simple. However, real-world scenes lack labeled three-dimensional data, and manually annotating three-dimensional data is both time-consuming and laborious. This limits the supervised learning training of 3D vision tasks and reduces the generalization ability and performance of the model. Self-supervised learning is one of the main methods to solve this problem and has been proven effective in the 2D field. Some studies have begun to explore self-supervised representation learning for point clouds and have achieved promising results.

[0003] Self-supervised representation methods based on point clouds can be roughly divided into two categories: generative and discriminative. Generative self-supervised models use generative models (such as autoencoders or generative adversarial networks) to learn the data distribution. These distributions can capture abstract features in the point cloud and generate point cloud feature vectors. However, since the generated self-supervised models focus on learning the entire data distribution of the point cloud, they may not be able to effectively capture some local tiny details and structures in the point cloud. Discriminative models promote representation learning by generating discriminative labels. Most discriminative models enhance the input data and use contrastive learning methods to enrich the internal latent representation. Discriminative models do not require complex network architectures and can obtain good results using an encoder-like structure. However, discriminative methods usually rely on a large number of sample pairs and carefully designed learning strategies, making it challenging to effectively learn the features between local and global. Summary of the Invention

[0004] Aiming at the problem that existing self-supervised learning methods usually enhance the invariant perception of point cloud features by augmenting point cloud data and ignore the influence of part information on object perception, the present invention provides a self-supervised learning method for point clouds that combines part-whole perception, which makes full use of the component information of objects to enrich the overall representation of objects and effectively improves the performance of unsupervised point cloud representation in downstream classification and segmentation tasks.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A self-supervised learning method for point clouds that combines part-whole perception, comprising the following steps:

[0007] Step 1: First, perform the data preprocessing stage, which is divided into two parts:

[0008] (1) Perform data augmentation operation T on the overall point cloud to obtain two augmented point clouds P′ i and P″ i ;

[0009] (2) Divide the overall point cloud P i into 8 quadrants according to the orientation of the coordinate system; each quadrant is a part of the overall point cloud P i to obtain partial point clouds

[0010] Furthermore, in the above Step 1, the operation of dividing the overall point cloud P i into 8 quadrants according to the orientation of the coordinate system is specifically as follows:

[0011] First, take the centroid i of the overall point cloud P as the origin of the three-dimensional Cartesian coordinate system and move it to the coordinate (0, 0, 0); adjust the coordinates of all points in the overall point cloud P i to ensure that the relative positions of all points in the overall point cloud P i with respect to the centroid remain unchanged; then, divide the overall point cloud P i into several parts according to the 8 quadrants of the coordinate system to obtain partial point clouds

[0012] Step 2: Pass the two augmented point clouds P′ i and P″ i through the point cloud feature extractor network and the point cloud feature projection head network respectively to obtain the feature vectors z′ i and z″ i corresponding to the augmented point clouds P′ i and P″ i , and use the contrast loss to evaluate the similarity of the two feature vectors;

[0013] Furthermore, in the above Step 2, the point cloud feature extractor network is the PointNet and DGCNN models, and the point cloud feature projection head network consists of 2 layers of MLP and generates 512-dimensional feature vectors;

[0014] The calculation formula of the contrast loss is:

[0015]

[0016] Among them, σ represents the temperature parameter, N represents the batch size during training, and sim represents the cosine similarity function.

[0017] Step 3: Pass the partial point cloud C i through the point cloud feature extractor network and the point cloud feature projection head network to obtain partial feature vectors Then put the partial feature vectors into the implicit field to learn the implicit representation of the part;

[0018] Furthermore, in the said Step 3, learning the implicit representation of the part is specifically as follows:

[0019] For any point in the partial point cloud calculate the average distance between it and its M nearest neighbor points Then calculate the average local density of all points in the partial point cloud , and the formula is:

[0020]

[0021] where E b is the number of points in each point cloud; is used to calculate the unsigned distance of the point cloud; a represents the serial number of any point in the component point cloud , so the unsigned distance field of the partial point cloud is expressed as:

[0022]

[0023] where q i represents any query point in the three-dimensional space, and N i (*) is the nearest neighbor point to the query point q i in it; the shape contour of this part is obtained through the unsigned distance field; the self-supervised loss for learning the implicit representation is expressed as:

[0024]

[0025] where |Q| represents the number of query points, represents the feature after splicing the query point coordinate q i with the partial feature vector .

[0026] Step 4: Integrate the enhanced point clouds P′ i and P″i The eigenvectors are integrated to obtain an overall integrated eigenvector and the integrated partial eigenvectors Compare the differences between the two eigenvectors, and optimize the network through backpropagation according to the differences;

[0027] Furthermore, in step 4, the formula for comparing the differences between the two eigenvectors is:

[0028]

[0029] where N represents different point cloud data, and ||·||2 represents the 2-norm.

[0030] Compared with the prior art, the present invention has the following advantages:

[0031] The present invention shows superior detection performance in both two feature extraction backbone networks (PointNet and DGCNN) on the ModelNet40 dataset. For proxy tasks, generative models, modal methods, and contrastive learning methods, these two feature extractors are 0.3% and 0.2% higher than the best methods respectively in terms of test accuracy, and also show better performance than the original PointNet and DGCNN supervised learning benchmarks. In addition, for OcCo, GSPCon, Tsai, etc. and the PoCC method that focus on learning local feature information, although they capture local features through random masking or sampling, they emphasize learning specific attributes within a small range of the point cloud. Paying attention to the local information of specific key points may not effectively capture the overall structure or the relationship between different local parts. On the other hand, the present invention emphasizes the spatial features of the point cloud, enabling the model to understand the structure and function of the entire object. Description of the Drawings

[0032] Figure 1 The flowchart of the method provided for the embodiment;

[0033] Figure 2 (a) Divide it into 8 quadrants according to the coordinates of the points; (b) Visualization of the part division result; the circled part represents the change in the topological structure;

[0034] Figure 3 The generated implicit field map;

[0035] Figure 4 The component segmentation result diagram of the embodiment of the present invention. Detailed Description of the Invention

[0036] To gain a deep understanding of the present invention, we will describe it comprehensively and in detail. However, the present invention has multiple implementation manners and is not limited to the specific examples listed herein. The presentation of these examples aims to deepen the comprehensive understanding of the disclosed content of the present invention.

[0037] A self-supervised learning method for point clouds that combines part-whole perception, as Figure 1 shown, includes the following steps:

[0038] Step 1: First, perform the data preprocessing stage, which is divided into two parts:

[0039] (1) Perform data augmentation operation T on the whole point cloud to obtain two augmented point clouds P′ i and P″ i ;

[0040] (2) Divide the whole point cloud P i into 8 quadrants according to the orientation of the coordinate system; each quadrant is a part of the whole point cloud P i to obtain partial point clouds as Figure 2 shown;

[0041] Furthermore, in the above Step 1, the specific operation of dividing the whole point cloud P i into 8 quadrants according to the orientation of the coordinate system is as follows:

[0042] First, take the centroid i of the whole point cloud P as the origin of the three-dimensional Cartesian coordinate system and move it to the coordinate (0, 0, 0); adjust the coordinates of all points in the whole point cloud P i to ensure that the relative positions of all points in the whole point cloud P i with respect to the centroid remain unchanged; then, divide the whole point cloud P i into several parts according to the 8 quadrants of the coordinate system to obtain partial point clouds

[0043] Step 2: Pass the two augmented point clouds P′ i and P″ i through the point cloud feature extractor network and the point cloud feature projection head network respectively to obtain the feature vectors z i and z i corresponding to the augmented point clouds P′ i and P″ i , and use the contrastive loss to evaluate the similarity of the two feature vectors;

[0044] Furthermore, in the above Step 2, the point cloud feature extractor network is the PointNet and DGCNN models, and the point cloud feature projection head network is composed of 2 layers of MLP and generates 512-dimensional feature vectors;

[0045] The calculation formula of the contrastive loss is as follows:

[0046]

[0047] Among them, σ represents the temperature parameter, N represents the batch size during training, and sim represents the cosine similarity function.

[0048] Step 3: Pass the partial point cloud C i through the point cloud feature extractor network and the point cloud feature projection head network to obtain partial feature vectors Then put the partial feature vectors into the implicit field to learn the implicit representation of the part. The implicit field is as Figure 3 shown;

[0049] Furthermore, in the above Step 3, learning the implicit representation of the part is specifically as follows:

[0050] For any point in the partial point cloud calculate the average distance between it and its M nearest neighbor points Then calculate the average local density of all points in the partial point cloud . The formula is:

[0051]

[0052] Among them, E b is the number of points in each point cloud; is used to calculate the unsigned distance of the point cloud; a represents the serial number of any point in the part point cloud . Therefore, the unsigned distance field of the partial point cloud is expressed as:

[0053]

[0054] Among them, q i represents any query point in three-dimensional space, N i (*) is the nearest neighbor point to the query point q i in; the shape contour of this part is obtained through the unsigned distance field; the self-supervised loss for learning the implicit representation is expressed as:

[0055]

[0056] Among them, |Q| represents the number of query points, represents the query point coordinate qi The features after splicing with partial feature vectors Step 4: Integrate the enhanced point cloud P′

[0057] and P″ i to obtain the integrated overall feature vector i and the integrated partial feature vector Compare the differences between the two feature vectors, and optimize the network through backpropagation according to the differences;

[0058]

[0059]

[0060]

[0061] where N represents different point cloud data, and ||·||2 represents the 2-norm.

[0061] In the experiment, PointNet and DGCNN are used as point cloud feature extractors At the same time, a 2-layer MLP is used as the projection head to generate a 512-dimensional feature vector. The Adam optimizer is used, and the weight decay is 1×10 -4 , and the initial learning rate is 1×10 -3 . The learning rate scheduler is cosine annealing, and the model is trained end-to-end within 200 epochs. During pre-training, the weights of the shared point cloud feature extractor and the point cloud feature projection head are shared. After pre-training, all downstream tasks are performed on the pre-trained point cloud feature extractor . The correlation analysis of the experimental part and the overall features is shown in Table 1, and the segmentation results are as Figure 4 shown.

[0062] Table 1 Correlation analysis of part and whole

[0063]

[0064] Table 2 Comparison of linear classification results on the ModelNet40 dataset

[0065]

[0066] The present invention is compared with the current mainstream self-supervised representation learning methods: Jigsaw3D and Rotation are both based on proxy tasks. STRL, Self-Contrast, PoCCA, and CLR-GAM are all based on contrastive learning methods. OcCo, GSPCon, and Tsai et al. are based on generative models. CrossPoint, Lee et al., TCMSS, and CrossNet are all based on cross-modal methods. In addition, for the sake of convenience of representation, the abbreviations Sup, PT, CL, GM, CM, and PCF are used in the experiments to represent the methods based on supervised learning, proxy tasks, contrastive learning, generative models, cross-modal methods, and feature analysis respectively. The experimental results are shown in Table 2;

[0067] It can be seen that the present invention shows superior detection performance in both of the two feature extraction backbone networks (PointNet and DGCNN) on the ModelNet40 dataset. For the proxy task, generative model, modal method, and contrastive learning method, these two feature extractors are respectively 0.3% and 0.2% higher than the best method in terms of test accuracy, and also show better performance than the original PointNet and DGCNN supervised learning benchmarks. In addition, for the OcCo, GSPCon, Tsai et al. and PoCC methods that focus on learning local feature information, although they capture local features through random masking or sampling, they emphasize learning specific attributes within a small range of the point cloud. Focusing on the local information of specific key points may not effectively capture the overall structure or the relationship between different locals. On the other hand, the present invention emphasizes the spatial features of the point cloud, enabling the model to understand the structure and function of the entire object.

[0068] The content not detailedly described in the specification of the present invention belongs to the prior art well-known to those skilled in the art. Although the illustrative specific embodiments of the present invention are described above for the convenience of those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.

Claims

1. A self-supervised learning method for point clouds that combines part-whole perception, characterized in that, The method includes the following steps: Step 1: First, perform the data preprocessing stage, which is divided into two parts: (1)Perform data augmentation operation T on the overall point cloud to obtain two augmented point clouds P i ' and P i ''; (2) Divide the overall point cloud P according to the orientation of the coordinate system i into 8 quadrants; each quadrant is a part of the overall point cloud P i and obtain partial point clouds Step 2: Pass the two enhanced point clouds P i ′ and P i ″ through the point cloud feature extractor network and the point cloud feature projection head network respectively to obtain the feature vectors z′ i ′ and P i ″ corresponding to z′ i and z″ i , and use the contrastive loss to evaluate the similarity of the two feature vectors; Step 3: Take partial point cloud C i through the point cloud feature extractor network and the point cloud feature projection head network to obtain partial feature vectors Then put the partial feature vectors into the implicit field to learn the partial implicit representation; Step 4: Integrate the enhanced point clouds P i ′ and P i ″ to obtain the integrated overall feature vector and the integrated partial feature vector Compare the differences between the two feature vectors and optimize the network through backpropagation according to the differences.

2. The self-supervised learning method for point clouds integrating part-whole perception according to claim 1, characterized in that, The said Step 1: First, perform the data preprocessing stage, which is divided into two parts: (1) Perform data augmentation operation T on the overall point cloud to obtain two augmented point clouds P i ' and P i "; (2) Divide the overall point cloud P according to the orientation of the coordinate system i into 8 quadrants; each quadrant is a part of the overall point cloud P i to obtain a partial point cloud The specific operation is as follows: First, move the centroid i of the overall point cloud P to the origin of the three-dimensional Cartesian coordinate system, i.e., to the coordinates (0, 0, 0); adjust the coordinates of all points in the overall point cloud P i to ensure that the relative positions of all points in the overall point cloud P i with respect to the centroid remain unchanged; then, divide the overall point cloud P i into several parts according to the 8 quadrants of the coordinate system to obtain partial point clouds 3. The self-supervised learning method for point clouds with combined part-whole perception according to claim 2, wherein Step 2: Pass the two enhanced point clouds P i ′ and P i ″ through the point cloud feature extractor network and the point cloud feature projection head network respectively to obtain the feature vectors z′ i ′ and P i ″ corresponding to z′ i and z″ i , and use the contrastive loss to evaluate the similarity of the two feature vectors; Among them Point cloud feature extractor network For the PointNet and DGCNN models, the point cloud feature projection head network Consists of 2 layers of MLP, generating a 512-dimensional feature vector; The calculation formula of the contrastive loss is: Where, σ represents the temperature parameter, N represents the batch size during training, and sim represents the cosine similarity function.

4. A self-supervised learning method for point clouds that combines part-whole perception, characterized in that, Step 3: The partial point cloud C i passes through the point cloud feature extractor network and the point cloud feature projection head network respectively to obtain partial feature vectors Then, the partial feature vectors are put into the implicit field to learn the partial implicit representation; wherein, the learning part of the implicit representation is specifically: For a partial point cloud For any one point Calculate The average distance between it and its M nearest neighbor points Then calculate the partial point cloud The average local density of all points in, the formula is: Among them, E b is the number of points in each point cloud; is used to calculate the unsigned distance of the point cloud; a represents any point in the component point cloud sequence number, therefore, the unsigned distance field of part of the point cloud is expressed as: is expressed as: where q i represents an arbitrary query point in three-dimensional space, N i (*) is the nearest neighbor to the query point q i ; the shape contour of this part is obtained through the signed distance field; the self-supervised loss for learning the implicit representation is expressed as: Among them, |Q| represents the number of query points, represents the query point coordinate q i and the partial feature vector after splicing.

5. The self-supervised learning method of a point cloud jointly perceiving part and whole according to claim 4, characterized in that, Step 4: Integrate the enhanced point clouds P i ′ and P i ″ to obtain the integrated overall feature vector and the integrated partial feature vector Compare the differences between the two feature vectors, and optimize the network by backpropagation according to the differences. Among them, the formula for comparing the differences between the two feature vectors is: Where, N represents different point cloud data, and ||·||2 represents the 2-norm.