Point cloud feature extraction method based on isovariant group representation

Through the method based on isovariant group representation, point cloud data is encoded using spherical harmonic function and the isovariant feature extractor EFE is designed, combined with the traditional point cloud neural network G, the problem of insufficient degeneration such as rotation in point cloud feature extraction is solved, and stable feature extraction and efficient downstream task application is achieved.

CN120388183APending Publication Date: 2025-07-29NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510443287.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing point cloud feature extraction methods lack rotation and other degeneration, poor generalization and robustness, difficult to effectively integrate with classical neural networks, and insufficient feature expression capabilities.

Method used

Using a method based on isovariant group representation, point cloud data is encoded using spherical harmonic function, and an isovariant feature extractor EFE is designed, combined with traditional point cloud neural network G, to construct a combined neural network EFE-G to achieve rotationally unchanged feature extraction.

Benefits of technology

It realizes stable feature extraction on rotating data, improves feature expression capabilities, can seamlessly integrate with classic neural networks, and is suitable for various downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388183A_ABST
    Figure CN120388183A_ABST
Patent Text Reader

Abstract

The invention discloses a point cloud feature extraction method based on isovariant group representation. The method comprises the steps of 1, collecting a 3D point cloud data set; step 2, enhancing and labeling the obtained point cloud data, performing sampling scaling, rotation, noise addition and other operations on point cloud three-dimensional coordinates, and labeling the data according to a specific downstream task; 3, encoding the data by using a spherical harmonic function to obtain a point embedding X; 4, performing feature extraction on the point embedding X by using an equivariant feature extractor (EFE), and performing connection and addition on the point embedding X and a spring layer of the point embedding to obtain an intermediate result T; 5, repeating the step 4 for multiple times, and adding the obtained result and the point-embedded spring layer to obtain an extracted isovariant feature A; step 6, inputting the isovariant feature A into a point cloud neural network G according to a downstream task demand, and training to obtain a combined neural network EFE-G; and step 7, performing feature extraction on the point cloud data by using a combined neural network EFE-G with equal denaturation and completing a subsequent downstream task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for feature extraction of 3D point clouds, and particularly to a neural network method with rotational equivariance. Background Art

[0002] In recent years, the rapid development of high-precision sensors such as lidar and Kinect has significantly promoted the acquisition of 3D data, enabling its wide application in various fields. As a commonly used 3D data representation form, point clouds exhibit excellent spatial expression capabilities. Therefore, point clouds have become the main data format for representing the 3D world and a key research tool in 3D graphics tasks.

[0003] Among many time series analysis methods, recent models such as PointNet and DGCNN use convolutional operations to extract point cloud features, ensuring high-quality features while addressing the challenge of permutation invariance. However, the features extracted by these methods have limited generalization ability and robustness because they are highly sensitive to rotational transformations and lack SO(3) invariance or equivariance. For spatial objects, regardless of the angle from which the point cloud data is collected, the results obtained from the final downstream tasks should be the same, but methods without equivariance perform poorly on rotated data. Designing methods with equivariance is a current popular issue. Literature: Wang Y, Sun Y, Liu Z, et al. Dynamic graph cnn for learning on pointclouds. ACM Transactions on Graphics (tog), 2019, 38(5): 1 - 12.

[0004] Some recent point cloud studies have adopted equivariant methods for feature extraction and network construction. These methods use point-by-point feature analysis and use geometric features such as included angles and dihedral angles for feature extraction. These methods have good representation quality for local and global features and have rotational equivariance, performing well on rotated data. However, these methods only rely on low-order equivariant representations and only use partial geometric information mathematically, which has insufficient expression ability in linear spaces and cannot capture the complex geometric properties of point clouds in detail. The extracted features often have poor semantic quality and low accuracy. At the same time, it is difficult for these methods to be effectively integrated with existing methods.

[0005] To overcome the bottleneck in feature quality of these methods, Congyue Deng et al. improved the functions of traditional neural networks by designing these functions to be equivariant and replacing traditional neurons with vector neurons, so that the neural networks designed by these improved functions also have rotational equivariance. This method does have equivariance and improves the performance of downstream tasks at the same time, but its feature expression ability is still limited and it is difficult to combine with complex classical networks. Literature: Deng C, Litany O, Duan Y, et al. Vector neurons: A general framework for so(3)-equivariant networks. Proceedings of the IEEE / CVF International Conference on Computer Vision. 2021: 12200-12209. Summary of the Invention

[0006] Object of the Invention: To overcome the problems that existing classical methods basically do not have rotational equivariance, poor generalization ability and robustness, and are sensitive to rotational data. At the same time, solve the problems of insufficient expression ability, low feature quality of existing equivariant methods, and difficulty in embedding with classical neural networks.

[0007] To solve the above technical problems, the present invention discloses a method for extracting point cloud features based on equivariant group representation. This method draws inspiration from the concept of equivariant group representation in physics, uses spherical harmonic functions to encode 3D direction vectors into a high-dimensional space, thereby obtaining an equivariant representation and further deriving equivariant features of points. This method uses only a very small number of learnable parameters, and the extracted equivariant features can be easily integrated into classical neural networks, making it applicable to a wide range of downstream tasks in various scenarios and playing an important role in fields such as point cloud classification, point cloud segmentation, and point cloud registration. This method includes the following steps:

[0008] Step 1, collect a 3D point cloud dataset;

[0009] Step 2, perform augmentation and annotation on the obtained point cloud data, perform operations such as sampling, scaling, rotation, and adding noise on the three-dimensional coordinates of the point cloud, and annotate the data according to specific downstream tasks;

[0010] Step 3, use spherical harmonic functions to encode the data to obtain point embeddings X;

[0011] Step 4, use the designed equivariant feature extractor EFE to extract features from the point embeddings X, and add them to the skip connection of the point embeddings to obtain an intermediate result T;

[0012] Step 5: Repeat Step 4 multiple times, and add the obtained results to the skip connections of the point embeddings to obtain the extracted equivariant feature A.

[0013] Step 6: Input the equivariant feature A into the traditional point cloud neural network G according to the requirements of the downstream task, and perform training to obtain the combined neural network EFE-G.

[0014] Step 7: Use the combined neural network EFE-G with equivariance to extract features from the point cloud data and complete subsequent downstream tasks.

[0015] In Step 1, the 3D point cloud dataset will be collected. The LiDAR technology can be used to scan the object from multiple angles to obtain the point cloud data; a synthetic dataset can also be used if necessary. Finally, multiple point cloud datasets with the shape of n*3 are obtained, where n is the number of points in the point cloud, and each point contains three-dimensional spatial coordinates.

[0016] In Step 2, the original data is processed. The necessary processing is centering, that is, taking the weighted average coordinates of all points in each point cloud as the origin, which can eliminate the impact of translation operations on equivariance. Subsequently, some data augmentation operations can be selected, such as randomly selecting some points to add noise, and the noise can be sampled from a normal distribution with a mean of 0 and a small variance. And according to the requirements of the downstream task, such as classification and segmentation tasks, the data is labeled.

[0017] In Step 3, for each sampled point p in the augmented data i , the embedding vector x is represented using spherical harmonics. i,K Under the setting of the hyperparameter maximum order l max , for each order l, there are 2l + 1 spherical harmonic functions. Therefore, for 0 ≤ l ≤ l max , there are a total of (l max + 1) 2 basis functions. In addition, each basis function has C l channels. Therefore, the dimension of the point embedding is (l max + 1) 2 × C l . Initially, the point embedding is represented as x i,0 , where l = 1, the first basis vector is initialized to a constant 1, and the second to fourth basis vectors are defined by the eigenvectors corresponding to the minimum eigenvalue of the covariance matrix between the point p i and its neighboring points p j ∈ N i . After passing through K equivariant modules, the obtained feature is x i,K .

[0018] In step 4, the designed equivariant feature extractor (EFE) is used to extract features from the point embedding X, and the skip connection of the point embedding is added to obtain the intermediate result T. Several equivariant functions are designed in the EFE to replace the functions of the traditional neural network.

[0019] Equivariant linear layer:

[0020] Linear(x) = [x 0 W 0 + b; x 1 W 1 ; …; x l W l ,

[0021] where W and b represent the weight and bias respectively, the superscript l represents the equivariant order, and x l = x[l:2l + 1]. This formula indicates that when a point embedding passes through the equivariant linear layer, it is split and matrix multiplication is performed using different weights according to the superscript l.

[0022] Equivariant layer normalization:

[0023]

[0024] where RMS represents the root mean square calculation and norm represents the L2 normalization. The formula indicates that when a point embedding passes through the equivariant layer normalization, normalization is performed using different weights according to the superscript l.

[0025] Equivariant gate activation function:

[0026]

[0027] where Sigmoid represents the activation function. The formula indicates that when a point embedding passes through the equivariant layer gate activation function, calculations are performed using different weights according to the superscript l.

[0028] Specifically, for the point embedding X and the equivariant feature extractor EFE, the intermediate result is calculated as T = EFE(X) + X.

[0029] In step 5, step 4 is repeated multiple times, and the obtained result is added to the skip connection of the point embedding to obtain the extracted equivariant feature A. Specifically, for the initial point embedding x i,0 , after repeating step 4 K times, the point embedding obtained is x i,K . If the repetition is stopped at this time, x i,K is output as the equivariant feature A.

[0030] In step 6, the equivariant feature A is input into the traditional point cloud neural network G according to the requirements of the downstream task and trained to obtain the combined neural network EFE-G. The traditional point cloud neural network G can be any previously proposed point cloud network. As long as it can accept the input of n*3 point cloud data, it can be directly connected to the equivariant feature extractor EFE to obtain an equivariant combined network EFE-G. After being trained, this network can still be applied to the original downstream task and has equivariance.

[0031] In step 7, the equivariant combined neural network EFE-G is used to extract features from the point cloud data, output the results, and complete the subsequent downstream tasks. No matter what rotation is performed on the input data, the output results remain unchanged.

[0032] Beneficial effects: The significant advantage of the present invention is that it can conveniently and efficiently extract the equivariant features of point cloud data. A new equivariant function is designed to construct the network. The extracted features are stable for rotated data and have strong expressive ability. At the same time, the network structure is simple, uses extremely few learnable parameters, has good portability, and can be combined with many classic point cloud neural networks to form an equivariant neural network, so as to be applied to various downstream tasks. Description of the Drawings

[0033] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0034] Figure 1 It is a flow chart for constructing the equivariant feature extractor EFE and training the equivariant neural network EFE-G of the present invention.

[0035] Figure 2 It is a schematic structural diagram of the equivariant feature extractor EFE in the present invention.

[0036] Figure 3 It is the accuracy rate of the classification experiment of the present invention on the ModelNet40 dataset and its rotated test set.

[0037] Figure 4 It is the accuracy rate of the segmentation experiment of the present invention on the ShapeNet dataset and its rotated test set.

[0038] Figure 5 It is a comparison of the model parameter quantities of the present invention. Detailed Embodiments

[0039] In order to make the purpose, technical solutions, and advantages of the present invention clearer and more distinct, this chapter further describes the invention in detail in conjunction with the drawings.

[0040] Figure 1This is the flowchart for constructing the equivariant feature extractor EFE and training the equivariant neural network EFE-G in the present invention, including six steps.

[0041] In the first step, a 3D point cloud dataset will be collected. The Lidar technology can be used to scan an object from multiple angles to obtain point cloud data; synthetic datasets can also be used if necessary. Eventually, multiple point cloud datasets with the shape of n*3 are obtained, where n is the number of points in the point cloud and contains the three-dimensional spatial coordinates of each point.

[0042] In the second step, the original data is processed. The necessary processing is centering, that is, taking the weighted average coordinates of all points in each point cloud as the origin, which can eliminate the impact of translation operations on equivariance. Specifically, the three-dimensional coordinates of each point cloud are added according to the corresponding dimensions and then averaged to obtain the average coordinates, and then the coordinates of each point are subtracted by the average coordinates to complete centering. Subsequently, some data augmentation operations can be selected, such as randomly selecting some points to add noise, and the noise can be sampled from a normal distribution with a mean of 0 and a small variance. And according to the needs of downstream tasks, such as classification and segmentation tasks, the data is labeled.

[0043] In the third step, for each sampled point p in the augmented data i , the spherical harmonics are used to represent the embedding vector x i,K . Under the setting of the hyperparameter maximum order l max , for each order l, there are 2l + 1 spherical harmonic functions. Therefore, for 0 ≤ l ≤ l max , there are a total of (l max + 1) 2 basis functions. In addition, each basis function has C l channels. Therefore, the dimension of the point embedding is (l max + 1) 2 × C l . Initially, the point embedding is represented as x i,0 , where l = 1, the first basis vector is initialized to a constant 1, and the second to fourth basis vectors are defined by the eigenvectors corresponding to the minimum eigenvalue of the covariance matrix between the point p i and its neighboring points p j ∈ N i . After passing through K equivariant modules, the resulting feature is x i,K .

[0044] In the fourth step, the designed equivariant feature extractor EFE is used to extract features from the point embedding X and add them to the skip connection of the point embedding to obtain the intermediate result T. The schematic diagram of the basic structure of EFE is as Figure 2As shown, it takes point embeddings as input. After passing through a carefully designed equivariant module and some equivariant functions, an intermediate output is obtained, which is regarded as completing one layer of feature extraction. After K times of feature extraction with the set hyperparameters, the final equivariant feature A is output, which is used to combine with a classical neural network and complete downstream tasks later.

[0045] In EFE, several functions with equivariance are designed to replace the functions of traditional neural networks.

[0046] Equivariant linear layer:

[0047] Linear(x) = [x 0 W 0 + b; x 1 W 1 ; …; x l W l ,

[0048] where W and b represent the weight and bias respectively, the superscript l represents the equivariant order, and x l = x[l:2l + 1]. This formula indicates that when a point embedding passes through the equivariant linear layer, it is split and matrix multiplications are performed using different weights according to the superscript l. The linear layer is a basic component of the neural network and plays an important role in transmitting information and updating neurons. Since the point embedding is already a designed feature with equivariance, an equivariant linear layer needs to be designed. Performing matrix multiplications on the features using different weights according to the superscript l can maintain equivariance channel by channel.

[0049] Equivariant layer normalization:

[0050]

[0051] where RMS represents the root mean square calculation and norm represents the L2 normalization. The formula shows that when a point embedding passes through the equivariant layer normalization, normalizations are performed using different weights according to the superscript l. Normalization is an important method in neural networks. As the network depth increases, the sum of the information accumulated by each node becomes larger and larger, which may cause problems such as the increase in the order of magnitude of features, ultimately leading to poor learning effects and low computational efficiency. Using normalization can control the numerical size of features, and at the same time, performing normalizations on the features using different weights according to the superscript l can maintain equivariance channel by channel, which is helpful for feature learning.

[0052] Equivariant gate activation function:

[0053]

[0054] Among them, Sigmoid represents the activation function. The formula shows that when a point embedding passes through the equivariant layer gate activation function, different weights are used for calculation according to the superscript l. The activation function is an essential part of the neural network. The superposition of linear combinations is still a linear combination and cannot fit more information. Therefore, an activation function is needed to introduce non-linearity. Non-linearity can synthesize the information of linear functions, fit more complex function structures, enhance the expression ability of the model, and calculating the activation function for the features with different weights according to the superscript l can maintain equivariance channel by channel, which is helpful for feature learning.

[0055] In addition to the above functions, the CG tensor product (Clebsch-Gordan tensor product) is used in EFE to replace convolution or matrix operations for information transmission. The CG tensor product is a vector calculation method with equivariance, and its formula is as follows:

[0056]

[0057] where f and g are feature vectors, and l and m represent the splitting method of the vectors and the indices of the vectors after splitting respectively. C is the Clebsch-Gordan coefficient, which is non-zero only when |l1 - l2| ≤ l ≤ |l1 + l2|. This method ensures the efficient transmission of feature information while not destroying the equivariance of the structure.

[0058] In the fifth step, the equivariant feature A is input into the traditional point cloud neural network G according to the requirements of the downstream task and trained. The training method and loss function can directly use the methods of the neural network G. Finally, the combined neural network EFE-G is obtained. The traditional point cloud neural network G can be any previously proposed point cloud network. As long as it can accept the input of n*3 point cloud data, it can be directly connected to the equivariant feature extractor EFE to obtain an equivariant combined network EFE-G. After training, this network can still be applied to the original downstream task and has equivariance.

[0059] In the sixth step, the equivariant combined neural network EFE-G is used to extract features from the point cloud data, output the results, and complete the subsequent downstream tasks.

[0060] After training the model, when predicting new point cloud data, the new data is encoded through point embedding and then input into the EFE equivariant feature extractor. After obtaining the equivariant features, they are input into the classical point cloud neural network G to complete the downstream task. No matter what rotation transformation is performed on the input data, the output results remain unchanged.

[0061] Embodiment

[0062] To verify the effectiveness of the model, instance verification was carried out on the ModelNet40 dataset and the ShapeNet dataset, and two core tasks in point cloud processing were evaluated: 3D shape classification and part segmentation. To highlight the effectiveness of the method in equivariant feature extraction, three different training / testing settings were adopted: training and testing the network by rotating around the z-axis (z / z), where both the training set and the testing set were rotated around the z-axis in this setting; training by rotating around the z-axis and testing under arbitrary rotations (z / SO(3)), where the training set was rotated around the z-axis and the testing set was rotated arbitrarily in space in this setting; and training and testing under arbitrary rotations (SO(3) / SO(3)), where both the training set and the testing set were rotated arbitrarily in space in this setting. Here, z refers to data augmentation with rotation only around the z-axis, and SO(3) represents arbitrary rotation. All rotations were dynamically generated during training to compare the equivariance achieved through the structure with the equivariance learned through data augmentation. During testing, each shape was presented with a single rotation.

[0063] In the classification and segmentation experiments, EFE was integrated with classical point cloud network architectures, introducing only a small number of learnable parameters without modifying the original network structure. This formed a network architecture with rotational equivariance and achieved satisfactory performance. Considering the stability and influence of various classical networks, PointNet and DGCNN were mainly selected as the basic architectures of the hybrid model because they have excellent stability, robustness, and wide influence, and can verify the effectiveness of the method.

[0064] First, the model performance was evaluated on the synthetic ModelNet40 dataset, which contains CAD models from 40 categories, such as airplanes, bottles, chairs, dressers, vases, etc. The data preprocessed by PointNet was used, including 9843 models for training and 2468 models for testing. In this experiment, the point cloud with a sampling number of 1024 was used. Each point was represented by (x, y, z) coordinates in Euclidean space. Figure 3 The performance of the EFE-G hybrid model was reported and compared. Compared with non-equivariant networks, the EFE model always achieved good results in all three settings, showing its robustness to rotation. Especially in the z / SO(3) setting, where the testing set contains rotations not seen in the training set and classical methods performed poorly, the EFE model remained stable. Even in the SO(3) / SO(3) setting, where extensive rotational data augmentation was applied during training, the performance of rotation-sensitive networks was still inferior to that of the EFE equivariant network, which proves the effectiveness of the method of the present invention.

[0065] Shape part segmentation is a more challenging task compared to object classification. Experiments were conducted on the ShapeNet dataset to evaluate its performance in the part segmentation task. The dataset contains 16,881 models from 16 categories, annotated with 50 parts. Additionally, there is no overlap between the training set and the test set, and 2,048 points with (x, y, z) coordinates are sampled as the model input. Figure 4 The segmentation results of different methods in the z / SO(3) and SO(3) / SO(3) scenarios are shown. Classic methods, such as PointNet and DGCNN, show vulnerability to rotation. This further confirms that, despite the application of rotational data augmentation, these methods lacking inherent rotational equivariance still perform poorly. In contrast, the equivariant model of the present invention achieves excellent results.

[0066] Figure 5 This is a comparison of the model parameters of the present invention. Different from other fixed architectures, the present invention is lightweight and highly portable, and can be easily integrated with classic methods to achieve rotational equivariance. The VN method is similar to the present invention and is an equivariant method that can be conveniently applied to classic models. Figure 5 The model size was evaluated, showing that the present invention introduces only a small increase in the number of learnable parameters compared to classic models. Specifically, the EFE models corresponding to PointNet and DGCNN only increase by 0.85% and 0.50% respectively, while their corresponding VN models show a significantly larger increase in the number of parameters. In contrast, the present invention provides a more lightweight design while being seamlessly integrated with classic models.

[0067] The present invention provides a method for extracting point cloud features based on equivariant group representation. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using existing technologies.

Claims

1. A method for extracting point cloud features based on equivariant group representation, characterized in that, It includes the following steps: Step 1, collect a 3D point cloud dataset; Step 2, perform augmentation and annotation on the obtained point cloud data, perform operations such as sampling and scaling, rotation, and adding noise on the three-dimensional coordinates of the point cloud, and annotate the data according to specific downstream tasks; Step 3, use spherical harmonics to encode the data to obtain point embedding X; Step 4, use the designed equivariant feature extractor EFE to extract features from the point embedding X, and add it to the skip connection of the point embedding to obtain an intermediate result T; Step 5, repeat Step 4 multiple times, and add the obtained result to the skip connection of the point embedding to obtain the extracted equivariant feature A; Step 6, input the equivariant feature A into the traditional point cloud neural network G according to the downstream task requirements, and perform training to obtain the combined neural network EFE-G; Step 7, use the combined neural network EFE-G with equivariance to extract features from the point cloud data and complete subsequent downstream tasks.

2. The method according to claim 1, wherein In Step 1, the 3D point cloud dataset will be collected. The LiDAR technology can be used to scan the object from multiple angles to obtain point cloud data; if necessary, a synthetic dataset can also be used. Eventually, multiple point cloud datasets with the shape of n*3 are obtained, where n is the number of points in the point cloud, and each point contains three-dimensional spatial coordinates.

3. The method according to claim 2, wherein In Step 2, the original data is processed. The necessary processing is centering, that is, taking the weighted average coordinates of all points in each point cloud as the origin, which can eliminate the impact of translation operations on equivariance. Subsequently, some data augmentation operations can be selected, such as randomly selecting some points to add noise, and the noise can be sampled from a normal distribution with a mean of 0 and a small variance. And according to the downstream task requirements, such as classification and segmentation tasks, etc., the data is annotated.

4. The method according to claim 3, characterized in that In step 3, for each sampling point p in the enhanced data i , the embedding vector x is represented using spherical harmonics i,K . Under the setting of the hyperparameter maximum order l max , for each order l, there are 2l + 1 spherical harmonic functions. Therefore, for 0 ≤ l ≤ l max , there are a total of (l max + 1) 2 basis functions. In addition, each basis function has C l channels. Therefore, the dimension of the point embedding is (l max + 1) 2 × C l . Initially, the point embedding is represented as x i,0 , where l = 1. The first basis vector is initialized to the constant 1, and the second to fourth basis vectors are defined by the eigenvectors corresponding to the minimum eigenvalue of the covariance matrix between the point p i and its neighboring points p j ∈ N i . After passing through K equivariant modules, the resulting feature is x i,K .

5. The method according to claim 4, characterized in that, In Step 4, use the designed equivariant feature extractor EFE to extract features from the point embedding X, and add it to the skip connection of the point embedding to obtain an intermediate result T. Several equivariant functions are designed in EFE to replace the functions of traditional neural networks, Equivariant linear layer: Linear(x) = [x 0 W 0 + b; x 1 W 1 ; …; x l W l , where \(W\) and \(b\) represent weights and biases respectively, and the superscript \(l\) represents the order of equivariance, \(x l = x[l:2l + 1], indicating that when a point embedding passes through an equivariant linear layer, it is split according to the superscript and matrix multiplications are performed using different weights respectively; Equivariant layer normalization: where RMS represents the root mean square calculation, and norm represents L2 normalization. The formula shows that when a point embedding passes through equivariant layer normalization, different weights are used for normalization according to the superscript; Equivariant gate activation function: where Sigmoid represents the activation function. The formula shows that when a point embedding passes through the equivariant layer gate activation function, different weights are used for calculation according to the superscript l; Specifically, for the point embedding X and the equivariant feature extractor EFE, the intermediate result is calculated as T = EFE(X) + X.

6. The method according to claim 5, characterized in that, In step 5, step 4 is repeated multiple times, and the obtained result is added to the skip connection of the point embedding to obtain the extracted equivariant feature A. Specifically, for the initial point embedding x i,0 , after repeating step 4 K times, the point embedding obtained is x i,K . If the repetition is stopped at this time, then x i,K is output as the equivariant feature A.

7. The method according to claim 6, characterized in that, In Step 6, the equivariant feature A is input into the traditional point cloud neural network G according to the downstream task requirements, and training is performed to obtain the combined neural network EFE-G. The traditional point cloud neural network G can be any previously proposed point cloud network. As long as it can accept the input of n*3 point cloud data, it can be directly connected to the equivariant feature extractor EFE to obtain an equivariant combined network EFE-G. After training, this network can still be applied to the original downstream tasks and has equivariance.

8. The method according to claim 7, wherein In step 7, the combined neural network EFE-G with equivariance is used to extract features from the point cloud data, output the results, and complete subsequent downstream tasks. Regardless of any rotation of the input data, the output results remain unchanged.