A lightweight multi-stage point cloud classification method based on a fusion Transformer network

By integrating the lightweight multi-stage method of Transformer network in point cloud classification tasks, combining efficient feature aggregation branches and learnable position coding, the problem of large amount of parameters in point cloud classification is solved, and efficient point cloud classification accuracy and model stability are achieved.

CN119445181BActive Publication Date: 2025-05-30HOHAI UNIV

Patent Information

Application Number
CN202411057437.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2025-05-30
Estimated Expiration
2044-08-02

AI Technical Summary

Technical Problem

The existing Transformer network has problems with structural complexity and large number of parameters in point cloud classification tasks, and the exploration of point cloud location encoding is not in-depth enough, resulting in the classification accuracy not high enough while ensuring low parameter volume.

Method used

A lightweight multi-stage point cloud classification method based on a converged Transformer network is proposed. By integrating efficient feature aggregation branches and Transformer branches, learningable position coding is adopted, and SE module is introduced into the original point embedding module, and a four-stage learning strategy is adopted to improve classification accuracy.

Benefits of technology

It realizes improving point cloud classification accuracy under low parameter quantity conditions, enhances the stability and performance of the model, avoids the loss of point cloud information, and achieves a good balance between lightweight design and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119445181B_ABST
    Figure CN119445181B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight multi-stage point cloud classification method based on a fused Transformer network, including: constructing an original point embedding module, local feature extraction, global feature aggregation, and multi-stage learning. The network of the present invention consists of an efficient feature aggregation branch and a Transformer branch. The efficient feature aggregation branch is equipped with learnable position encoding, which can fully capture the position information that is often ignored by most point cloud classification networks. The Transformer branch clusters the sampled points with similar features to achieve a balance between capturing long-range dependencies and computational complexity. At the same time, a squeeze-excitation module is introduced into the original point embedding to enhance the representation of point cloud features and improve the subsequent point cloud learning performance. A four-stage learning strategy is adopted to effectively enhance the network's ability to understand point clouds. A large number of experiments prove that the present invention can achieve excellent classification performance while maintaining a low number of parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of laser radar point cloud data processing, and in particular relates to a lightweight multi-stage point cloud classification method based on a fused Transformer network. Background Art

[0002] Point cloud classification is an important task in point cloud data processing. It has shown great potential in a variety of applications such as urban information modeling, autonomous driving, robotics, and 3D mapping. Unlike 2D images, 3D point clouds are characterized by disorder, irregularity, and sparsity, which poses challenges to the direct application of deep learning-based methods. To address these challenges, many point cloud classification architectures have been proposed in recent years and have achieved remarkable results in point cloud classification tasks. However, most methods unilaterally focus on local features or global features without integrating the two features. In contrast, Transformer, as a new point cloud local and global feature extraction architecture, has achieved excellent results in point cloud classification tasks. Various Point Cloud Transformer methods have been continuously introduced, achieving state-of-the-art results in point cloud classification. However, existing Transformer networks have challenges in terms of structural complexity and the number of parameters, and the exploration of point cloud position encoding is not deep enough. Summary of the invention

[0003] Purpose of the invention: The purpose of the present invention is to provide a lightweight multi-stage point cloud classification method based on a fused Transformer network. The lightweight multi-stage point cloud Transformer classification network (Lightweight Multi-Stage Point Cloud Classification Network Integratedwith Transformer, LMPT for short) of the present invention cleverly integrates efficient feature aggregation branches with Transformer branches, and can effectively extract and integrate local and global information. At the same time, a learnable position encoding is introduced to effectively integrate point cloud geometric information that is often overlooked by many point cloud Transformer architectures. In addition, the SE module contained in the original point embedding significantly enhances the feature extraction capability of the model. The network architecture of the present invention adopts a four-stage learning strategy to improve classification accuracy while ensuring a low number of parameters.

[0004] Technical solution: A lightweight multi-stage point cloud classification method based on a fusion Transformer network according to the present invention takes the original Lidar point cloud data as input and performs the following steps S1-S4. A high-dimensional representation of point features is obtained through the original point embedding module, and then local feature extraction and global feature aggregation are sequentially performed to obtain the residual results of each stage. Finally, multi-stage learning is performed on each residual result to obtain the final classification result output. The steps are as follows:

[0005] Step S1: Construct an original point embedding module to project the original point cloud into a higher dimension. The original point embedding module includes a linear layer and a squeeze-excitation (SE) module Squeeze-Excitation;

[0006] Step S2: Local feature extraction. The embedded high-dimensional point cloud features are passed through two branches to respectively extract their local information by using farthest point sampling (FPS) and k-nearest neighbor algorithm (kNN) as well as sorting and kNN algorithms, and downsampling of the number of point clouds is achieved;

[0007] Step S3: Global feature aggregation. The local features extracted by the two branches are respectively integrated and then passed into a learnable position encoding layer and a self-attention layer to achieve global feature aggregation. Then, the results of these two branches are integrated to obtain the residual results of each stage;

[0008] Step S4: Multi-stage learning. A four-stage learning strategy is adopted, and the residual result of each stage is used as the input of the next stage. The downsampling rate of each stage is 1 / 2. The results after four-stage learning are passed into a classifier to obtain the final classification result.

[0009] Further, the specific steps of step S1 are as follows:

[0010] Step S1.1: Project the point cloud information into a higher-dimensional channel. In order to minimize unnecessary parameter quantities, one-dimensional convolution operation is used, as shown in the following formula:

[0011]

[0012] where x n is the input channel, h k is the convolution kernel size, y n is the output channel, represents the summation of the convolution kernel indices from k = 0 to k = K - 1;

[0013] Step S1.2: After one-dimensional convolution, the SE module is adopted to solve the problem of low discrimination degree between channels, and the utilization of network for different channel information is enhanced by learning the importance weights of each channel. First, through the compression operation, the spatial information of each channel is compressed into a single value;

[0014]

[0015] Wherein, H and W are respectively the height and width of the feature map, z is the output of the result, and x ij is a specific value in the feature map; then, the importance of each channel is recalibrated through a two-layer fully connected neural network as follows:

[0016] s = σ(W 2 δ(W 1 z)) (3)

[0017] Wherein, W 1 and W 2 are the weight matrices of two fully connected layers, δ represents the ReLU non-linear layer, σ represents the sigmoid activation function, and finally, in order to obtain the output feature map the calibration vector s is multiplied element-wise with the original input feature map X on a channel-by-channel basis as follows:

[0018]

[0019] Wherein, × represents element-wise multiplication.

[0020] Furthermore, the specific steps of step S2 are as follows:

[0021] Step S2.1: Local feature extraction in the efficient feature aggregation branch. After passing through the initial point embedding layer, the FPS method is used to select a subset of local center points, so as to obtain the downsampling process of N point clouds, as shown in equation (5), and finally N / 2 point clouds are obtained;

[0022]

[0023] Wherein, F i and P i respectively represent the feature and coordinate of the i-th original point, F c and P c respectively represent the feature and coordinate of the c-th center point in the subset; subsequently, the kNN algorithm is used to group the points into k adjacent spaces to form k local 3D regions for local feature extraction, as follows:

[0024]

[0025] Wherein, N k represents the k-th neighbor space. The coordinates of N original points and the subset of local center points output by FPS are used as the input of kNN; through these two steps, local features and coordinates are obtained;

[0026] Step S2.2: Local feature extraction in the Transformer branch. Compared with the previous part, the distance between two point clouds is directly calculated to extract the position information of the point cloud in more detail;

[0027] P n =||xyz i -xyz j || 2 (7)

[0028] In the formula, P n is the distance between point i and point j, and xyz i and xyz j represent the coordinates of the i-th and j-th points respectively; subsequently, similar to Step S2.1, the calculated point cloud distance is used as the input for subsequent kNN analysis; first, the point cloud distances are sorted and the k nearest neighbors are selected, and then the coordinates of the corresponding point cloud are extracted according to the distance index, as shown below:

[0029]

[0030] knn_xyz k =xyz j |j∈knn_idx k (9)

[0031] In the formula, argsort k represents sorting the distances of all points and selecting the indices of the first k points as knn_idx k ; then, knn_idx k is used as an index to find the corresponding point coordinates, thereby obtaining knn_xyz k .

[0032] Furthermore, the specific steps of Step S3 are as follows:

[0033] Step S3.1: Global feature aggregation in the efficient feature aggregation branch. The function of this step is to aggregate the local features extracted in Step S2.1 and obtain the global information of the point cloud. First, the neighbor points and the central points from the FPS and kNN operations are concatenated to prevent the loss of global information, as follows:

[0034] F kc =Concat(f k ,F k ) (10)

[0035] In the formula, F kc effectively aggregates global and local information and serves as the input to the LPE layer;

[0036] Subsequently, learnable parameters and kNN point coordinates are introduced into the trigonometric positional encoding and weighted. Due to the characteristics of the sine function, the trigonometric function can not only describe the exact position in the embedding space but also implicitly capture the relative position details between two points in the three-dimensional environment. Therefore, starting from a parameter-free trigonometric positional encoding, a learnable positional encoding design is carried out as follows:

[0037]

[0038]

[0039] In the formula, pos represents the coordinate information of one of the x, y, or z axes of the point, which is embedded into the feature dimension of c / 3. i represents the dimension of the trigonometric function calculation. α and β control the amplitude and wavelength respectively. Then, they are concatenated along the last dimension as follows:

[0040] PE = Concat(PE (pos,2i) , PE (pos,2i+1) ) (13)

[0041] To further improve the learning effect of the point cloud positional encoding, learnable parameters are introduced to gradually enhance its ability to learn position information along with the deep learning of the network as follows:

[0042] LPE = PE × LP (14)

[0043] F w = (LPE + p i ) × LPE (15)

[0044] In the formula, LPE is the positional encoding of the learnable parameter, LP represents the learnable parameter, and the parameters of these positional encodings are initialized using a random function. As the network depth increases, these parameters will be updated by the optimizer. × represents the element-wise product. Equation (14) illustrates the specific calculation process of LPE, and the LPE result obtained from equation (14) is used as the input of equation (15); p i represents the coordinates of the i-th k-nearest neighbor, and F w is the output result of the learnable parameter positional encoding. Finally, the features of F w are further aggregated through a pooling operation to obtain the final output F x of the efficient feature extraction module at each stage as follows:

[0045] F x = Max{F w}+ Ave{F w} (16)

[0046] In the formula, Max{·} and Ave{·} represent max pooling and average pooling respectively, and the combination of max pooling and average pooling is used to achieve more accurate point cloud learning.

[0047] Step S3.2: Global feature aggregation in the Transformer branch, which mainly consists of a position encoding layer and a self-attention layer, to efficiently fuse the position information and global feature information in the Transformer branch;

[0048] Step S3.3: Integration of the learning results of the two branches. Adding the results learned by the two branches respectively can maximize the performance of the model at each stage, as shown in the following formula:

[0049] res = F y +F x (17)

[0050] In the formula, res represents the output of the residual result at each stage, F x is the final output of the efficient feature aggregation branch, while F y is the final output of the Transformer branch.

[0051] Furthermore, the specific steps of Step S3.2 are as follows:

[0052] Step S3.2.1: Position encoding layer. Since the 3D point cloud has coordinate position information, a multi-layer perceptron composed of two linear layers and one ReLU non-linear layer is used to learn the relative coordinates of the point cloud and serve as the position encoding of the Transformer branch. The position encoding strategy adopted here is different from that of the efficient feature aggregation branch, and by integrating their respective advantages, more comprehensive point cloud position information can be learned. Specifically as follows:

[0053] PE = W 2 δ(W 1 (pos i -pos j )) (18)

[0054] In the formula, pos i and pos j respectively represent the three-dimensional position coordinates of the i-th and j-th points, W 1 and W 2 respectively represent the weight matrices of the two linear layers, δ represents the ReLU non-linear layer, and PE is the calculated position encoding result;

[0055] Step S3.2.2: Self-attention layer. Vector self-attention is used to ensure the integrity of the learned point cloud information, specifically as follows:

[0056]

[0057] In the formula, Q, K, and V respectively represent the query vector, key vector, and value vector, d represents the dimension of the input vector. To enhance the global ability of the Transformer branch, a residual connection is made between the output of the self-attention layer and the original point cloud features, as shown in Equation (20):

[0058] F y = Wf y + F(20)

[0059] In the formula, f y represents the output of the self-attention layer, W represents the weight matrix, F represents the original features, and F y is the final output of the Transformer branch. Through this step, the Transformer branch can more effectively aggregate global features.

[0060] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method of the present invention.

[0061] The present invention also discloses a computer-readable storage medium, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, the steps of the method of the present invention are implemented.

[0062] The present invention also discloses a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the method of the present invention are implemented.

[0063] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages:

[0064] 1. The original point embedding module of the present invention effectively captures the features of the original points, further enhancing the stability and performance of the model and facilitating subsequent network processing.

[0065] 2. The efficient feature extraction branch with learnable position encoding in the present invention enhances the effect of global feature learning and ensures the learning ability of the model under low parameter conditions.

[0066] 3. The present invention adopts the Transformer branch to better capture global information and long-range dependencies, further avoiding the loss of point cloud information.

[0067] 4. The present invention adopts a four-stage learning strategy to make the learning of point cloud features more comprehensive, enhancing the performance of the model and the accuracy of the classification result.

[0068] 5. The model of the present invention achieves an excellent balance between lightweight design and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 is a schematic diagram of a lightweight multi-stage point cloud classification network integrating Transformer according to the present invention;

[0070] Figure 2 is a detailed schematic diagram of each stage in a lightweight multi-stage point cloud classification network integrating Transformer according to the present invention;

[0071] Figure 3 is a detailed schematic diagram of the original point embedding module according to the present invention;

[0072] Figure 4 is a detailed schematic diagram of the learnable position encoding in the efficient feature aggregation branch according to the present invention;

[0073] Figure 5 is a detailed schematic diagram of the Transformer according to the present invention;

[0074] Figure 6 is a visualization diagram of the dataset retrieval example provided by the example of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0075] The technical solutions of the present invention will be further described below with reference to the accompanying drawings.

[0076] A lightweight multi-stage point cloud classification network integrating Transformer provided by an embodiment of the present invention,

[0077] combines Figure 1 and Figure 2 As shown first, the original points are projected into a high dimension suitable for point cloud learning through the original point embedding module; then, through the efficient feature aggregation branch and the Transformer branch, local feature extraction and original feature aggregation are respectively performed, so as to enhance the effect of global feature learning while maintaining the low parameters and complexity of the model; finally, the residual results learned by the two branches are added and aggregated, and input into the final classifier to obtain the final result output.

[0078] In this embodiment, two Lidar point cloud datasets, ModelNet40 and ScanObjectNN, are used to verify the model. The ModelNet dataset can be obtained through the link (http: / / modelnet.cs.princeton.edu) on the Princeton University website and is used to evaluate the effectiveness of the network in 3D classification tasks. The ModelNet40 dataset contains 12,311 CAD models and is divided into 40 object categories. Specifically, the training set contains 9,843 models and the test set contains 2,468 models. ScanObjectNN is a real 3D point cloud classification dataset. Due to characteristics such as cluttered backgrounds, partial missing, and deformation, this dataset is extremely challenging, with a large number of outliers in the point cloud. The dataset consists of occluded objects extracted from real indoor scans, including 15,000 real objects from 15 categories, for a total of 2,902 3D objects. We train and test the classification task on the most challenging variant (PB T50 RS) of the dataset. The specific implementation of this embodiment adopts the same data preparation strategy as PointNet, including data sampling and data augmentation strategies.

[0079] A lightweight multi-stage point cloud classification network integrating Transformer according to the present invention is specifically as follows:

[0080] Step S1: Original point embedding module. A new type of original point embedding method is proposed, which consists of a linear layer and an SE module and can project the original point cloud into a higher dimension to facilitate the full learning of subsequent point cloud features.

[0081] The specific steps of step S1 are as follows:

[0082] Step S1.1: Project the point cloud information into channels of a higher dimension. To minimize unnecessary parameter quantities, a one-dimensional convolution operation is used, as shown in the following formula:

[0083]

[0084] In the formula, x n is the input channel, h k is the convolution kernel size, y n is the output channel, represents the summation of the convolution kernel indices from k = 0 to k = K - 1.

[0085] Step S1.2: After one-dimensional convolution, the distinguishability between channels is low. To solve this problem, the SE module is adopted to enhance the network's utilization of information in different channels by learning the importance weights of each channel. First, through a compression operation (global average pooling), as shown in Equation 2, the spatial information of each channel is compressed into a single value.

[0086]

[0087] Wherein, H and W are the height and width of the feature map respectively, z is the output of the result, and x ij is the specific value in the feature map. Then, the importance of each channel is recalibrated through a two-layer fully connected neural network as follows:

[0088] s = σ(W 2 δ(W 1 z)) (3)

[0089] Wherein, W 1 and W 2 are the weight matrices of two fully connected layers, δ represents the ReLU non-linear layer, and σ represents the sigmoid activation function. Finally, in order to obtain the output feature map the calibration vector s is multiplied element-wise with the original input feature map X on a channel-by-channel basis as follows:

[0090]

[0091] Wherein, × represents element-wise multiplication.

[0092] Step S2: Local feature extraction. The embedded high-dimensional point cloud features are respectively used in two branches to extract their respective local information by using the FPS and kNN algorithms as well as the sorting and kNN algorithms, and downsampling of the number of point clouds is achieved.

[0093] The specific steps of Step S2 are as follows:

[0094] Step S2.1: Local feature extraction in the efficient feature aggregation branch. After passing through the initial point embedding layer, the FPS method is used to select a subset of local center points, so as to obtain the downsampling process of N point clouds, as shown in Equation 5, and finally N / 2 point clouds are obtained.

[0095]

[0096] Wherein, F i and P i respectively represent the feature and coordinate of the i-th original point, and F c and P c respectively represent the feature and coordinate of the c-th center point in the subset. Subsequently, the kNN algorithm is used to group the points into k adjacent spaces to form k local 3D regions for local feature extraction as follows:

[0097]

[0098] Wherein, N kRepresents the k-th neighbor space. The coordinates of N original points and the subset of local center points output by FPS are used as the input of kNN. Through these two steps, local features and coordinates are obtained.

[0099] Step S2.2: Local feature extraction in the Transformer branch. Compared with the previous part, we directly calculate the distance between two point clouds to extract the position information of the point cloud in more detail.

[0100] P n =||xyz i -xyz j || 2 (7)

[0101] In the formula, P n is the distance between point i and point j, and xyz i and xyz j represent the coordinates of the i-th and j-th points respectively. Subsequently, similar to step S2.1, the calculated point cloud distance is used as the input for subsequent kNN analysis. First, sort the point cloud distances and select the k nearest neighbors, and then extract the coordinates of the corresponding point cloud according to the distance index, as shown below:

[0102]

[0103] knn_xyz k =xyz j |j∈knn_idx k (9)

[0104] In the formula, arg sort k means sorting the distances of all points and selecting the indices of the first k points as knn_idx k . Then, use knn_idx k as the index to find the corresponding point coordinates, so as to obtain knn_xyz k .

[0105] Step S3: Global feature aggregation. After integrating the local features extracted by the two branches respectively, they are input into the learnable position encoding layer and the self-attention layer to achieve global feature aggregation, and then the results of these two branches are integrated to obtain the residual results of each stage.

[0106] The specific steps of step S3 are as follows:

[0107] Step S3.1: Global feature aggregation in the efficient feature aggregation branch. The role of this step is to aggregate the local features extracted in Step S2.1 and obtain the global information of the point cloud. First, the neighbor points and the central points from the FPS and kNN operations are concatenated to prevent the loss of global information, as follows:

[0108] F kc = Concat(f k , F k ) (10)

[0109] In the formula, F kc effectively aggregates global and local information and serves as the input to the LPE layer.

[0110] Subsequently, learnable parameters and the kNN point coordinates are introduced into the trigonometric positional encoding and weighted. Due to the characteristics of the sine function, the trigonometric function can not only describe the exact position in the embedding space but also implicitly capture the relative position details between two points in the three-dimensional environment. Therefore, starting from a parameter-free trigonometric positional encoding, a learnable positional encoding is designed, as shown below:

[0111]

[0112]

[0113] In the formula, pos represents the coordinate information of one of the x, y, or z axes of the point, which is embedded into the feature dimension of c / 3. i represents the dimension of the trigonometric function calculation. α and β control the amplitude and wavelength respectively. These two trigonometric function formulas can effectively learn the relative position information of the point cloud. Then, they are concatenated along the last dimension, as shown below:

[0114] PE = Concat(PEx pos,2i) , PE (pos,2i+1) ) ((13)

[0115] To further improve the learning effect of the point cloud positional encoding, learnable parameters are introduced to gradually enhance its ability to learn the position information along with the deep learning of the network, as shown below:

[0116] LPE = PE × LP (14)

[0117] F w = (LPE + p i ) × LPE (15)

[0118] Wherein, LPE is the positional encoding of learnable parameters. LP represents learnable parameters, and the parameters of these positional encodings are initialized using a random function, and these parameters will be updated by the optimizer as the network depth increases. × represents the element-wise product. Equation 14 illustrates the specific calculation process of LPE, and the LPE result obtained from Equation 14 is used as the input of Equation 15. p i represents the coordinates of the i-th k-nearest neighbor. F w is the output result of the positional encoding of learnable parameters. Finally, we further aggregate the features of F w to obtain the final output F x of the efficient feature aggregation module at each stage, as follows:

[0119] F x = Max{F w}+ Ave{F w}} (16)

[0120] Wherein, Max{·} and Ave{·} represent max pooling and average pooling respectively. The combination of max pooling and average pooling is adopted to achieve more accurate point cloud learning.

[0121] Step S3.2: Global feature aggregation in the Transformer branch, which mainly consists of a positional encoding layer and a self-attention layer, to efficiently fuse the positional information and global feature information in the Transformer branch.

[0122] Step S3.2.1: Positional encoding layer. Since the 3D point cloud has coordinate position information, a multi-layer perceptron composed of two linear layers and one ReLU non-linear layer is used to learn the relative coordinates of the point cloud and then use them as the positional encoding of the Transformer branch, as follows:

[0123] PE = W 2 δ(W 1 (pos i - pos j )) (17)

[0124] Wherein, pos i and pos j represent the three-dimensional position coordinates of the i-th and j-th points respectively, W 1 and W 2 represent the weight matrices of the two linear layers respectively, δ represents the ReLU non-linear layer, and PE is the calculated positional encoding result. The positional encoding strategy adopted here is different from that of the efficient feature aggregation branch, and their respective advantages are integrated, so that more comprehensive point cloud position information can be learned.

[0125] Step S3.2.2: Self-attention layer, using vector self-attention to ensure the integrity of the learned point cloud information, specifically as follows:

[0126]

[0127] In the formula, Q, K, and V represent the query vector, key vector, and value vector respectively, and d represents the dimension of the input vector. To enhance the global ability of the Transformer branch, a residual connection is made between the output of the self-attention layer and the original point cloud features, as shown in Equation 20:

[0128] F y = Wf y + F(19)

[0129] In the formula, f y represents the output of the self-attention layer, W represents the weight matrix, F represents the original features, and F y is the final output of the Transformer branch. Through this step, the Transformer branch can more effectively aggregate global features.

[0130] Step S3.3: Integration of the learning results of the two branches. Adding the learning results of the two branches respectively can maximize the performance of the model at each stage, specifically as follows:

[0131] res = F y + F x (20)

[0132] In the formula, res represents the output of the residual result at each stage, F x is the final output of the efficient feature aggregation branch, and F y is the final output of the Transformer branch.

[0133] Step S4: Multi-stage learning, as shown in Figure 1 , adopt a 4-stage learning strategy, use the residual result of each stage as the input of the next stage, the downsampling rate of each stage is 1 / 2, and input the results after four-stage learning into the classifier to obtain the final classification result.

[0134] Table 1: Quantitative comparison of classification performance (mAcc, OA, parameters) on two datasets, distinguishing non-Transformer methods and Transformer-based methods, and highlighting the best performance in bold for each metric.

[0135]

[0136] Transformer-based methods

[0137]

[0138] It can be seen that the invention has only 1.44 MB of parameters (half of the number of PCT network parameters), indicating that the network of the invention has the lowest complexity among all Transformer-based methods, thus achieving the lightest architecture. On the ModelNet40 dataset, although the overall accuracy (OA) of the network of the invention is 93.8%, slightly lower than that of PointTransformerV2, its mean accuracy (mAcc) is 91.9%, which is the highest among all methods (0.3% higher than PointTransformerV2). It is worth noting that on the ScanObjectNN dataset, the overall accuracy (OA) of the network of the invention is 85.0%, and the mean accuracy (mAcc) is 86.4%, significantly superior to other Transformer-based methods, where the mAcc is 5.2% higher than that of TPNC and the OA is 4.8% higher than that of ACE-Transformer.

[0139] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. A lightweight multi-stage point cloud classification method based on a fused Transformer network, characterized in that: Taking the original Lidar point cloud data as input, execute the following steps S1 to S4, obtain the high-dimensional representation of the point features through the original point embedding module, then perform local feature extraction and global feature aggregation in sequence to obtain the residual results of each stage, and finally learn each residual result in multiple stages to obtain the final classification result output. The steps are as follows: Step S1: construct an original point embedding module to project the original point cloud to a higher dimension. The original point embedding module includes a linear layer and a squeeze-excitation SE module; Step S2: local feature extraction, the embedded high-dimensional point cloud features are respectively extracted from the local information by using the farthest point sampling FPS and k nearest neighbor algorithm kNN and sorting and kNN algorithm through two branches, and the number of point clouds is downsampled; Step S3: Global feature aggregation: After the local features extracted by the two branches are integrated accordingly, they are respectively passed to the learnable position encoding layer and the self-attention layer to achieve global feature aggregation. Then, the results of the two branches are integrated to obtain the residual results of each stage. Step S4: Multi-stage learning, adopting a four-stage learning strategy, taking the residual result of each stage as the input of the next stage, the downsampling rate of each stage is 1 / 2, and the results after the four-stage learning are passed into the classifier to obtain the final classification result.

2. According to claim 1, a lightweight multi-stage point cloud classification method based on a fused Transformer network is characterized in that: The specific steps of step S1 are as follows: Step S1.1: Project the point cloud information into a higher dimensional channel to minimize the amount of unnecessary parameters, as shown below: In the formula, x n is the input channel, h k is the convolution kernel size, y n For the output channel, represents the sum of the convolution kernel indices from k = 0 to k = K-1; Step S1.2: After the one-dimensional convolution, the SE module is used to solve the problem of low discrimination between channels. The importance weight of each channel is learned to enhance the network's utilization of different channel information. First, the spatial information of each channel is compressed into a single value through compression operation. In the formula, H and W are the height and width of the feature map, z is the output of the result, and x ij is the specific value in the feature map; then, a two-layer fully connected neural network is used to recalibrate the importance of each channel, as shown below: s=σ(W2δ(W1z)) (3) Where W1 and W2 are the weight matrices of the two fully connected layers, δ represents the ReLU nonlinear layer, and σ represents the sigmoid activation function. Finally, in order to obtain the output feature map The calibration vector s is multiplied element-wise with the original input feature map X on a channel-by-channel basis as follows: Here, × represents element-by-element multiplication.

3. According to claim 1, a lightweight multi-stage point cloud classification method based on a fused Transformer network is characterized in that: The specific steps of step S2 are as follows: Step S2.1: Local feature extraction in the efficient feature aggregation branch. After the initial point embedding layer, the FPS method is used to select a subset of local center points to obtain the downsampling process of N point clouds, as shown in equation (5), and finally N / 2 point clouds are obtained; In the formula, F i and P i Respectively represent the features and coordinates of the i-th original point, F c and P c They represent the features and coordinates of the cth center point in the subset respectively; then, the kNN algorithm is used to group the points into k neighboring spaces to form k local 3D regions for local feature extraction, as shown below: Where N k represents the kth neighbor space; the coordinates of the N original points and the local center point subset output by FPS are used as the input of kNN; through these two steps, local features and coordinates are obtained; Step S2.2: Local feature extraction in the Transformer branch. Compared with the previous part, the distance between two point clouds is directly calculated to extract the location information of the point clouds in more detail. P n =||xyz i -xyz j || 2 (7) Where P n is the distance between point i and point j, xyz i and xyz j Represent the coordinates of the i-th and j-th points respectively; then, similar to step S2.1, the calculated point cloud distance is used as the input of the subsequent kNN analysis; first, the point cloud distances are sorted and the k nearest neighbors are selected, and then the coordinates of the corresponding point cloud are extracted according to the distance index, as shown below: knn_xyz k =xyz j |j∈knn_idx k (9) In the formula, argsort k Indicates that all Sort the distances of the points and select the indexes of the first k points as knn_idx k ; Then, use knn_idx k As an index to find the corresponding point coordinates, so as to obtain knn_xyz k .

4. According to claim 1, a lightweight multi-stage point cloud classification method based on a fused Transformer network is characterized in that: The specific steps of step S3 are as follows: Step S3.1: Global feature aggregation in the efficient feature aggregation branch. The function of this step is to aggregate the local features extracted in step S2.1 and obtain the global information of the point cloud. First, the neighbor points and the center point from the FPS and kNN operations are connected to prevent the loss of global information, as follows: F kc =Concat(f k ,F k ) (10) In the formula, F kc It effectively aggregates global and local information and serves as the input of the LPE layer; Subsequently, learnable parameters and kNN point coordinates are introduced into the trigonometric position encoding and weighted, starting from a parameter-free trigonometric position encoding, and the learning position encoding design is performed as follows: In the formula, pos represents the coordinate information of one of the x, y or z axes of the point, which is embedded into the feature dimension of c / 3, i represents the dimension of trigonometric function calculation, α and β control the amplitude and wavelength respectively, and then it is spliced ​​along the last dimension, as shown below: PE=Concat(PE (pos,2i) ,ON (pos,2i+1) ) (13) The introduction of learnable parameters improves the learning effect of point cloud position encoding and gradually enhances its ability to learn position information as the network deepens, as shown below: LPE=PE×LP (14) F w =(LPE+p i )×LPE (15) Where LPE is the position encoding of the learnable parameters, LP represents the learnable parameters, and the parameters of these position encodings are initialized using a random function. As the network depth increases, these parameters will be updated by the optimizer. × represents the product of the elements. Equation (14) describes the specific calculation process of LPE. The LPE result obtained from equation (14) is used as the input of equation (15); p i represents the coordinates of the i-th k-nearest neighbor, F w is the output of the learnable parameter position encoding. Finally, the pooling operation further aggregates F w features to obtain the final output F of the efficient feature extraction module at each stage x , as shown below: F x =Max{F w }+Ave{F w } (16) In the formula, Max{·} and Ave{·} represent maximum pooling and average pooling respectively. The combination of maximum pooling and average pooling is used to achieve more accurate point cloud learning; Step S3.2: Global feature aggregation in the Transformer branch, which mainly consists of two parts: the position encoding layer and the self-attention layer, to efficiently fuse the position information and global feature information in the Transformer branch; Step S3.3: Integration of the learning results of the two branches. Adding the results learned by the two branches can maximize the performance of the model at each stage, as shown in the following formula: res=F y +F x (17) In the formula, res represents the output of the residual result at each stage, F x is the final output of the efficient feature aggregation branch, and F y It is the final output of the Transformer branch.

5. According to claim 4, a lightweight multi-stage point cloud classification method based on a fused Transformer network is characterized in that: The specific steps of step S3.2 are as follows: Step S3.2.1: Position encoding layer. Since the 3D point cloud has coordinate position information, a multilayer perceptron consisting of two linear layers and one ReLU nonlinear layer is used to learn the relative coordinates of the point cloud and then use it as the position encoding of the Transformer branch, as shown below: PE=W2δ(W1(pos i Possibly j )) (18) In the formula, pos i and pos j Represent the three-dimensional position coordinates of the i-th and j-th points respectively, W1 and W2 represent the weight matrices of the two linear layers respectively, δ represents the ReLU nonlinear layer, and PE is the calculated position encoding result; Step S3.2.2: Self-attention layer, using vector self-attention to ensure the integrity of the learned point cloud information, as follows: In the formula, Q, K and V represent the query vector, key vector and value vector respectively, d represents the dimension of the input vector. In order to enhance the global capability of the Transformer branch, a residual connection is made between the output of the self-attention layer and the original point cloud features, as shown in Equation (20): F y =Wf y +F (20) In the formula, f y represents the output of the self-attention layer, W represents the weight matrix, F represents the original feature, and F y It is the final output of the Transformer branch.

6. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method of claim 1.

7. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to claim 1 are implemented.

8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to claim 1 are implemented.

Citation Information

Patent Citations

  • Automatic assembly method based on cross-source point cloud and multi-modal information

    CN117523206A

  • Techniques for modifying and training a neural network

    US20210374518A1

Cited By

  • Universal efficient low-parameter fine tuning method for point cloud analysis

    CN121937834A