Lightweight point cloud feature extraction method based on deep learning network
Through a lightweight point cloud feature extraction method, using one-dimensional convolution and lightweight feature extraction modules, combined with Polarization Pooling, the problem of high computational and storage overhead in point cloud processing is solved, and efficient and accurate point cloud classification and segmentation is achieved, which is suitable for resource-constrained environments.
Patent Information
- Application Number
- CN202510650822.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-05
AI Technical Summary
Existing deep learning methods have too high computational and storage overhead in point cloud processing, making it difficult to achieve real-time processing on embedded and mobile devices. In addition, the performance of existing local feature extractors has reached a bottleneck, resulting in large computational workload and slow inference speed.
A lightweight point cloud feature extraction method based on deep learning network is adopted. Through shared one-dimensional convolution and lightweight feature extraction modules such as PPConv and PPConv Seg, combined with the Polarization Pooling method, local and global features of the point cloud are extracted, which reduces the computational complexity and improves the network generalization ability.
It significantly improves the processing efficiency and accuracy of point cloud classification and segmentation tasks, reduces computing and storage requirements, and is suitable for resource-constrained environments and for application scenarios with high real-time requirements such as autonomous driving and robot perception.
Smart Images

Figure CN120599282A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning and three-dimensional object surface feature extraction, and in particular to a lightweight point cloud feature extraction method based on a deep learning network. Background Art
[0002] With the development of 3D sensing technology, point cloud data, as an important representation of the morphology and structure of objects in 3D space, has been widely used in fields such as autonomous driving, robotic perception, and spatial modeling. Point cloud data has the following characteristics in practical applications: First, its sparsity and disorder make it complex to process; second, point cloud data is typically high-dimensional and large in volume, requiring significant computing resources and storage space to process; and finally, noise and missing data in point clouds can affect data quality and accuracy.
[0003] In recent years, deep learning technology has been gradually introduced into the field of point cloud processing. Point cloud feature extraction methods based on deep learning significantly improve the performance of tasks such as point cloud classification, segmentation, and object recognition by automatically learning high-level features in point cloud data. However, existing deep learning methods usually face the problem of excessive computational and storage overhead. The number of layers and parameters of deep neural networks has increased significantly, resulting in excessive computational complexity, which makes real-time processing difficult to achieve, especially in resource-constrained environments such as embedded and mobile devices. In addition, existing point cloud processing methods based on deep learning mostly rely on complex local feature extractors, and their performance has reached a certain bottleneck, but it also causes problems such as large computational load and slow inference speed. Therefore, how to design a method that can fully extract the spatial features of point clouds while ensuring high feature extraction accuracy while ensuring low computational and storage requirements has become an important research direction in the current field of point cloud processing. Summary of the Invention
[0004] The purpose of the present invention is to overcome the shortcomings and deficiencies of the existing technology and propose a lightweight point cloud feature extraction method based on a deep learning network. The surface features of the point cloud are extracted by shared one-dimensional convolution, and the global features of the point cloud are extracted by designing an ultra-lightweight feature extraction module, thereby achieving effective point cloud classification and segmentation accuracy.
[0005] To achieve the above-mentioned purpose, the technical solution provided by the present invention is: a lightweight point cloud feature extraction method based on a deep learning network, wherein the deep learning network is composed of a classification network and a segmentation network, the classification network is composed of a one-dimensional convolution, a lightweight classification feature extraction module and a global feature extraction module, and the segmentation network is composed of a lightweight segmentation feature extraction module, a lightweight classification feature extraction module and a global feature extraction module, wherein the one-dimensional convolution is for extracting point cloud features, the lightweight classification feature extraction module is called PPConv, which is designed based on low-rank theory and is used to extract local features of point clouds, the lightweight segmentation feature extraction module is called PPConv Seg, which is used to obtain local features of point clouds, and the global feature extraction module adopts the Polarization Pooling method, which extracts high-frequency and low-frequency features of point clouds, highlights high-frequency features, and can better extract global features of point clouds;
[0006] The specific implementation of the lightweight point cloud feature extraction method includes the following steps:
[0007] 1) Acquire point cloud data and preprocess it to obtain multi-scale point cloud data, and then divide the multi-scale point cloud data into training set and test set;
[0008] 2) The data in the training set is fed into a deep learning network for training. During the training process, the classification network uses one-dimensional convolution and PPConv to extract local features of the point cloud, then uses the global feature extraction module to extract global features of the point cloud, and finally uses the global features for classification; the segmentation network uses PPConv Seg and PPConv to extract local features of the point cloud, then uses the global feature extraction module to extract global features of the point cloud, and finally uses the extracted local features and global features to complete the segmentation task; wherein, when performing the classification and segmentation tasks, cross entropy is used to calculate the loss between the network prediction result and the label, and label smoothing is also performed to form soft labels to enhance the generalization ability of the network;
[0009] 3) Input the data in the test set into the trained deep learning network to obtain the prediction results, and then use the prediction results to obtain the classification and segmentation results of the point cloud data, and visualize the segmentation results to complete the point cloud classification and segmentation task.
[0010] Further, the step 1) includes the following steps:
[0011] 1.1) Image Preprocessing: Data enhancement is performed on the point cloud data, including random rotation, farthest point sampling, and random scaling. Random rotation and random scaling are performed to increase the richness of samples and enhance network generalization. Farthest point sampling is performed to preserve the structural information of the point cloud, which helps to better complete downstream tasks. The enhanced point cloud data is represented as follows:
[0012] P={p i |i=1,2,3,...,N}
[0013] Where N represents the number of sampled point clouds, and pi represents the characteristics of the i-th point cloud;
[0014] 1.2) Dataset division: For subsequent network training, the preprocessed point cloud data is divided into training set and test set according to the proportion.
[0015] Furthermore, in step 2), in the classification task, one-dimensional convolution and PPConv are used to extract local features of the point cloud data, and the Polarization Pooling method is used to extract global features, which are then used for classification. The specific processing flow is as follows:
[0016] First, feature extraction is performed on each point using shared one-dimensional convolution with different convolution kernels:
[0017]
[0018] Where, Represents the feature extraction process, MLP(p i ,K=j) means that the one-dimensional convolution layer K with the convolution kernel size j is used to extract features from the i-th point cloud;
[0019] Secondly, the features extracted from linear layers with different kernel sizes are fused:
[0020]
[0021] Where y i represents the fusion feature, f k1 (·),f k3 (·),f k5 (·) represents the use of one-dimensional convolution layers with convolution kernel sizes of 1, 3, and 5 to extract features from point clouds. Here, the point cloud features are linearly changed twice with different convolution kernel sizes. This is based on the low-rank theory to compress and reconstruct the point cloud features, which helps to reduce parameters and extract effective features. At the same time, different convolution kernel sizes can also improve the receptive field and help extract multi-scale information. Indicates splicing of features;
[0022] Next, the fusion feature y i After the linear layer processing, the skip residual connection is used:
[0023] y f =MLP(y i )+p i
[0024] Where y f It is the local feature of the point cloud extracted by PPConv. MLP represents the linear layer. The residual connection here makes the network more generalizable and can build a deeper neural network without overfitting.
[0025] Finally, the Polarization Pooling method is used to extract global features, and the classification task is performed with the help of the global features. The Polarization Pooling method is expressed as:
[0026] PPool(y f )=(V max -V mean )·V max
[0027] Where PPool(·) represents the Polarization Pooling method, where V max 、V mean Represents the high-frequency and low-frequency features extracted by max-pooling and mean-pooling. This method is to highlight the high-frequency features and extract the structural features of the point cloud.
[0028] Furthermore, in step 2), in the segmentation task, PPConv Seg and PPConv are used to extract local features of the point cloud data. Compared with PPConv, PPConv Seg adds a step of extracting global features to fuse the global features with the local features of the point cloud, thereby better completing the segmentation task. The specific processing flow is as follows:
[0029] First, input the point cloud into PPConv to obtain the feature y f ;
[0030] Then, the global features are extracted by using the max-pooling method, and then the global features are combined with the feature y f Perform splicing;
[0031] Finally, with the help of shared one-dimensional convolution, the fused features are obtained;
[0032] After using PPConv Seg to extract local features, PPConv is used to further refine the local features. Finally, the Polarization Pooling method is used to extract global features. The extracted local features are combined with the global features to complete the segmentation task.
[0033] Furthermore, in the deep learning network, for the network prediction results, the cross entropy with label smoothing is used to calculate the loss, which is expressed as:
[0034]
[0035] In the formula, L represents the loss value, C represents the number of categories, target class represents the target label, and y' i′ represents the target label after smoothing, h i′ Represents the network's predicted probability for category i', and ε represents the label smoothing factor; through smoothing, the network's generalization ability is enhanced.
[0036] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0037] 1. This invention significantly improves the processing efficiency and accuracy of point cloud classification and segmentation tasks. By adopting innovative lightweight feature extraction modules (such as PPConv and PPConv Seg), not only does it effectively reduce computational complexity, but also, through precise feature extraction, it greatly improves the adaptability and generalization capabilities of the network in different scenarios. Compared with traditional deep learning methods, this invention significantly reduces the computational overhead of the network while maintaining high accuracy, making it more suitable for resource-constrained environments such as embedded systems and mobile devices.
[0038] 2. This invention optimizes the network structure by designing an ultra-lightweight feature extraction module, significantly reducing the network's computational and storage requirements. Compared to existing methods involving computationally intensive deep neural networks, the network structure of this invention has a lower parameter count and enables efficient real-time processing on embedded platforms and mobile devices. It is particularly suitable for applications with high real-time requirements, such as autonomous driving and robotic perception. The efficiency of this method makes point cloud data processing more convenient and significantly improves the response speed in practical applications.
[0039] 3. The method of the present invention not only excels in classification and segmentation tasks but also possesses strong versatility and scalability. Through its lightweight design, the method can be widely applied in a variety of fields, including autonomous driving, robotic perception, 3D modeling, and environmental monitoring, and can be adapted and expanded according to actual needs. Whether it is for complex 3D surface feature extraction or simple scene reconstruction, the present invention provides an efficient and accurate solution and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 Flowchart of the method of the present invention.
[0041] Figure 2 It is a flowchart of the deep learning network. In the figure, PPConv SegConv1d represents one-dimensional convolution, MLP represents linear layer, Input represents point cloud input, Output Scores represents classification network output probability, Output represents segmentation network segmentation result, Global Feature represents global feature of point cloud, Shared represents shared convolution layer, Classification represents classification network, and Segmentation represents segmentation network.
[0042] Figure 3 This is the structural diagram of PPConv; in the figure, Plus represents the feature addition operation, Concat represents the feature splicing operation, Conv represents one-dimensional convolution, CR represents two convolutional layers for point cloud feature compression and reconstruction, and CompressionResconstruction Module represents the feature compression and reconstruction module.
[0043] Figure 4 This is the structural diagram of PPConv Seg; in the figure, Global Feature represents the global feature of the point cloud, Max-pooling represents maximum pooling, and Repeat represents feature repetition.
[0044] Figure 5 This is a flow chart of the Polarization Pooling method. In the figure, Max Pool represents maximum pooling, MeanPool represents average pooling, Minus represents feature subtraction, and Times represents feature multiplication. DETAILED DESCRIPTION
[0045] The present invention will be further described below with reference to specific examples, but the embodiments of the present invention are not limited thereto.
[0046] like Figures 1 to 5As shown, this embodiment discloses a lightweight point cloud feature extraction method based on a deep learning network, wherein the deep learning network is composed of a classification network and a segmentation network, wherein the classification network is composed of a one-dimensional convolution, a lightweight classification feature extraction module and a global feature extraction module, and the segmentation network is composed of a lightweight segmentation feature extraction module, a lightweight classification feature extraction module and a global feature extraction module, wherein the one-dimensional convolution is used to extract point cloud features, the lightweight classification feature extraction module is called PPConv, which is designed based on low-rank theory and is used to extract local features of point clouds, the lightweight segmentation feature extraction module is called PPConv Seg, which is used to obtain local features of point clouds, and the global feature extraction module adopts the Polarization Pooling method, which can better extract the global features of point clouds by extracting high-frequency and low-frequency features of point clouds and highlighting high-frequency features;
[0047] The specific implementation of the lightweight point cloud feature extraction method includes the following steps:
[0048] 1) Obtain point cloud data and preprocess it to obtain multi-scale point cloud data. Then divide the multi-scale point cloud data into training set and test set, including the following steps:
[0049] 1.1) Image Preprocessing: Data enhancement is performed on the point cloud data, including random rotation, farthest point sampling, and random scaling. Random rotation and random scaling are performed to increase the richness of samples and enhance network generalization. Farthest point sampling is performed to preserve the structural information of the point cloud, which helps to better complete downstream tasks. The enhanced point cloud data is represented as follows:
[0050] P={p i |i=1,2,3,...,N}
[0051] Where N represents the number of sampled point clouds, p i Represents the features of the i-th point cloud;
[0052] 1.2) Dataset Division: For subsequent network training, the preprocessed point cloud data is divided into training set and test set in a ratio of 7:3.
[0053] 2) The data in the training set is fed into a deep learning network for training. During the training process, the classification network uses one-dimensional convolution and PPConv to extract local features of the point cloud, then uses the global feature extraction module to extract global features of the point cloud, and finally uses the global features for classification; the segmentation network uses PPConv Seg and PPConv to extract local features of the point cloud, then uses the global feature extraction module to extract global features of the point cloud, and finally uses the extracted local features and global features to complete the segmentation task; wherein, when performing the classification and segmentation tasks, cross entropy is used to calculate the loss between the network prediction result and the label, and label smoothing is also performed to form soft labels to enhance the generalization ability of the network;
[0054] In the classification task, one-dimensional convolution and two PPConvs are used to extract local features on the point cloud data, increasing the feature dimension of the point cloud to 128. Polarization Pooling is then used to extract global features, which are then used for classification. The specific processing flow is as follows:
[0055] First, feature extraction is performed on each point using shared one-dimensional convolution with different convolution kernels:
[0056]
[0057] Where, Represents the feature extraction process, MLP(p i ,K=j) means that the one-dimensional convolution layer K with the convolution kernel size j is used to extract features from the i-th point cloud;
[0058] Secondly, the features extracted from linear layers with different kernel sizes are fused:
[0059]
[0060] Where y i represents the fusion feature, f k1 (·),f k3 (·),f k5 (·) represents the use of one-dimensional convolution layers with convolution kernel sizes of 1, 3, and 5 to extract features from point clouds. Here, the point cloud features are linearly changed twice with different convolution kernel sizes. This is based on the low-rank theory to compress and reconstruct the point cloud features, which helps to reduce parameters and extract effective features. At the same time, different convolution kernel sizes can also improve the receptive field and help extract multi-scale information. Indicates splicing of features;
[0061] Next, the fusion feature y i After the linear layer processing, the skip residual connection is used:
[0062] y f =MLP(y i )+p i
[0063] Where y f It is the local feature of the point cloud extracted by PPConv. MLP represents the linear layer. The residual connection here makes the network more generalizable and can build a deeper neural network without overfitting.
[0064] Finally, the Polarization Pooling method is used to extract global features, and the classification task is performed with the help of the global features. The Polarization Pooling method is expressed as:
[0065] PPool(y f )=(V max -V mean )·V max
[0066] Where PPool(·) represents the Polarization Pooling method, where V max 、V mean Represents the high-frequency and low-frequency features extracted by max-pooling and mean-pooling. This method is to highlight the high-frequency features and extract the structural features of the point cloud;
[0067] In the segmentation task, two PPConv Seg and one PPConv are used to extract local features of point cloud data. Compared with PPConv, PPConv Seg adds a step of extracting global features to fuse global features with local features of point clouds, thereby better completing the segmentation task. The specific processing flow is as follows:
[0068] First, input the point cloud into PPConv to obtain the feature y f ;
[0069] Then use the max-pooling method to extract the global features, and then combine the global features with the feature y f Perform splicing;
[0070] Finally, with the help of shared one-dimensional convolution, the fused features are obtained;
[0071] After using PPConv Seg to extract local features, PPConv is used to further refine the local features. Finally, the Polarization Pooling method is used to extract global features. The extracted local features are combined with the global features and spliced to complete the segmentation task.
[0072] In deep learning networks, for the results of network predictions, the loss is calculated using cross entropy with label smoothing, which is expressed as:
[0073]
[0074]
[0075] In the formula, L represents the loss value, C represents the number of categories, target class represents the target label, and y' i′ represents the target label after smoothing, h i′ Represents the network's predicted probability for category i', and ε represents the label smoothing factor; through smoothing, the network's generalization ability is enhanced.
[0076] Loss function and optimization strategy: Cross entropy loss was used, and a label smoothing strategy was implemented. The SGD optimizer was adopted, the learning rate was set to 0.1, and the entire network was trained through 300 rounds of training.
[0077] 3) Input the data in the test set into the trained deep learning network to obtain the prediction results, and then use the prediction results to obtain the classification and segmentation results of the point cloud data. Different segmentation categories can be divided into different colors, and the segmentation results can be visualized to complete the point cloud classification and segmentation task.
[0078] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. A lightweight point cloud feature extraction method based on deep learning network, characterized by: The deep learning network is composed of a classification network and a segmentation network. The classification network is composed of a one-dimensional convolution, a lightweight classification feature extraction module and a global feature extraction module. The segmentation network is composed of a lightweight segmentation feature extraction module, a lightweight classification feature extraction module and a global feature extraction module. Among them, the one-dimensional convolution is used to extract point cloud features. The lightweight classification feature extraction module is called PPConv, which is designed based on low-rank theory and is used to extract local features of point clouds. The lightweight segmentation feature extraction module is called PPConv Seg, which is used to obtain local features of point clouds. The global feature extraction module adopts the Polarization Pooling method, which extracts high-frequency and low-frequency features of point clouds, highlights high-frequency features, and can better extract global features of point clouds. The specific implementation of the lightweight point cloud feature extraction method includes the following steps: 1) Acquire point cloud data and preprocess it to obtain multi-scale point cloud data, and then divide the multi-scale point cloud data into training set and test set; 2) The data in the training set is fed into a deep learning network for training. During the training process, the classification network uses one-dimensional convolution and PPConv to extract local features of the point cloud, then uses the global feature extraction module to extract global features of the point cloud, and finally uses the global features for classification; the segmentation network uses PPConv Seg and PPConv to extract local features of the point cloud, then uses the global feature extraction module to extract global features of the point cloud, and finally uses the extracted local features and global features to complete the segmentation task; wherein, when performing the classification and segmentation tasks, cross entropy is used to calculate the loss between the network prediction result and the label, and label smoothing is also performed to form soft labels to enhance the generalization ability of the network; 3) Input the data in the test set into the trained deep learning network to obtain the prediction results, and then use the prediction results to obtain the classification and segmentation results of the point cloud data, and visualize the segmentation results to complete the point cloud classification and segmentation task.
2. The lightweight point cloud feature extraction method based on deep learning network according to claim 1 is characterized in that: Step 1) The following steps are involved: 1.1) Image Preprocessing: Data enhancement is performed on the point cloud data, including random rotation, farthest point sampling, and random scaling. Random rotation and random scaling are performed to increase the richness of samples and enhance network generalization. Farthest point sampling is performed to preserve the structural information of the point cloud, which helps to better complete downstream tasks. The enhanced point cloud data is represented as follows: P={p i |i=1,2,3,...,N} Where N represents the number of sampled point clouds, p i Represents the features of the i-th point cloud; 1.2) Dataset division: For subsequent network training, the preprocessed point cloud data is divided into training set and test set according to the proportion.
3. The lightweight point cloud feature extraction method based on deep learning network according to claim 2 is characterized in that: In step 2), in the classification task, one-dimensional convolution and PPConv are used to extract local features of the point cloud data, and the Polarization Pooling method is used to extract global features, which are then used for classification. The specific processing flow is as follows: First, feature extraction is performed on each point using shared one-dimensional convolution with different convolution kernels: Where, Represents the feature extraction process, MLP(p i ,K=j) means that the one-dimensional convolution layer K with the convolution kernel size j is used to extract features from the i-th point cloud; Secondly, the features extracted from linear layers with different kernel sizes are fused: Where y i represents the fusion feature, f k1 (·),f k3 (·),f k5 (·) represents the use of one-dimensional convolution layers with convolution kernel sizes of 1, 3, and 5 to extract features from point clouds. Here, the point cloud features are linearly changed twice with different convolution kernel sizes. This is based on the low-rank theory to compress and reconstruct the point cloud features, which helps to reduce parameters and extract effective features. At the same time, different convolution kernel sizes can also improve the receptive field and help extract multi-scale information. Indicates splicing of features; Next, the fusion feature y i After the linear layer processing, the skip residual connection is used: and f =MLP(and i )+p i Where y f It is the local feature of the point cloud extracted by PPConv. MLP represents the linear layer. The residual connection here makes the network more generalizable and can build a deeper neural network without overfitting. Finally, the Polarization Pooling method is used to extract global features, and the classification task is performed with the help of the global features. The Polarization Pooling method is expressed as: PPool(y f )=(V max -V mean )·V max Where PPool(·) represents the Polarization Pooling method, where V max 、V mean Represents the high-frequency and low-frequency features extracted by max-pooling and mean-pooling. This method is to highlight the high-frequency features and extract the structural features of the point cloud.
4. The lightweight point cloud feature extraction method based on deep learning network according to claim 3 is characterized in that: In step 2), in the segmentation task, PPConv Seg and PPConv are used to extract local features of the point cloud data. Compared with PPConv, PPConv Seg adds a step of extracting global features to fuse the global features with the local features of the point cloud, thereby better completing the segmentation task. The specific processing flow is as follows: First, input the point cloud into PPConv to obtain the feature y f ; Then, the global features are extracted by using the max-pooling method, and then the global features are combined with the feature y f Perform splicing; Finally, with the help of shared one-dimensional convolution, the fused features are obtained; After using PPConv Seg to extract local features, PPConv is used to further refine the local features. Finally, the Polarization Pooling method is used to extract global features. The extracted local features are combined with the global features to complete the segmentation task.
5. The lightweight point cloud feature extraction method based on deep learning network according to claim 4 is characterized in that: In deep learning networks, for the results of network predictions, the loss is calculated using cross entropy with label smoothing, which is expressed as: In the formula, L represents the loss value, C represents the number of categories, target class represents the target label, and y' i′ represents the target label after smoothing, h i′ Represents the network's predicted probability for category i', and ε represents the label smoothing factor; through smoothing, the network's generalization ability is enhanced.