Three-dimensional point cloud classification method based on local space structure sensing module

By introducing the Local Spatial Structure Perception Module (LSSPM) into the PointNet++ network, the accuracy and robustness issues of existing 3D point cloud classification methods on complex data and sparse point clouds are solved, achieving more efficient feature extraction and classification accuracy.

CN120747597APending Publication Date: 2025-10-03CHANGCHUN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510836299.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-21
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing 3D point cloud classification methods have poor generalization ability and low recognition accuracy, especially when performing poorly on complex data. In addition, traditional methods and deep learning networks such as PointNet++ suffer from spatial information loss or low computational efficiency when processing sparse and irregular point clouds.

Method used

A 3D point cloud classification method based on the Local Spatial Structure Perception Module (LSSPM) is adopted. By introducing a differentiable weight generator into the PointNet++ network, spatial weight coefficients are generated based on differential geometric characteristics, thereby enhancing the feature extraction capability of local geometric structures. Key feature information is retained through multi-scale region partitioning and progressive downsampling technology.

Benefits of technology

The accuracy and robustness of three-dimensional point cloud classification are significantly improved, especially in sparse point clouds and complex scenes, showing good accuracy and stability, improving the average accuracy and overall accuracy, and enhancing the network's ability to learn spatial structural features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747597A_ABST
    Figure CN120747597A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional point cloud classification method and system based on a local space structure sensing module, and provides the local space structure sensing module (LSSPM) aiming at the defects of an existing Point Net + + network in learning point cloud space information, an LSSPM-Point Net + + classification network is constructed, the network takes Point Net + + as a framework, original point clouds are processed layer by layer through a four-stage LSSPM module, and the point cloud space information is classified. Spatial structure weight information is fused in the feature extraction dimension raising process, an LSSPM module combines spatial distribution features and input features, spatial distribution weights are generated through a multi-layer perceptron, dynamic embedding of local geometric constraints is achieved, and the spatial feature representation capacity is enhanced. The result shows that on a ModelNet40 data set, compared with a traditional network architecture, the method has the advantages that the classification accuracy is remarkably improved, the average accuracy (MA) is improved by 1.4%, the overall accuracy (OA) is improved by 1.9%, and good robustness is achieved for different point cloud density input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of three-dimensional point cloud processing technology, and in particular to a three-dimensional point cloud classification system and method based on a local spatial structure perception module. Background Art

[0002] With the development of 3D sensor technology, 3D point cloud data has been widely used in various fields. It contains rich, multidimensional information and can accurately describe the 3D shape of objects. Traditional point cloud classification algorithms rely on manually designed features, have poor generalization capabilities, and suffer from low recognition accuracy when faced with complex data. In recent years, deep learning has been used for point cloud classification, but due to the sparse and irregular nature of point clouds, CNNs are difficult to directly apply. Strategies such as projection and voxel methods suffer from spatial information loss or low computational efficiency. PointNet and its improved PointNet++ architecture also suffer from insufficient modeling of local geometric structures and limitations in capturing spatial geometric features, which affect point cloud classification performance.

[0003] Purpose

[0004] In order to achieve the above objectives, the present invention aims to solve the above problems existing in the existing three-dimensional point cloud classification methods, and provide a three-dimensional point cloud classification system and method based on a local spatial structure perception module to enhance the network's ability to learn the spatial structure features of three-dimensional point clouds and improve the accuracy and robustness of point cloud classification. Summary of the Invention

[0005] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:

[0006] A 3D point cloud classification method based on a local spatial structure perception module comprises the following steps:

[0007] S1. Build the network framework: A network based on PointNet++ is constructed, consisting of two parts: encoding and decoding. The encoding network converts the raw, unordered, sparse point cloud data into high-dimensional feature vectors containing rich semantic and geometric information. The decoding network uses the features extracted by the encoding network to perform classification predictions on the point cloud data. During construction, the PointNet++ hierarchical feature learning framework is utilized, integrating the Farthest Point Sampling (FPS) algorithm with a multi-scale region partitioning strategy to pave the way for efficient point cloud data processing.

[0008] S2, Design of the Local Spatial Structure Perception Module (LSSPM): In the encoding network, the original geometric feature abstraction layer is replaced with the SAM module. Leveraging a differentiable weight generator and based on the properties of differential geometry, this module adaptively generates spatial weight coefficients based on the local curvature tensor. This dynamically incorporates 3D spatial structural features into the feature extraction process, establishing an implicit association between point cloud data and the underlying geometric manifold. This significantly enhances the network's ability to represent spatial features and enables in-depth learning of local, fine-grained features of point clouds.

[0009] S3, encoding network processing: The raw point cloud data is processed layer by layer through four levels of LSSPM. Through progressive downsampling, the point cloud is compressed to 1 / 256 of its original size, while the feature channels are expanded to 512 dimensions. This allows each encoded point to carry multi-scale geometric details and deep semantic information, minimizing data volume while preserving key features to the greatest extent possible.

[0010] S4, Decoding and Classification: The decoding stage directs the high-dimensional representations aggregated by the encoding network to the classification prediction layer. The classification network consists of a PointNet layer, a fully connected FC layer, and a Softmax function. The PointNet layer integrates the high-dimensional representations to extract global features. The FC layer maps the features to the classification space to generate category scores. The Softmax function normalizes the scores to obtain probabilities, thereby determining the category to which the point cloud belongs.

[0011] Furthermore, in step 1, the encoding network adopts the PointNet++ hierarchical feature learning framework, which can adaptively select representative points for feature extraction through the farthest point sampling (FPS) algorithm and multi-scale region partitioning strategy, ensuring that spatial features can be effectively captured on point cloud data of different densities, providing a high-quality feature foundation for subsequent processing.

[0012] Furthermore, in step 2, the differentiable weight generator of the local spatial structure perception module (LSSPM) is based on differential geometry properties and can adaptively generate spatial weight coefficients through precise calculation of the local curvature tensor. This coefficient can be dynamically adjusted according to changes in the local geometric structure of the point cloud, enabling the network to focus more accurately on key feature areas and significantly enhance the ability to learn spatial structural features.

[0013] Furthermore, in step 3, the progressive downsampling process strictly controls the sampling ratio, compressing the number of point cloud points to 1 / 256 of the initial scale. At the same time, through a carefully designed feature channel expansion mechanism, the number of channels is expanded to 512 dimensions. While reducing data redundancy, it effectively retains multi-scale geometric details and deep semantic information, improving network processing efficiency and feature expression capabilities.

[0014] Furthermore, in step 4, the PointNet layer of the classification network uses a symmetric function to implement permutation invariant feature learning, which can effectively handle the disorder of point cloud data. The FC fully connected layer maps high-dimensional features to the classification space through a large number of neuron connections, realizing the conversion of features to category scores. The Softmax function normalizes the category scores and outputs the classification results in a probabilistic form to ensure the accuracy and reliability of the results.

[0015] Beneficial effects

[0016] Compared with the existing technology, the 3D point cloud classification system and method based on the local spatial structure perception module provided by the present invention has the following beneficial effects:

[0017] By designing a local spatial structure perception module (LSSPM), this paper dynamically integrates three-dimensional spatial structure features into the feature extraction process, significantly enhancing the network's ability to learn the spatial structure features of three-dimensional point clouds. In the ModelNet40 dataset, the average accuracy (MA) is improved by 1.4% and the overall accuracy (OA) is improved by 1.9%, providing more reliable technical support for three-dimensional point cloud classification tasks.

[0018] In the present invention, after processing with a four-level local spatial structure perception module and progressive downsampling, the method exhibits good robustness when facing inputs of different point cloud densities. When the number of input point clouds is as low as 64, the accuracy can still be maintained at 89%, and it can work stably in complex and changeable actual scenarios.

[0019] During the data processing and model training process, this paper adopts the ModelNet40 benchmark dataset, scientific experimental parameter settings, and multi-dimensional experimental comparative analysis, combined with ablation experiments to verify the effectiveness of the local spatial structure perception module. After the introduction of this module, MA increased by 3.8% and OA increased by 2.5%, providing important theoretical and practical basis for the development of 3D point cloud classification technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 : Presents the pipeline of the LSSPM-PointNet++ network from point cloud input to classification output, including four-level LSSPM processing and progressive downsampling.

[0021] Figure 2 : Demonstrates the process of LSSPM fusing spatial distribution and input features, including coordinate difference calculation, MLP weight generation and feature fusion.

[0022] Figure 3 : Comparing the accuracy of LSSPM-PointNet++ with other networks under different number of point clouds, showing its robustness under sparse point clouds. DETAILED DESCRIPTION

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0024] Example

[0025] The 3D point cloud classification method based on the local spatial structure perception module (LSSPM-PointNet++) proposed in this invention includes the following specific implementation steps:

[0026] S1: Multi-dimensional point cloud data acquisition and preprocessing

[0027] The ModelNet40 benchmark dataset was selected as the core data source. This dataset contains 40 object categories and 12,311 3D models, with a training set to test set ratio of 8:2. Each model consists of 2048 unordered points, and the original features are 3D coordinates (X, Y, Z). Data preprocessing is completed through the following steps:

[0028] Coordinate normalization: calculating the point cloud centroid

[0029]

[0030] The coordinates are translated to the center of mass and scaled to the unit sphere, the formula is:

[0031]

[0032] Ensure that the point cloud data is distributed within a unit sphere centered at the origin to eliminate the effects of scale and position differences.

[0033] Normal vector calculation: For each point's K = 50 neighborhood, principal component analysis (PCA) is used to calculate the covariance matrix. The eigenvector corresponding to the minimum eigenvalue is taken as the normal vector. After normalization, it is used as an additional feature to form a 6-dimensional input (X, Y, Z, NX, NY, NZ) to enhance the geometric structure representation capability.

[0034] Data augmentation: Random rotation (0-360° around the X / Y / Z axes), scaling (randomly selected with a scale factor of 0.8-1.2), and Gaussian noise (mean 0, standard deviation 0.01) are used to augment the training data to generate a 4x larger dataset and improve the model's generalization ability.

[0035] The data collection process strictly controls the consistency of input feature dimensions to ensure that each sample contains 2048 points and corresponding normal vectors, such as Figure 3As shown, it provides standardized input for subsequent model training.

[0036] S2: Local Spatial Structure Perception Module (LSSPM) network architecture construction

[0037] The LSSPM-PointNet++ network is built based on the PyTorch framework. The core architecture is as follows Figure 1 As shown, it includes two parts: encoding and decoding:

[0038] Encoding network (four-level hierarchical feature extraction):

[0039] Farthest Point Sampling (FPS): Downsample by 1 / 4 at each level, reducing the number of point clouds from 2048 to 512, 128, 32, and 8 points in sequence, to ensure that multi-scale features are extracted at different densities, such as Figure 3 shown.

[0040] Ball query: With the current sampling point as the center, the search radius is increased hierarchically (0.2, 0.4, 0.8, 1.6). K = 32 neighboring points are selected for each point to construct a local area to capture the contextual geometric relationship.

[0041] Local Spatial Structure Perception Module (LSSPM): Calculates the coordinate difference between neighboring points and the center point

[0042] Δp i =(x i -x c ,y i -y c ,z i -z c )

[0043] and the Euclidean distance d i =||Δp i ||, spliced ​​into 4-dimensional spatial features [Δp i ,d i After being processed by two layers of MLP (128 dimensions in the first layer and 1 dimension in the second layer), the spatial weights are generated by the Sigmoid function.

[0044] w i =σ(MLP([Δp i ,d i ]))

[0045] Multiply element-wise with the input features to achieve weighted fusion:

[0046] f fusion =f input ⊙w i

[0047] This operation enables the network to dynamically focus on key areas where curvature changes significantly, enhancing the expressiveness of local geometric features.

[0048] Decoding network (classification prediction):

[0049] The 8 sampling points output by the encoding network each carry 512-dimensional features, which are aggregated into global features through global maximum pooling, mapped to 1024 dimensions through the fully connected layer (FC), and then processed by Dropout (0.5) and Softmax functions to output 40 classification probabilities to complete the point cloud category prediction.

[0050] S3: Model training

[0051] Model training and optimization:

[0052] The Adam optimizer is used (initial learning rate 0.001, decay step size 100, minimum learning rate 1e-5), and the loss function is cross entropy loss combined with label smoothing (ε=0.1). The specific training process is as follows:

[0053] Parameter configuration:

[0054] The batch size is set to 24 to balance GPU memory usage and training stability.

[0055] The number of training rounds is 200, and the early stopping mechanism (termination if the validation set loss remains unchanged for 15 consecutive rounds) is combined to avoid overfitting.

[0056] The weight decay coefficient is set to 1e-4 to suppress the excessive growth of model parameters.

[0057] Tiered training:

[0058] Each LSSPM module gradually expands the feature channels (6 → 64 → 128 → 256 → 512), preserving the original features through residual connections to ensure efficient gradient propagation. During training, TensorBoard was used to monitor loss curves and accuracy changes in real time, with model checkpoints saved every 10 epochs. The optimal model achieved an overall accuracy (OA) of 93.2% on the validation set.

[0059] Distributed training:

[0060] By using data parallel technology, training acceleration is achieved on an 8-card GPU cluster, reducing the time for a single round of training from 120 seconds to 25 seconds, with an acceleration ratio of 4.8 times, ensuring efficient processing of large-scale data.

[0061] S4: Signal processing and feature fusion verification

[0062] Interference elimination and pulse compression (comparative experiment):

[0063] Compared with the voxel method (VoxNet) and the projection method (MVCNN), LSSPM-PointNet++ directly processes the original point cloud, avoiding the spatial information loss of voxelization and the depth information loss of the projection method. Figure 3 As shown in the figure, when the number of input point clouds is reduced to 64, the accuracy of the method of the present invention still reaches 89%, which is significantly improved compared with 75% of VoxNet and 80% of MVCNN, demonstrating its excellent processing ability for sparse point clouds.

[0064] The effectiveness of the LSSPM module was verified through ablation experiments: after removing LSSPM, the overall accuracy of the model dropped to 91.5%, which is consistent with the original PointNet++, while the accuracy of the complete model including LSSPM reached 93.4%, indicating that the SAM module contributed 1.9% accuracy improvement, verifying the necessity of dynamic fusion of spatial structure features.

[0065] Cross-modal validation (KITTI dataset):

[0066] KITTI point cloud data collected by a vehicle-mounted lidar was transformed and noise filtered before being fed into the LSSPM-PointNet++ model. The classification accuracy for cars, pedestrians, and cyclists reached 97.2%, 92.5%, and 95.3%, respectively. These are significant improvements over the 93.7%, 88.2%, and 91.1% achieved by traditional PointNet++, demonstrating the model's generalization capabilities in real-world scenarios.

[0067] S5: Gesture Recognition and Performance Testing

[0068] Classification decision process:

[0069] After the input point cloud is processed by the four-stage LSSPM module, the decoding network outputs the classification probabilities using the Softmax function. When the probability of a category exceeds 0.95, it is directly classified as that category. If the highest probability is less than 0.95, a secondary feature re-extraction is triggered, combining the spatial weights of neighboring point clouds to recalculate the classification results to ensure the reliability of the results.

[0070] Multi-scenario testing:

[0071] ModelNet40 benchmark test: LSSPM-PointNet++ achieved an overall accuracy (OA) of 93.4% and an average accuracy (MA) of 91.3% on the test set, which are 1.9% and 1.4% higher than PointNet++'s 91.5% and 89.9%, respectively. In particular, the accuracy exceeded 95% for complex geometric categories such as "cabinet" and "telephone".

[0072] Robustness test: When 5% Gaussian noise is added to the point cloud data, the overall accuracy of the proposed method remains at 92%, significantly better than PointCNN's 88.7% and DRNet's 89.5%, demonstrating its strong robustness to noise interference.

[0073] Industrial application verification: In industrial CT scanning parts classification, the classification accuracy rate for parts such as gears and bearings with 10% missing points reached 94.6%, meeting the high-precision requirements of automated inspection equipment.

[0074] Finally, it should be noted that

[0075] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein with equivalents. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A three-dimensional point cloud classification method based on local spatial structure perception, characterized in that: The following steps are involved: S1, data preprocessing: obtain 3D point cloud data and perform standardization processing. The 3D point cloud data contains 3D coordinates and normal vector information. A single data contains 2048 points. The 3D point cloud data is downsampled to obtain point cloud data of different densities for subsequent experiments. S2, building a network model: Using PointNet++ as the framework, the local spatial structure perception module (LSSPM) is introduced. The LSSPM is used to fuse spatial distribution features and input features to generate spatial distribution weights. The network model consists of two parts: encoding and decoding. The encoding network is composed of a feature extractor. In the decoding stage, the aggregated high-dimensional representation is directly directed to the classification prediction layer. S3, Feature Extraction: Point cloud data is processed layer by layer through a four-level LSSPM module. Each level of the module dynamically aggregates contextual features within the spatial neighborhood through the farthest point sampling (FPS) algorithm and the spherical neighborhood search mechanism, performs progressive downsampling, compresses the number of points to 1 / 256 of the initial scale, and simultaneously completes the expansion of the feature channel to 512 dimensions, so that each encoded point carries multi-scale geometric details and deep semantic information. S4, model training: The network model was trained on the ModelNet40 dataset, which contains 40 categories and 12311 data models. The ratio of training set to test set was 8:

2. The parameters were optimized. The optimizer was Adam, the learning rate was 0.001, the decay step size was 0.0001, the momentum was 0.1, the batch size was 24, 200 epochs were trained, and 5-fold cross validation was used to evaluate the model performance. S5, classification prediction: The point cloud to be classified is input into the trained model, and the classification score is obtained through the PointNet layer, the fully connected layer and the Softmax function to obtain the classification result. The evaluation indicators of the classification result include the average accuracy MA and the overall accuracy OA.

2. The classification method according to claim 1, characterized in that The input-output relationship of the local spatial structure perception module (LSSPM) is expressed as: y = f(x, w), where x is the input feature, w is the spatial distribution weight, and y is the output feature. The LSSPM module combines the spatial distribution feature and the input feature so that the weight of each point is not only related to the vector from itself to the center point, but also to the weights of all points in the local area.

3. The classification method according to claim 1, characterized in that The method for generating the spatial distribution weight is as follows: using a multi-layer perceptron (MLP) to adjust the dimension of the spatial distribution feature to match the dimension of the input feature, and applying a Sigmoid function to generate the spatial distribution weight, which is defined as w = σ(MLP([x space ;x local ])), where x space is the concatenation of the spatial feature (x, y, z) and the Euclidean distance of a certain point, x local is the concatenation of spatial features of all local regions, [·;·] represents the concatenation operation, σ is the Sigmoid function, and the parameters of the MLP are optimized according to the training data.

4. The classification method according to claim 1, characterized in that The expression of the fused feature is z=x⊙w, where x is the input feature, w is the spatial distribution weight, ⊙ is the element-by-element multiplication, and the fused feature includes spatial distribution feature information.

5. The classification method according to claim 1, characterized in that: The method has an average accuracy MA ≥ 91.3% and an overall accuracy OA ≥ 93.4% on the ModelNet40 dataset. When the number of input point clouds is 64, the classification accuracy remains above 89%.

6. The classification method according to claim 1, characterized in that In the four-level LSSPM, the structure of each level module includes a feature extraction layer, a spatial weight generation layer and a feature fusion layer. The feature extraction layer uses a multi-layer perceptron (MLP) to extract point cloud features, the spatial weight generation layer generates spatial distribution weights, and the feature fusion layer multiplies the input features by the spatial distribution weights element by element to obtain fused features.

7. The classification method according to claim 1, characterized in that During the model training process, the cross-entropy loss function is used to measure the difference between the predicted classification result and the true label. The network parameters are updated through the back-propagation algorithm to minimize the loss function.

8. The classification method according to claim 1, characterized in that: The method is robust to inputs of different point cloud densities. The robustness is measured by changing the number of point clouds in the input network. When the point cloud sparsity is reduced, the recognition rate decreases slightly, but still maintains a high classification accuracy.

9. The classification method according to claim 1, characterized in that: The method also includes an ablation experiment step, using the original network as a baseline model, introducing LSSPM, and comparing and analyzing the performance differences under various configurations to verify the effectiveness of LSSPM. The ablation experiment results show that after introducing LSSPM, MA is improved by 3.8% and OA is improved by 2.5%.

10. The classification method according to claim 1, characterized in that: The method can be applied to scenarios such as smart city construction, unmanned driving systems, precision agriculture monitoring, and augmented reality applications to classify and process three-dimensional point cloud data, thereby improving the parsing accuracy and utilization efficiency of point cloud data.