Graph Convolutional Neural Network Module with Attention Mechanism for 3D Point Cloud Scenes
By introducing attention mechanism and K-NN algorithm into the graph convolution neural network module, the local feature map structure is constructed and feature weights are calculated, and the problem of insufficient local feature extraction and feature aggregation capabilities in the existing technology is solved, and efficient point cloud processing and semantic segmentation effects are achieved.
Patent Information
- Application Number
- CN202111618088.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-27
AI Technical Summary
When processing three-dimensional point cloud data, the local feature extraction capability is poor and the feature aggregation capability is insufficient, resulting in important information being filtered and it is impossible to clearly distinguish between effective and invalid features.
A graph convolution neural network module based on attention mechanism is proposed. K neighboring points of the center point are obtained through the K-NN algorithm, a local feature map structure is constructed, and the weights of different features are calculated using attention pooling strategy to extract the most important local features.
It realizes efficient point cloud classification and segmentation methods, improves the performance of semantic segmentation tasks, is suitable for industrial needs, and performs well in applications such as land change analysis, urban modeling, and road segmentation.
Smart Images

Figure CN114358246B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of point cloud data, and in particular to a graph convolutional neural network module with an attention mechanism for three-dimensional point cloud scenes. Background Art
[0002] Point cloud is a collection of discrete points in three-dimensional space. Compared with ordinary remote sensing images, point cloud data carries more spatial information. Therefore, it is of great value for tasks such as surface monitoring. The research on three-dimensional point cloud data has a wide range of applications in society and other fields. It mainly includes road segmentation, 3D city modeling, autonomous driving, face recognition, forest monitoring, etc. The research on three-dimensional point cloud data mainly focuses on exploring the three-dimensional information and deep features carried by point cloud data. The scientific research on three-dimensional point cloud data has a long history. From artificial geometry to deep learning, it has widely promoted the progress of hard science and society. Three-dimensional point cloud data has more spatial information than traditional images, which poses more challenges and brings opportunities in the field of point cloud. At the same time, due to the successful application of convolutional neural networks (CNN) in image classification, target detection, semantic segmentation, etc., deep learning methods have emerged. At the same time, inspired by these excellent results, point cloud research has also shifted from traditional machine learning to flexible neural network structures. This patent also analyzes the application of point cloud data in actual industry and commerce from the perspective of deep learning.
[0003] Prior art 1
[0004] Hang Su et al. processed 3D point cloud data by projecting the point cloud into 2D. They also used existing 2D image processing methods to perform tasks such as data classification and segmentation. Charles et al. first proposed using symmetric functions to process point cloud data through mathematical theory derivation to meet the permutation invariance of point cloud data. Daniel Maturana et al. applied deep learning methods such as CNN by voxelizing point clouds. Generally speaking, early research on 3D point cloud data mainly used theoretical knowledge from various disciplines such as geometry to analyze shallow information. The most advanced methods always imply deep learning represented by convolutional neural networks (CNN). In addition, it has also achieved great achievements and outstanding performance in semantic segmentation, classification, and target detection.
[0005] It is undeniable that CNN is likely to be the de facto standard in deep learning. However, its parameters grow exponentially with the increase of convolutional layers, and its size increases with the increase of computing power. Projection and voxelization usually bring huge memory usage and computational consumption. In addition, due to the persistence of multiplication and addition operations, computational consumption is a bottleneck for industrial applications and cannot meet the real-time requirements of industry.
[0006] Prior art 2:
[0007] SEGCloud processes the point cloud data by dividing it into several small point clouds and applying trilinear interpolation and conditional random fields. Charles et al. proposed a method to gradually expand the receptive field by improving the PointNet network, and also improved its slightly large computational complexity. Recently, benefiting from the successful expansion of graphs and other nonlinear structures in the field of deep learning, graph convolutional neural networks have achieved state-of-the-art performance in computer vision and have attracted the attention of many researchers. Inspired by this, AdaptConv uses dynamic convolution kernels to make convolution operations more flexible. 3D-GCN also designed a learnable convolution kernel to obtain local features and showed good learning ability. DGCNN proposed a convolution method called EdgeConv, which can dynamically calculate the graph structure of each network layer and aggregate the features of the central nodes in the local graph and their corresponding edge features. The graph structure is obtained by the K-NN method, which extracts local features better. With the success of the attention mechanism in the field of natural language processing, more and more people are applying it in the field of computer vision. GAPNet, GACNet and LAE-Conv all obtain point cloud features by designing attention modules. These graph-based point cloud processing methods and attention-based point cloud processing methods all use maximum pooling to aggregate the features of local feature maps.
[0008] Disadvantages of the second prior art
[0009] The above methods all use a simple max pooling strategy to aggregate local feature information. Therefore, it leads to many disadvantages, such as important information being filtered and the inability to clearly distinguish between valid and invalid features. Max pooling directly selects the largest data among all features, but other data do contribute to feature extraction during the calculation process, so a lot of useful information is often discarded by the max pooling strategy.
[0010] In view of the defect of poor local feature extraction ability of existing models, the present invention proposes a graph convolutional neural network module based on the attention mechanism, which uses the K-NN algorithm to obtain K neighboring points of the center point and sequentially constructs a local feature graph structure. The local feature graph structure can extract the local topological structure of the point cloud data to better represent the local features, thus well solving the defects of the existing model.
[0011] In view of the poor feature aggregation ability of the DGCNN model, this paper proposes a graph convolutional neural network based on the attention pooling strategy. It obtains several neighboring points of each center point by the K-NN algorithm, calculates different attention weights, and calculates and extracts the most important local features of the current input data through the weights. Summary of the invention
[0012] In order to solve the problems existing in the prior art, the present invention provides a graph convolutional neural network module with an attention mechanism for three-dimensional point cloud scenes, which solves the defects of poor local feature extraction capability of existing models and poor feature aggregation capability of DGCNN models.
[0013] The technical solution provided by the present invention is as follows:
[0014] The graph convolutional neural network module of the attention mechanism of the three-dimensional point cloud scene includes: an attention graph encoding module, namely the AGEM module, and an attention pooling module, namely the AP module.
[0015] Preferably, the AGEM module comprises the following steps:
[0016] S1: First, use the K-NN algorithm to obtain the k nearest points of the center point through the given k value and form a local point cloud set;
[0017] S2: The input point cloud features are then raised to the same size as the k-neighboring point set through the Repeat operation;
[0018] S3: Encode the k-neighboring point set and the original point set to obtain high-dimensional features;
[0019] S4: Concatenate the acquired high-dimensional features and pass them into the AP module as input features.
[0020] Preferably, the AP module includes: attention weight calculation, attention weight mask, and MPL module.
[0021] The beneficial effects of the graph convolutional neural network module of the attention mechanism of the three-dimensional point cloud scene of the present invention are as follows:
[0022] 1. The present invention realizes a highly efficient point cloud classification and segmentation method, which exceeds previous methods and therefore has certain commercial value.
[0023] 2. For semantic segmentation tasks, the model size is only 2.03M, which is well suited to industrial needs.
[0024] 3. This method can be well applied to the point cloud field, such as land change analysis, urban modeling, road segmentation, etc.
[0025] 4. Monitor forest changes and the impact of terrain on forest dynamics, and classify forest tree species. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is the AGM attention graph convolution module of the present invention.
[0027] Figure 2 It is the AP attention pooling module of the present invention.
[0028] Figure 3 The data flow chart of the present invention.
[0029] Figure 4 This is a visualization diagram of the results of the method of the present invention. DETAILED DESCRIPTION
[0030] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0031] The present invention uses the existing deep learning framework Pytorch and the corresponding programming library, mainly including Numpy, Pandas, Tensor, etc. Among them, Pytorch mainly uses deep learning models, including linear modules, convolution modules, parameter penalty modules, etc.
[0032] The specific implementation principle of the scheme is as follows:
[0033] We convert segmentation into a classification task in the proposed method, performing pixel-level classification instead of patch segmentation, as Figure 1 As shown. The main method is AGM, the attention graph convolution module. The method includes AGEM (attention graph encoding) module and AP, the attention pooling, module. For the AGEM module, it changes the features of the data to meet the attention calculation needs of the AP module. Specifically, the K-NN algorithm is first used to obtain the k nearest points of the center point through the given k value, and form a local point cloud set. Then the input point cloud features are increased to the same size as the k neighboring point set through the Repeat operation. Next, the k neighboring point set is encoded with the original point set to obtain high-dimensional features, such as relative distance, relative coordinates, etc. Finally, the obtained high-dimensional features are spliced and passed into the AP module as input features.
[0034] For the AP module, it consists of attention weight calculation, attention weight mask, and MPL module. Figure 2 Shown
[0035] The data flow of this patent Figure 3 As shown below;
[0036] Finally, the present invention conducts extensive experiments on three widely adopted public datasets: ModelNet40 dataset for object classification, ShapeNetPart dataset for part segmentation, and S3DIS dataset for semantic segmentation. AGNet comprehensively outperforms the state-of-the-art methods, with an accuracy improvement of 4.2% on PointNet, 6.0% on ECC, 7.5% on VoxNet, and 8.7% on 3DShapeNets in the ModelNet40 dataset.
[0037] The results of this patented method can be visualized as follows Figure 4 As shown:
[0038] The network architecture is implemented as key protection points, such as Figure 1 As shown in the figure, the network consists of a single-layer MLP transformation, AGEM (attention graph convolutional encoding module) and AP (attention pooling module). The specific protection technology focuses are as follows:
[0039] Single-layer MLP transformation
[0040] The clean input data has great information redundancy between channels due to different dimensions, so a layer of MLP is used to adjust part of the dimensional information and match the input of the downstream model at the same time.
[0041] Local feature aggregation module
[0042] After obtaining data that has changed to a certain extent, for point cloud features, the K-NN algorithm can be used to obtain the k nearest points, and the semantic information of the searched neighboring points can be used to obtain high-dimensional features such as relative distance, and local feature maps can be constructed in turn to complete the aggregation of local feature information.
[0043] Attention Pooling Module
[0044] The obtained aggregated features are filtered and pooled in the channel dimension, and the attention mechanism is used to calculate the weights of different features to obtain important information in the features.
[0045] Output
[0046] The most common linear fully connected layer is used to input the final classification result. This part is a public technology and is not included in the key protection technology of this method.
Claims
1. A method for realizing 3D point cloud scene classification based on an attention mechanism graph convolutional neural network, characterized in that, the attention mechanism graph convolutional neural network includes an attention graph convolutional module AGM and a single-layer MLP transformation; the AGM includes an attention graph convolutional encoding module AGEM and an attention pooling module AP; the AGEM module comprises the following steps: S1: First, use the K-NN algorithm to obtain k nearest neighbors of the center point through a given k value and form a local point cloud set; S2: Subsequently, lift the input point cloud features to the same size as the k-nearest neighbor set through the Repeat operation; S3: Encode the k-nearest neighbor set and the original point set to obtain high-dimensional features; S4: Concatenate the obtained high-dimensional features and input them as input features into the AP module; the AP module includes: attention weight calculation, attention weight mask, MPL module Output: Use a linear fully connected layer to output the final classification result.
Citation Information
Patent Citations
Point cloud feature extraction model based on graph neural network and classification segmentation method
CN113554654A