A Point Cloud Classification / Segmentation Method Based on Local Geometric Contours and Global Structure Preservation

Through the multi-level feature extractor and cross-attention mechanism, the problem of insufficient integration of local geometric contours and global structure information in point cloud classification/segmentation is solved, and higher classification/segmentation accuracy and robustness are achieved.

CN120014363BActive Publication Date: 2025-08-01UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510188304.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-08-01
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The existing point cloud classification and segmentation methods are insufficient in processing point cloud surface contour information, especially inadequate integration of local geometric contours and global structural semantic information, resulting in insufficient classification/segmentation accuracy and robustness.

Method used

A multi-level tandem feature extractor is adopted to extract local features and global features through local neighborhood and global neighborhood construction modules, and feature fusion is used to combine local geometric contours and global structural information to improve the accuracy and robustness of feature extraction.

Benefits of technology

It significantly improves the accuracy and robustness of point cloud classification/segmentation, improves the processing ability of complex point cloud data, and has good online adaptability and real-time processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014363B_ABST
    Figure CN120014363B_ABST
Patent Text Reader

Abstract

The present invention discloses a point cloud classification / segmentation method based on local geometric contours and global structure preservation. First, local and global neighborhood relation graphs of the point cloud are extracted from the perspectives of spatial distance and feature space respectively. Among them, spatial distance calculation and the k-nearest neighbor algorithm accurately capture complex geometric and semantic information in the point cloud data, enhancing the accuracy and robustness of feature extraction. At the same time, the neighborhood relation graph self-diffusion mechanism further strengthens the feature diffusion of the point cloud surface contour to fully utilize the local geometric contour information and improve the expression ability of the local structure. Then, local features and global features are respectively extracted from the local neighborhood relation graph and the global neighborhood relation graph and sent into the feature fusion module. The cross-attention mechanism is used to effectively fuse the local features and global features, adaptively strengthening the key information, thereby improving the accuracy of classification and segmentation. This fusion method enables the classification / segmentation output head to more accurately capture features of different scales, improves the processing ability of complex point cloud data, enables it to generate accurate classification / segmentation results according to the input fused features, and significantly improves the classification / segmentation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of dialogue recommendation, and more specifically, relates to a point cloud classification / segmentation method based on local geometric contours and global structure preservation. Background Art

[0002] With the continuous development of 3D scanning technology, point cloud data is increasingly widely used in fields such as computer vision, robotics, autonomous driving, and urban modeling. Point cloud data usually consists of a large number of three-dimensional coordinate points and contains rich geometric shapes and structural information. Therefore, point cloud classification and segmentation tasks play an important role in these applications. However, existing point cloud classification and segmentation methods generally have some technical bottlenecks, especially in dealing with the surface contour information of point clouds.

[0003] Traditional point cloud data processing methods, especially deep learning-based graph convolutional networks (GCNs) and convolutional neural networks (CNNs), often focus on extracting global or local features of point clouds. For example, in the invention patent application published on October 20, 2023, with publication number CN116912561A, a point cloud data classification and segmentation method based on a multi-view adaptive graph convolutional network is disclosed. By constructing a graph network in the offline stage, the collected point cloud data is input into an adaptive rotation matrix generator to obtain adaptively rotated point cloud data and generate three multi-view projection images; the point cloud data and the three multi-view projection images are respectively constructed to obtain global information graph data and local information graph data, which are then respectively input into a global feature extraction network and a local feature extraction network. After the global features and local features are extracted, the two are fused and input into the functional neural network output head. The loss function is calculated based on the obtained results to realize the training of the graph network; in the online stage, the point cloud data collected in real time is input into the trained graph network to obtain the point cloud classification or segmentation results. The present invention adaptively rotates the point cloud data to obtain an optimized viewing angle, flexibly adjusts the position of the viewing angle according to the geometric features and spatial distribution of the point cloud data, and introduces connections between points of different depths on a specific projection surface, thereby better capturing the local and global features of the point cloud data and making the composition more accurate and robust. The global feature extraction network of this invention patent application is a six-layer tandem neural network with a maximum pooling layer, and the local feature extraction network is a two-layer tandem neural network. It simply uses a convolutional neural network to extract global features and local features and splice them together. Some other methods extract local features by constructing a neighborhood relationship graph and using spatial distance calculations. However, these methods ignore the fact that local neighborhood construction methods based on spatial distance usually rely on the geometric distance between points to select neighbors. This method is easily affected by changes in the density of point cloud distribution in point clouds. In dense areas, neighbor points may be too concentrated, resulting in information redundancy and deviation in feature expression; while in sparse areas, neighbor points may be insufficient, resulting in incomplete local feature information. Due to the sparsity and complexity of point cloud data, especially in surface contour areas such as the edges of objects, existing methods often do not process these areas sufficiently, resulting in limited improvements in point cloud classification / segmentation accuracy.

[0004] Furthermore, existing point cloud data processing often simply concatenates local and global features, leaving them independent and potentially leading to inconsistent or redundant information. This is particularly true for scenes with complex shapes and rich semantics, where existing technologies often fail to fully integrate local geometric contour information with global structural semantics. This results in insufficient robustness and generalization capabilities for point cloud classification / segmentation models when processing diverse, unstructured point cloud data.

[0005] Therefore, how to overcome the limitations of existing technologies, make full use of the local geometric contour information and global structural semantic information in point cloud data, and improve the accuracy and robustness of point cloud classification and segmentation has become a key issue that needs to be urgently solved in the current field of point cloud analysis. Summary of the Invention

[0006] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a point cloud classification / segmentation method based on local geometric contours and global structure preservation, so as to fully consider the local geometric contour information and global structural semantic information, and use the cross-attention mechanism to perform multi-scale feature fusion of local features and global features to improve the accuracy and robustness of point cloud classification / segmentation.

[0007] To achieve the above-mentioned object of the invention, the present invention provides a point cloud classification / segmentation method based on local geometric contours and global structure preservation, which is characterized by comprising the following steps:

[0008] (1) Use multi-stage cascade feature extractors to extract fusion features from point cloud data

[0009] The first-level feature extractor takes the three-dimensional coordinates of the midpoints in the point cloud data as input features, and the subsequent feature extractors take the fused features output by the previous-level feature extractor as input features. Each level of feature extractors includes a local neighborhood construction module, a global neighborhood construction module, a local feature extraction network, a global feature extraction network, and a feature fusion module. In each level of feature extractors, the input features are input into the local neighborhood construction module and the global neighborhood construction module respectively.

[0010] 1.1) Obtain local neighborhood relationship graph and global neighborhood relationship graph

[0011] In the local neighborhood construction module, the spatial distance between each point and other points is first calculated based on the input features. Then, based on the spatial distance, the neighborhood point set of each point is determined from the point cloud data, that is, the neighborhood relationship graph is constructed by selecting the first k minimum spatial distances. The obtained neighborhood relationship graph is then propagated and enhanced using the self-diffusion mechanism, thereby quickly expanding to the contour points on the point cloud surface. Finally, the first k maximum values and indices on the specified dimension in the diffused neighborhood relationship graph are determined as the local neighborhood relationship graph.

[0012] In the global neighborhood construction module, the spatial distance between each point and other points is first calculated based on the input features, and then each point finds the k nearest neighbor indices as the global neighborhood relationship graph;

[0013] 1.2) Extract local features and global features

[0014] Both the local feature extraction network and the global feature extraction network are multi-layer cascade neural networks. Each layer of the neural network consists of a two-dimensional convolution layer and a maximum pooling layer. The feature update function of each layer is:

[0015]

[0016] Among them, h i (n) is the input feature of the i-th point of the n-th layer neural network, h i (n+1) is the output feature of the i-th point of the n-th layer neural network, Θ is the parameter of the two-dimensional convolutional layer, max() is the maximum pooling function, N(i) represents the neighborhood point set of the i-th point, and j represents the j-th point in the neighborhood point set N(i);

[0017] The local neighborhood relationship graph is used as the input feature to the first layer of the local feature extraction network. The output feature of the last layer of the local feature extraction network is the extracted local feature F. local The global neighborhood relationship graph is used as the input feature to the first layer of the global feature extraction network. The output feature of the last layer of the global feature extraction network is the extracted global feature F global ;

[0018] 1.3) Use the cross attention mechanism to fuse local features and global features to obtain fused features

[0019] In the feature fusion module, the cross attention mechanism is used to fuse local features and global features: First, a set of learnable global query values Q is introduced, and for a given local feature F local Use the cross attention mechanism to aggregate it with the global query value Q to obtain the intermediate local feature CA(Q,F local ), then, the intermediate local feature CA(Q,F local ) and the global feature F global Further fusion is performed to obtain the fusion feature F fusion :

[0020] F fusion =CA(F global ,CA(Q,F local ))

[0021] Where CA(·,·) represents the cross attention operation;

[0022] (2) Point cloud classification / segmentation

[0023] Using the classification / segmentation output head, the fusion feature F obtained from each point in the point cloud data is fusion Perform classification / segmentation to obtain point cloud classification / segmentation results.

[0024] The object of the present invention is achieved as follows.

[0025] The point cloud classification / segmentation method based on local geometric contour and global structure preservation of the present invention first extracts the local and global neighborhood relationship graphs of the point cloud from the perspectives of spatial distance and feature space through a local neighborhood construction module and a global neighborhood construction module. Among them, spatial distance calculation and the k-nearest neighbor algorithm accurately capture the complex geometric and semantic information in the point cloud data, enhancing the accuracy and robustness of feature extraction. At the same time, the neighborhood relationship graph self-diffusion mechanism further strengthens the feature diffusion of the point cloud surface contour to make full use of the local geometric contour information and improve the expression ability of the local structure. Then, local features and global features are respectively extracted from the local neighborhood relationship graph and the global neighborhood relationship graph and sent into the feature fusion module, where the cross-attention mechanism is used to effectively fuse the local features and global features, adaptively strengthening the key information, thereby improving the accuracy of classification and segmentation. This fusion method enables the classification / segmentation output head to more accurately capture features of different scales, improves the processing ability of complex point cloud data, enables it to generate accurate classification / segmentation results according to the input fusion features, significantly improves the classification / segmentation accuracy, has good online adaptability and real-time processing ability, and meets the high-efficiency and reliability requirements in practical applications.

[0026] Through the organic combination of modular design and innovative technologies, the present invention shows significant advantages in task adaptability, expansion ability, and combining multi-view composition. The modular design enables each functional module to be independently optimized and flexibly combined, thereby improving the adaptability of the system under different tasks. Whether it is the classification or segmentation task of point cloud data, it can be efficiently configured and adjusted according to specific requirements. In addition, the feature fusion mechanism enhances the processing ability of the system for complex point cloud data. Especially in the application scenario of multi-view composition, by effectively fusing information from different views, the robustness and generalization ability of the model for diverse data are improved. Therefore, the present invention not only performs excellently in traditional point cloud analysis tasks but also has good scalability and can adapt to the needs of more complex tasks in the future. Brief Description of the Drawings

[0027] Figure 1 is a flowchart of a specific implementation manner of the point cloud classification / segmentation method based on local geometric contour and global structure preservation of the present invention;

[0028] Figure 2 is the overall network structure diagram of the point cloud classification / segmentation method based on local geometric contour and global structure preservation of the present invention;

[0029] Figure 3 is the effect diagram of the self-diffusion of the neighborhood relationship graph for any point in the point cloud data, where (a) is before self-diffusion and (b) is after self-diffusion;

[0030] Figure 4 They are the point cloud segmentation effect diagrams of 'airplane', 'chair', 'table' and 'lamp' in the ShapNetPart dataset. Specific implementation manners

[0031] The specific implementation manners of the present invention will be described below with reference to the accompanying drawings so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.

[0032] Embodiment 1

[0033] Figure 1 It is a flowchart of a specific implementation manner of the point cloud classification / segmentation method based on local geometric contours and global structure preservation of the present invention.

[0034] In this embodiment, as Figure 1 shown, the point cloud classification / segmentation method based on local geometric contours and global structure preservation of the present invention includes the following steps:

[0035] Step S1: Use a multi-stage cascaded feature extractor to perform fused feature extraction on the point cloud data

[0036] In this embodiment, as Figure 2 shown, use a four-stage cascaded feature extractor 1-4 to perform fused feature extraction on the point cloud data. The first-stage feature extractor uses the three-dimensional coordinates of the points in the point cloud data as the input feature, and the subsequent feature extractors use the fused feature output by the previous-stage feature extractor as the input feature. Each stage of the feature extractor includes a local neighborhood construction module, a global neighborhood construction module, a local feature extraction network, a global feature extraction network, and a feature fusion module. In each stage of the feature extractor, the input feature is respectively input into the local neighborhood construction module and the global neighborhood construction module.

[0037] In this embodiment, through a lidar (LiDAR) scan or a three-dimensional model generation tool, point cloud data with three-dimensional coordinates P = [x, y, z] is obtained, and the data size is N×3, where N is the number of points in the point cloud.

[0038] Step S1.1: Obtain a local neighborhood relationship graph and a global neighborhood relationship graph

[0039] In the local neighborhood construction module, first, the spatial distance between each point and other points is calculated based on the input features. Then, according to the spatial distance, the neighborhood point set of each point is determined from the point cloud data, that is, a neighborhood relationship graph is constructed by selecting the first k minimum spatial distances. Then, the obtained neighborhood relationship graph is used with the self-diffusion mechanism for feature propagation and enhancement, so as to quickly expand to the surface contour points of the point cloud. Finally, the first k maximum values and indices in the specified dimension of the neighborhood relationship graph after diffusion are determined as the local neighborhood relationship graph.

[0040] In this embodiment, the local neighborhood construction module consists of a spatial distance calculation algorithm, a k-nearest neighbor algorithm, a neighborhood relationship graph symmetrization method, a neighborhood relationship graph normalization algorithm, a neighborhood relationship graph self-diffusion algorithm, and a tensor maximum k-element selection operation. Among them, the spatial distance calculation algorithm is used to calculate the spatial distance between each point and other points based on the input features. The k-nearest neighbor algorithm is used to determine the neighborhood point set of each point from the point cloud data according to the spatial distance. This algorithm constructs a neighborhood relationship graph by selecting the first k minimum values. The neighborhood relationship graph symmetrization method is used to convert the asymmetric neighborhood relationship graph into a symmetric matrix to satisfy the positive definiteness of the matrix diffusion kernel. The neighborhood relationship graph normalization algorithm is used to normalize the neighborhood relationship graph of the graph structure, so as to alleviate the numerical instability problem that may occur during the point diffusion process of the neighborhood relationship graph.

[0041] The normalization algorithm of the neighborhood relationship graph is:

[0042]

[0043] where A represents the neighborhood relationship graph and D represents the degree matrix of the neighborhood relationship graph.

[0044] The neighborhood relationship graph self-diffusion algorithm uses the self-diffusion mechanism to perform feature propagation and enhancement on the obtained point cloud neighborhood relationship graph, so as to quickly expand to the surface contour points of the point cloud. The tensor maximum k-element selection operation is used to determine the first k maximum values and indices in the specified dimension of the neighborhood relationship graph after diffusion.

[0045] The self-diffusion mechanism calculation formula of the neighborhood relationship graph self-diffusion algorithm is:

[0046]

[0047] where n represents the number of diffusion times. In this embodiment, n = 5, and the self-diffusion effect diagram of the neighborhood relationship graph is as Figure 3As shown, the neighborhood before diffusion is the nearest neighborhood based on the central point. Since the local geometric structure is ignored, some points belonging to different components are included in the neighborhood of the central point, resulting in noise in feature extraction. After using the self-diffusion method, the neighborhood of the central point changes adaptively according to the local geometric structure of the point cloud, avoiding noise while expanding the neighborhood range.

[0048] In the global neighborhood construction module, first, the spatial distance between each point and other points is calculated based on the input features, and then each point finds the indices of the k nearest neighbors as the global neighborhood relationship graph.

[0049] In this embodiment, the global neighborhood construction module consists of a feature space distance calculation algorithm and a k-nearest neighbor algorithm. Among them, the feature space distance calculation algorithm is used to calculate the distance between each point and other points in the feature space, and the k-nearest neighbor algorithm finds the indices of the k nearest neighbors for each point.

[0050] Step S1.2: Extract local features and global features

[0051] Both the local feature extraction network and the global feature extraction network are multi-layer cascaded neural networks. Each layer of the neural network consists of a two-dimensional convolutional layer and a max pooling layer. The feature update function for each layer is:

[0052]

[0053] where h i (n) is the input feature of the i-th point in the n-th layer of the neural network, h i (n+1) is the output feature of the i-th point in the n-th layer of the neural network, Θ is the parameter of the two-dimensional convolutional layer, max() is the max pooling function, N(i) represents the set of neighborhood points of the i-th point, and j represents the j-th point in the neighborhood point set N(i).

[0054] The local neighborhood relationship graph is used as the input feature to the first layer of the neural network of the local feature extraction network, and the output feature of the last layer of the local feature extraction network is the extracted local feature F local The global neighborhood relationship graph is used as the input feature to the first layer of the neural network of the global feature extraction network, and the output feature of the last layer of the global feature extraction network is the extracted global feature F global .

[0055] In this embodiment, both the local feature extraction network and the global feature extraction network are four-layer cascaded neural networks. Among them, the two-dimensional convolutional layer in the first layer is a graph convolutional layer with an input dimension of 3 and an output dimension of 64, the two-dimensional convolutional layer in the second layer is a graph convolutional layer with an input dimension of 64 and an output dimension of 64, the two-dimensional convolutional layer in the third layer is a graph convolutional layer with an input dimension of 64 and an output dimension of 128, and the two-dimensional convolutional layer in the fourth layer is a graph convolutional layer with an input dimension of 128 and an output dimension of 256.

[0056] Step S1.3: Use the cross-attention mechanism to fuse the local feature and the global feature to obtain the fused feature

[0057] In this embodiment, as Figure 2 shown, in the feature fusion module, the cross-attention mechanism is used to fuse the local feature and the global feature: First, a set of learnable global query values Q are introduced, and for the given local feature F local use the cross-attention mechanism to aggregate it with the global query value Q to obtain the intermediate local feature CA(Q,F local ), then, the intermediate local feature CA(Q,F local ) is further fused with the global feature F global to obtain the fused feature F fusion :

[0058] F fusion = CA(F global ,CA(Q,F local ))

[0059] where CA(·,·) represents the cross-attention operation.

[0060] Step S2: Point cloud classification / segmentation

[0061] Adopt a classification / segmentation output head, and perform classification / segmentation according to the fused feature F fusion obtained from each point in the point cloud data to obtain the point cloud classification / segmentation result.

[0062] In this embodiment, the classification / segmentation output head is a functional neural network output head including several layers of multi-layer perceptrons, and the output dimension is the number of categories in the classification task or the number of parts to be segmented in the segmentation task.

[0063] According to the fused feature F fusion obtained from each point in the point cloud data, input it into the functional neural network output head, and update the network weights after calculating the loss function based on the obtained result. The functional neural network output head is a multi-layer perceptron with a series structure of two linear layers and one max pooling layer, and its output result is a vector with the dimension of the number of categories in the classification task, and the index corresponding to the maximum value is the final result predicted by the network.

[0064] The step of updating the network weights is as follows: calculate the cross-entropy error between the obtained vector and the dataset labels and backpropagate to update the network weights. For updating the network weights, the publicly available dataset ModelNet40 is selected, where the training set has 9,843 point clouds, the test set has 2,468 point clouds, and there is no validation set. 1,024 points are randomly sampled from each point cloud. The optimizer used during the training process is the Adam optimizer with a learning rate of 0.001, and the number of training epochs is 300. The finally trained graph network is obtained, which includes a local neighborhood construction module, a global neighborhood construction module, a feature extraction network, a local feature extraction network, a global feature extraction network, and a classification / segmentation output head, i.e., a functional neural network output head.

[0065] Through specific actual experiments, using PyTorch as the underlying deep learning framework and the hardware environment of NVIDIA RTX 4090 GPU (24GB), running the above method, on the officially divided test set, the average instance accuracy reaches 94.2%, which is 1.3% higher than the existing technology used as a baseline, such as Wang et al. in "Dynamic graph cnn for learning on point clouds" ([J]. Acm Transactions On Graphics (tog), 2019, 38(5): 1-12).

[0066] Example 2

[0067] In this example, the ShapeNetPart dataset is selected for training and testing, which includes 12,137 training set samples, 1,870 validation set samples, and 2,874 test set samples. The number of points in each point cloud is about 2,000, and the specific number of point clouds may vary. During the training process, the Adam optimizer with a learning rate of 0.001 is adopted, and the number of training epochs is 200.

[0068] In the point cloud segmentation task, in this example, local geometric features are first extracted through the local neighborhood construction module, and then global semantic information is extracted through the global neighborhood construction module. The local features and global features after feature extraction are fused through the cross-attention mechanism to obtain a more accurate point cloud representation.

[0069] In the experiment, PyTorch is used as the underlying deep learning framework, and the hardware environment of NVIDIA RTX 4090 GPU (24GB) is used for training and testing. After verification on the training set and the test set, the average instance accuracy on the test set reaches 86.5%, which is 1.4% higher than the existing technology; the average class accuracy is 83.5%, which is 1.2% higher than the existing technology.

[0070] Figure 4 They are the point cloud segmentation effect diagrams of 'airplane', 'chair', 'table' and 'desk lamp' in the ShapNetPart dataset. Among them, different colors represent different components. From Figure 4 it can be seen that for categories such as 'airplane', 'chair' and 'desk lamp', each component has been accurately segmented, and the point cloud segmentation performance is better than existing methods.

[0071] Although the above description of the illustrative specific embodiments of the present invention is provided for the understanding of those skilled in the art of the present technology, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.

Claims

1. A point cloud classification / segmentation method based on local geometric contours and global structure preservation, characterized in that It includes the following steps: (1) Use a multi-level cascaded feature extractor to perform fused feature extraction on point cloud data The first-level feature extractor takes the three-dimensional coordinates of points in the point cloud data as input features. The subsequent feature extractors take the fused features output by the previous-level feature extractor as input features. Each level of the feature extractor includes a local neighborhood construction module, a global neighborhood construction module, a local feature extraction network, a global feature extraction network, and a feature fusion module. In each level of the feature extractor, the input features are respectively input into the local neighborhood construction module and the global neighborhood construction module; 1.1) Obtain the local neighborhood relationship graph and the global neighborhood relationship graph In the local neighborhood construction module, first calculate the spatial distance between each point and other points based on the input features. Then, according to the spatial distance, determine the neighborhood point set of each point from the point cloud data, that is, construct a neighborhood relationship graph by selecting the first k minimum spatial distances. Then, use the self-diffusion mechanism to propagate and strengthen the features of the obtained neighborhood relationship graph, so as to quickly expand to the surface contour points of the point cloud. Finally, determine the first k maximum values and indexes on the specified dimension in the diffused neighborhood relationship graph as the local neighborhood relationship graph; In the global neighborhood construction module, first calculate the spatial distance between each point and other points based on the input features. Then, each point finds the indexes of the k nearest neighbors as the global neighborhood relationship graph; 1.2) Extract local features and global features Both the local feature extraction network and the global feature extraction network are multi-layer cascaded neural networks. Each layer of the neural network consists of a two-dimensional convolutional layer and a max pooling layer. The feature update function of each layer is: where h i (n) is the input feature of the i-th point in the n-th layer of the neural network, and h i (n+1) is the output feature of the i-th point in the n-th layer of the neural network, Θ is the parameter of the two-dimensional convolutional layer, max() is the max pooling function, N(i) represents the set of neighborhood points of the i-th point, and j represents the j-th point in the neighborhood point set N(i); The local neighborhood relationship graph is used as an input feature to the first neural network layer of the local feature extraction network, and the output feature of the last layer of the local feature extraction network is the extracted local feature F local , and the global neighborhood relationship graph is used as an input feature to the first neural network layer of the global feature extraction network, and the output feature of the last layer of the global feature extraction network is the extracted global feature F global ; 1.3) Use the cross-attention mechanism to fuse local features and global features to obtain fused features In the feature fusion module, the cross-attention mechanism is used to fuse local features and global features: First, a set of learnable global query values Q are introduced, and for the given local feature F local The cross-attention mechanism is used to aggregate it with the global query value Q to obtain the intermediate local feature CA(Q, F local ), then, the intermediate local feature CA(Q, F local ) is further fused with the global feature F global to obtain the fused feature F fusion : F fusion = CA(F global , CA(Q, F local )) Among them, CA(·,·) represents the cross-attention operation; (2) Point cloud classification / segmentation Using a classification / splitting output header, perform classification / splitting based on the fusion feature F obtained from each point in the point cloud data to obtain the point cloud classification / splitting result. fusion Perform classification / splitting to obtain the point cloud classification / splitting result.

2. The point cloud classification / segmentation method based on local geometric contours and global structure preservation according to claim 1, wherein Before using the self-diffusion mechanism to propagate and strengthen the features of the neighborhood relationship graph described in step 1.1), it is necessary to convert the asymmetric neighborhood relationship graph into a symmetric matrix and perform normalization processing. The normalization of the neighborhood relationship graph is: Among them, A represents the neighborhood relationship graph, and D represents the degree matrix of the neighborhood relationship graph.

3. The point cloud classification / segmentation method based on local geometric contour and global structure preservation according to claim 1, wherein The calculation formula of the self-diffusion mechanism described in step 1.1) is: Among them, n represents the number of diffusion times.

4. The point cloud classification / segmentation method based on local geometric contours and global structure preservation according to claim 1, wherein Both the local feature extraction network and the global feature extraction network described in step 1.2) are four-layer cascaded neural networks. Among them, the two-dimensional convolutional layer of the first layer is a graph convolutional layer with an input dimension of 3 and an output dimension of 64. The two-dimensional convolutional layer of the second layer is a graph convolutional layer with an input dimension of 64 and an output dimension of 64. The two-dimensional convolutional layer of the third layer is a graph convolutional layer with an input dimension of 64 and an output dimension of 128. The two-dimensional convolutional layer of the fourth layer is a graph convolutional layer with an input dimension of 128 and an output dimension of 256.

5. The point cloud classification / segmentation method based on local geometric contours and global structure preservation according to claim 1, characterized in that, The classification / segmentation output head described in step (2) is a functional neural network output head including several layers of multi-layer perceptrons, and the output dimension is the number of categories in the classification task or the number of components to be segmented in the segmentation task.

Citation Information

Patent Citations

  • Point cloud data classification and segmentation method based on multi-view adaptive graph convolutional network

    CN116912561A

  • Point cloud classification and segmentation method based on point cloud channel attention feature fusion mechanism

    CN119131483A

  • Three-dimensional point cloud task analysis method fusing local features and global context information

    CN119445291A