A context-driven large-scale garden scene point cloud semantic segmentation method

By combining multi-scale spatial context information extraction with Transformer coding networks, the problem of insufficient network representation capabilities for semantic segmentation of point clouds in large-scale garden scenes is solved, achieving higher segmentation accuracy and stronger point cloud understanding capabilities.

CN117036694BActive Publication Date: 2025-11-04NANJING UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310907329.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-24
Publication Date
2025-11-04
Estimated Expiration
2043-07-24

AI Technical Summary

Technical Problem

Existing technologies are difficult to apply effectively to point cloud semantic segmentation in large-scale garden scenes, mainly because the representational capabilities of network structures are limited and cannot fully express the geometric complexity and unclear boundaries of objects in garden scenes.

Method used

We employ multi-scale spatial context information extraction, local Transformer encoding, and global Transformer encoding networks, combined with global-local context fusion and Transformer decoding, to improve the semantic segmentation capability of point clouds through the fusion of local and global features.

Benefits of technology

It improves the segmentation accuracy of point clouds in large-scale garden scenes, enhances the network's understanding of point clouds, and can more effectively extract contextual semantic information of segmentation tasks, reducing computational overhead while ensuring the integrity of point representation information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117036694B_ABST
    Figure CN117036694B_ABST
Patent Text Reader

Abstract

The application provides a context-driven large-scale garden scene point cloud semantic segmentation method, comprising the following steps: 1) multi-scale spatial context information extraction: extracting point cloud position coding, and aggregating local spatial information according to the near neighbor points of the point cloud, finally extracting point-by-point multi-scale spatial context information through feature splicing. 2) global-local context extraction: using a local Transformer encoding network to extract the local context of the point cloud, and using a global Transformer encoding network to extract the global context of the point cloud. 3) global-local context fusion and Transformer decoding: fusing the global and local features of the point cloud according to the correlation of the global features and the local features, and using a Transformer decoder to decode the features to obtain point-by-point class labels, completing the point cloud semantic segmentation. The method can accurately capture the global context information in the garden point cloud scene without increasing the calculation amount, improve the point-by-point representation ability, and thus improve the speed and accuracy of the point cloud segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of point cloud semantic segmentation, and particularly relates to a context-driven large-scale garden scene point cloud semantic segmentation method. BACKGROUND

[0002] Scene point cloud semantic segmentation aims to identify the semantic label of each point in the point cloud, which plays an important role in many real-life applications such as autonomous driving, robots, augmented reality and virtual reality. However, the large-scale garden scene point cloud has the characteristics of irregular point cloud structure, large spatial range, complex scene composition, large number of objects, irregular object geometry and unclear boundary between objects, and how to perform semantic segmentation on various scenes is still a challenging problem.

[0003] Prior point cloud semantic segmentation methods mostly focus on solving the problem of small indoor scene point cloud semantic segmentation, such as document 1: Shenlong Wang, Simon Suo, Wei-Chiu Ma, Andrei Pokrovsky, and Raquel Urtasun. Deep parametric continuous convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2589-2597, 2018.; document 2: Yangyan Li, Rui Bu, Mingchao Sun, Wei Wu, Xinhan Di, and Baoquan Chen. Pointcnn: Convolution on x-transformed points. In Advances in Neural Information Processing Systems, pages 820-830, 2018.; document 3: Pedro Hermosilla, Tobias Ristchel, Pere-Pau Vazquez, Alvaro Vinacua, and Timo Ropinski. Monte carlo convolution for learning on non-uniformly sampled point clouds. ACM Transactions on Graphics (TOG), 37(6):235-1, 2018.; document 4: Wu W, Qi Z, Fuxin L. Pointconv: Deep convolutional networks on 3d point clouds [C] / / Proceedings of the IEEE / CVF Conference on computer vision and pattern recognition. 2019:9621-9630.; document 5: Guo M H, Cai J X, Liu Z N, et al. PCT: Point cloud transformer [J]. Computational Visual Media, 2021, 7:187-199.; document 6: Zhao H, Jiang L, Jia J, et al.Point transformer [C] / / Proceedings of the IEEE / CVF international conference on computer vision. 2021:16259-16268.; Reference 7: X. Ye, J. Li, H. Huang, L. Du, and X. Zhang. 3D recurrent neural networks with context fusion for point cloud semantic segmentation. In The European Conference on Computer Vision (ECCV), September 2018; Reference 8: Chinese patent CN109410307B, a method for semantic segmentation of point clouds in a scene. These methods cannot be well applied to the semantic segmentation of point clouds in large-scale garden scenes, mainly because their network structure has limited representational capabilities.

[0004] Recently, some methods have been proposed for large-scale scene point cloud semantic segmentation, such as document 9: Landrieu L, Simonovsky M. Large-scale point cloud semantic segmentation with superpoint graphs [C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 4558-4567. proposed to convert large-scale scene point cloud into fewer superpoints and perform semantic segmentation on superpoints, but the encoding method of this method is relatively simple and cannot express the geometric shape of irregular objects. Document 10: Hu Q, Yang B, Xie L, et al. Randla-net: Efficient semantic segmentation of large-scale point clouds [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2020: 11108-11117. proposed to use random downsampling instead of time-consuming farthest point sampling in point cloud hierarchical encoding to improve the processing speed of the network. However, this method only aggregates local context information, which may not be enough for large-scale garden scene point cloud. Document 11: Fan S, Dong Q, Zhu F, et al. SCF-Net: Learning spatial contextual features for large-scale point cloud segmentation [C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021: 14504-14513. proposed to aggregate local and global features to enrich the representation ability of large-scale scene point cloud, but the global context they used is relatively low-level and cannot well express the global dependency relationship. Document 12: A large scene point cloud semantic segmentation method with publication number CN112819833A proposed a large scene semantic segmentation method, which uses dilated graph convolution and random sampling to improve the segmentation accuracy and inference speed of the model in large scenes.However, compared with previous indoor and outdoor scene point clouds, the garden point cloud has the characteristics of complex object geometry, unclear boundary between objects, and large space range occupied by objects, so that the above method is difficult to learn a robust feature representation in such a scene. SUMMARY

[0005] The technical problem to be solved by the present application is to provide a context-driven large-scale garden scene point cloud semantic segmentation method.

[0006] In order to solve the above technical problems, the present application discloses a context-driven large-scale garden scene point cloud semantic segmentation method, comprising the following steps:

[0007] Step 1, multi-scale spatial context information extraction: extract point cloud position encoding, and aggregate local spatial information according to the near neighbor points of the point cloud, and finally extract point-by-point multi-scale spatial context information through feature splicing;

[0008] Step 2, global-local context extraction: use local Transformer encoding network to extract local features of point cloud, and use global Transformer encoding network to extract global features of point cloud;

[0009] Step 3, global-local context fusion and Transformer decoding: fuse the global and local features of the point cloud according to the correlation of the global and local features, and use the Transformer decoder to decode the features to obtain the point-by-point class label, and complete the point cloud semantic segmentation.

[0010] Further, step 1 comprises the following steps:

[0011] Step 1-1, position encoding is performed on the input garden point cloud to obtain point-by-point position encoding features, and the input garden point cloud is denoted as N p , the position encoding feature is c represents the dimension of the point cloud position encoding feature;

[0012] Step 1-2, the near neighbor points include 8 near neighbors, 16 near neighbors and 32 near neighbors, and point cloud multi-scale neighborhood query is adopted to obtain point cloud 8 near neighbor encoding 16 near neighbor encoding and 32 near neighbor encoding

[0013] Step 1-3, local information aggregation, respectively, the point cloud 8 near neighbor point encoding 16 near neighbor encoding 32 near neighbor encoding are subjected to average pooling operation to obtain 8 near neighbor information 16 nearest neighbor information 32 nearest neighbor information

[0014] Step 1-4, feature concatenation, point-wise position encoding feature f pe 8 nearest neighbor information 16 nearest neighbor information and 32 nearest neighbor information Concatenate in the second dimension to obtain point cloud multi-scale spatial context information

[0015] Selecting multiple nearest neighbor information can make the context information contain multi-scale information, making the point-wise spatial representation more rich, which also makes the encoding network in the following text encode multi-scale information into the network.

[0016] Further, the point cloud multi-scale neighborhood query in step 1-2 includes the following steps:

[0017] Step 1-2-1, constructing a KD tree for the input garden point cloud;

[0018] Step 1-2-2, querying the 8 nearest neighbor points, 16 nearest neighbor points, and 32 nearest neighbor point indexes of the point cloud according to the KD tree

[0019] Step 1-2-3, querying position encoding features f according to the nearest neighbor point indexes pe 8 nearest neighbor encoding 16 nearest neighbor encoding and 32 nearest neighbor encoding

[0020] The method of using a KD tree to retrieve nearest neighbor points is advantageous in reducing the time complexity of the algorithm.

[0021] Further, step 2 includes the following steps:

[0022] Step 2-1, building a local Transformer encoding network

[0023] Step 2-2, building a global Transformer encoding network The global Transformer encoding network includes a key network F k , a value network F v , and a query network F q , the key network F k , the value network F v , and the value network F vEach of them includes a fully connected layer with 4c input channels and c2 output channels, where c2 represents the dimension of the global feature.

[0024] Step 2-3, using the farthest point sampling algorithm to extract the key points P of the garden point cloud key and the key point descriptor f key .

[0025] Step 2-4, the point cloud multi-scale spatial context information f mssca is input into the local Transformer encoding network to obtain the local feature f and the local point index I corresponding to the local feature loc , where m1 represents the number of local features, and c1 represents the dimension of the local feature; the key point descriptor f key is input into the global Transformer encoding network to obtain the global feature f and the key point index I corresponding to the global feature glb , where m2 represents the number of global features.

[0026] Further, step 2-3 includes the following steps:

[0027] Step 2-3-1, using the farthest point sampling algorithm to extract the point cloud key point index set I S .

[0028] Step 2-3-2, according to the point cloud key point index set I S , the key points P of the garden point cloud are extracted from the garden point cloud key = P[I S ]; according to the point cloud key point index set I S , the key point descriptor f mssca of the garden point cloud is extracted from the point cloud multi-scale spatial context information f key = f mssca [I S ].

[0029] Extracting key points and key descriptors can describe point cloud features with as few resources as possible while reducing computational overhead, ensuring the effectiveness of the attention mechanism while reducing the amount of subsequent calculations.

[0030] Further, the calculation steps of the global feature f glb in step 2-4 are as follows:

[0031] Step 2-4-1, key-value-query pair generation, input the key point descriptor f key into the key network F k to obtain the key The key point descriptor f key The input value network F v The obtained value The key point descriptor f key The input query network generates a query

[0032] Step 2-4-2, attention calculation, according to the key K ey , the query Q eury Calculate the attention value matrix between each pair of key points in the key point P key

[0033] Step 2-4-3, for each key point in the key point P key , feature weighting is performed to obtain global feature f glb and the key point index I glb corresponding to the global feature.

[0034] Further, step 2-4-2 includes: according to the key K ey , the query Q eury Calculate the attention value matrix between each pair of key points in the key point P key Where A[i1,i2] represents the correlation coefficient between the i1th key point and the i2th key point, i1,i2∈P key ; for the relationship calculation between the i1th key point and the i2th key point, the vector inner product of the query Q eury [i1] of the i1th key point and the key K ey [i2] of the i2th key point is obtained to obtain the relationship coefficient A[i1,i2] = Q eury [i1]·K ey [i2]; similarly, for the relationship calculation between the i2th key point and the i1th key point, the vector inner product of the query Q eury [i2] of the i2th key point and the key K ey [i1] of the i1th key point is obtained to obtain the relationship coefficient A[i2,i1] = Q eury [i2]·K ey [i1]; for all pairs of points in the key point P key , the above calculation is performed to obtain the attention value matrix A;

[0035] Step 2-4-3 includes: for the i1th key point, the global feature f is calculated as follows:

[0036]

[0037] For the key point P​​key Perform the above operations at each key point to obtain global features.

[0038] Calculating global features is beneficial for extracting global contextual information from point clouds, while the attention mechanism, as an important step in the Transformer network, can assign different levels of attention to different points, thereby improving the representational power of the encoding network.

[0039] Furthermore, step 3 includes the following steps:

[0040] Step 3-1: Calculate the correlation between local features and the relevant global features of the local features to obtain the correlation matrix Col;

[0041] Step 3-2, Build the weight network F w The weighted network F w It consists of 5 layers. The first four layers each consist of a fully connected layer, a batch normalization layer, and a ReLU layer. The input channel of the first layer is c2, and the output channel is int(c2 / 8), where int(·) is the floor function. The input and output channels of the second layer are both int(c2 / 8). The input channel of the third layer is int(c2 / 8), and the output channel is c2. The input and output channels of the fourth layer are both c2. The last layer is a softmax layer.

[0042] Step 3-3: Input the correlation matrix Col into the weight network F w Obtain feature weights

[0043] Steps 3-4: Construct the feature fusion network F fuse The feature fusion network F fuse It consists of a fully connected layer, a batch normalization layer, and a ReLU layer, with the number of input feature channels of the fully connected layer being c1+c2 and the number of output feature channels being c3.

[0044] Steps 3-5: Global-local feature input feature fusion network F fuse Perform fusion to obtain fusion characteristics;

[0045] Steps 3-6 involve inputting the fused features into the Transformer decoding network to obtain the final point-by-point category labels, thus completing the point cloud semantic segmentation.

[0046] Integrating global and local features can better combine global and local contexts, improve the representation ability of features, and enable the decoding network to more accurately infer the category label of the point.

[0047] Further, the calculating feature correlation of step 3-1 comprises the following steps:

[0048] Step 3-1-1, generating local points P loc from the local point index I loc = P[I loc ] according to the global point index I glb ; glb = P[I glb ];

[0049] Step 3-1-2, for each local point in P loc , querying the nearest global point from P glb , obtaining a local-global point index I l2g ;

[0050] Step 3-1-3, extracting relevant global features from global features f glb according to the local-global point index I l2g ;

[0051] Step 3-1-4, calculating the correlation between local features and relevant global features, obtaining a correlation matrix Col:

[0052]

[0053] wherein, 1≤i≤m1,1≤j≤c2.

[0054] This step provides a basis for the combination of global features and local features, and lays a foundation for the fusion of the two in the fusion network below.

[0055] Further, step 3-5 comprises the following steps:

[0056] Step 3-5-1, element-wise multiplying the relevant global features and the feature weight W, obtaining weighted relevant global features

[0057] Step 3-5-2, concatenating the weighted relevant global features and the local features f loc along the channel to obtain a concatenated feature

[0058] Step 3-5-3, inputting the concatenated feature f cat into the feature fusion network F fuse to obtain the final fusion feature c3 represents the number of output feature channels.

[0059] This step embodies the organic integration of local features and global features, and on the basis of ensuring the feature representation of both, a unified new feature is obtained, which can enable the subsequent decoding network to effectively output the style result.

[0060] Beneficial effects: the context-driven large-scale garden scene point cloud semantic segmentation method provided by the application uses a Transformer to extract local and global context information in the local and global of the original point cloud respectively, and finally supplements the global context information of the key points to the local context information of the original point cloud to improve the point-by-point representation ability, so that the context information can be fully represented on the point features, thereby improving the segmentation accuracy. The extraction of global features in the application is performed on key points rather than point by point, which greatly reduces the computational overhead while ensuring the completeness of point representation information. Secondly, the Transformer is used as a global feature extractor instead of using average pooling to extract global features, which makes the feature expression ability of the method stronger, and can effectively focus on key features using the attention mechanism, which greatly enhances the network's understanding ability of the point cloud, and can more efficiently extract the context semantic information related to the segmentation task in the point cloud, thereby improving the style ability of the network for large-scale garden scene point cloud. BRIEF DESCRIPTION OF DRAWINGS

[0061] The above and / or other aspects of the present application will become apparent and more readily appreciated from the following description, taken in conjunction with the accompanying drawings, in which:

[0062] Figure 1 is a schematic diagram of the processing flow of the application.

[0063] Figure 2 is a visualization result of the input point cloud.

[0064] Figure 3 is a visualization result of the key points obtained by sampling the farthest points.

[0065] Figure 4 is a visualization result of the final semantic segmentation result. DETAILED DESCRIPTION

[0066] Embodiments of the application will be described below with reference to the accompanying drawings.

[0067] As shown in Figure 1 , the context-driven large-scale garden scene point cloud semantic segmentation method disclosed by the embodiment of the application includes the following steps:

[0068] Step 1, Multi-scale spatial context information extraction: Extract point cloud position encoding, and aggregate local spatial information according to 8-neighbor, 16-neighbor and 32-neighbor of point cloud, finally extract point-wise multi-scale spatial context information through feature concatenation.

[0069] Step 1 includes the following steps:

[0070] Step 1.1, Point cloud position encoding, for an input garden point cloud As shown in Figure 2 , first use the position encoding function F pe to obtain the position encoding feature of each point , where c is the dimension of the position encoding feature. The position encoding function is a fully connected layer, where the input channel is 6 and the output channel is c, followed by a batch normalization layer and a ReLU activation function. In actual operation, c = 8.

[0071] Step 1.2, Point cloud multi-scale neighborhood query to obtain point cloud 8-neighbor encoding 16-neighbor encoding and 32-neighbor encoding includes the following steps:

[0072] Step 1.2.1, Construct a KD tree for the input garden point cloud. The method of constructing the KD tree uses the method described in document 13: Buitinck L, Louppe G, Blondel M, et al. API design for machine learning software: experiences from the scikit-learn project [J]. arXiv preprint arXiv: 1309.0238, 2013.

[0073] Step 1.2.2, Query the 8-neighbor points, 16-neighbor points and 32-neighbor points index of the point cloud according to the KD tree

[0074] Step 1.2.3, Query the position encoding feature according to the neighbor point index to obtain 8-neighbor encoding, 16-neighbor encoding and 32-neighbor encoding, the calculation method is as follows:

[0075]

[0076] f pe [I k8 ] represents the value of f k8 in f pe , f pe [I k16 ], fpe [I k32 ]Similarly.

[0077] Step 1.3, local information aggregation. Average pooling operation is performed on the point cloud 8-neighbor encoding, 16-neighbor encoding and 32-neighbor encoding respectively to obtain 8-neighbor information 16-neighbor information 32-neighbor information The calculation method is as follows:

[0078]

[0079] Wherein represents the i-th value in the second dimension of , Similarly.

[0080] Step 1.4, feature splicing. The point-wise position encoding feature f pe , 8-neighbor information 16-neighbor information and 32-neighbor information are spliced in the second dimension to obtain point cloud multi-scale spatial context information

[0081] Step 2, global-local context extraction: using a local Transformer encoding network to extract the local context of the point cloud, and using a global Transformer encoding network to extract the global context of the point cloud, including the following steps:

[0082] Step 2.1, building a local Transformer encoding network The local Transformer encoding network adopts the encoding network described in document 5: Guo M H, Cai J X, Liu Z N, et al. Pct: Point cloud transformer [J]. Computational Visual Media, 2021, 7: 187-199.

[0083] Step 2.2, building a global Transformer encoding network The global Transformer encoding network includes a key network F k , a value network F v and a query network F q . The key network F k includes a fully connected layer, and the input channel of the fully connected layer is 4c and the output channel is c2. The value network F vIt includes a fully connected layer with 4c input channels and c2 output channels. The query network F q It contains a fully connected layer with 4c input channels and c2 output channels. In actual operation, c2 = 128.

[0084] Step 2.3: Use the farthest point sampling algorithm to extract key points R from the garden point cloud. key and key point descriptor f key It includes the following steps:

[0085] Step 2.3.1: Use the farthest point sampling algorithm to extract the keypoint indices of the point cloud and add them to the keypoint index set I of the point cloud. S The execution process of the farthest point sampling algorithm is as follows: First, a point is selected from the point cloud set X as a seed point, and this point is added to the seed point set S, and the index value of this point is added to the point cloud key point index set I. S Then, calculate the distance between all points in the point cloud and all points in the seed point set S, and select the point with the largest distance. This process can be represented as follows:

[0086] I x =argmax i min j ||x i -x j ||,stx i ∈X,x j ∈S

[0087] Will I x Add the seed point to the set S, and simultaneously add the index of that point to the key point index set I of the point cloud. S Repeat the above process until the user-defined number of sampling points is met. Using the garden point cloud P as the point cloud set X, performing the above steps will yield the point cloud keypoint index set I. S

[0088] Step 2.3.2, based on the point cloud keypoint index set I S Extracting key points P from garden point cloud data key =P[I S ],like Figure 3 As shown. Based on the point cloud keypoint index set I S From point cloud multi-scale spatial context information f mssca Extracting key point descriptors f from garden point clouds key =f mssca [I S ].

[0089] Step 2.4, transfer the multi-scale spatial context information f of the point cloud. msscaThe local feature is obtained by inputting the local part of the Transformer encoding network and the local point index I corresponding to the local feature loc The global feature is obtained by inputting the key point descriptor into the global part of the Transformer encoding network and the key point index I corresponding to the global feature glb . Where m1 is the number of local features, and c1 is the dimension of the local feature. m2 is the number of global features, and c2 is the dimension of the global feature. In actual operation m2 = 256, c1 = 512, c2 = 128. int(·) represents the down rounding operation.

[0090] Step 2.4 Calculation of global feature f glb The steps are as follows:

[0091] Step 2.4.1, key-value-query pair generation, input the key point descriptor f key into the key network F k to obtain the key Input the key point descriptor f key into the value network F v to obtain the value Input the key point descriptor f key into the query network to generate the query Where c2 is the dimension of the key, value, and query.

[0092] Step 2.4.2, attention calculation, according to the key K ey , the query Q eury Calculate the attention value matrix between the two key points Where A[i1,i2] represents the correlation coefficient between the ith1 key point and the ith2 key point, i1,i2∈P key . For the relationship between the ith1 key point and the ith2 key point, the vector inner product of the query Q eury [i1] of the ith1 key point and the key K ey [i2] of the ith2 key point is taken to obtain the relationship coefficient A[i1,i2] = Q eury [i1]·K ey [i2]. Similarly, for the relationship between the ith2 key point and the ith1 key point, the vector inner product of the query Q eury [i2] of the ith2 key point and the key K ey [i1] of the ith1 key point is taken to obtain the relationship coefficient A[i2,i1] = Q eury [i2]·K ey [i1]. By analogy. For the key point P keyfor all pairs of points in P, perform the steps in step 2.4.2 to get the attention value matrix A.

[0093] Step 2.4.3, feature weighting, for the ithkey point, its global feature is calculated as follows,

[0094]

[0095] For each key point, perform the above operations, and stack along the first channel to get the global feature

[0096] Step 3, global-local context fusion and Transformer decoding: fuse the global and local features of the point cloud according to the correlation of the global and local features, and use the Transformer decoder to decode the features to get the point-by-point class label, including the following steps:

[0097] Step 3.1, calculate the correlation between the local feature and the relevant global feature. According to the global feature and the local feature, calculate the correlation between the two, and get the correlation matrix and the relevant global feature including the following steps:

[0098] Step 3.1.1, generate local point P loc from the global point index I loc , generate global point P glb from the global point index I glb = P[I glb ].

[0099] Step 3.1.2, for each local point in P loc , query the nearest global point from P glb to get the local-global point index I l2g .

[0100] Step 3.1.3, according to the local-global point index I l2g extract the relevant global feature glb from the global feature f

[0101] Step 3.1.4, calculate the correlation between the local feature and the relevant global feature to get the correlation matrix Col. The calculation method is as follows

[0102]

[0103] where 1≤i≤m1,1≤j≤c2.

[0104] Step 3.2, building the weight network F w The weight network F w contains 5 layers, and the first four layers each consist of a fully connected layer, a batch normalization layer, and a ReLU layer. Among them, the input channel of the first layer is c2, and the output channel is int(c2 / 8), where int(·) is the floor operation. The input and output channels of the second layer are both int(c2 / 8). The input channel of the third layer is int(c2 / 8), and the output channel is c2. The input and output channels of the fourth layer are both c2. The last layer is a softmax layer.

[0105] Step 3.3, input the correlation matrix Col into the weight network F w to obtain the feature weight

[0106] Step 3.4, building the feature fusion network F fuse The feature fusion network F fuse contains a fully connected layer, a batch normalization layer, and a ReLU layer in turn. Among them, the input feature channel number of the fully connected layer is c1+c2, and the output feature channel number is c3. In actual operation, c3=256.

[0107] Step 3.5, global-local feature fusion. According to the local feature f loc , the relevant global feature feature weight W, and the feature fusion network F fuse calculate the final fusion feature contains the following steps:

[0108] Step 3.5.1, multiply the relevant global feature and the feature weight W element by element to obtain the weighted relevant global feature

[0109]

[0110] Among them, is the element-wise multiplication.

[0111] Step 3.5.2, concatenate the weighted relevant global feature and the local feature f loc along the channel to obtain the spliced feature

[0112] Step 3.5.3, input the spliced feature into the feature fusion network F fuse to obtain the final fusion feature

[0113] Step 3.6, input the fusion feature f fuseThe input Transformer decoding network obtains the final point-by-point class label to complete the point cloud semantic segmentation, as shown in Figure 4

[0114] In specific implementations, the present application provides a computer storage medium and a corresponding data processing unit, wherein the computer storage medium can store a computer program, and the computer program can run the invention content of a context-driven large-scale garden scene point cloud semantic segmentation method and some or all steps in each embodiment of the present application when executed by the data processing unit. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), or the like.

[0115] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present application can be realized by means of a computer program and its corresponding general hardware platform. Based on such understanding, the technical solutions in the embodiments of the present application can be embodied in the form of a computer program, i.e., a software product, which can be stored in a storage medium and includes a plurality of instructions for causing a device (which can be a personal computer, a server, a single-chip microcomputer, a MUU, or a network device, etc.) containing a data processing unit to execute the method described in each embodiment or some parts of the embodiments of the present application.

[0116] The present application provides a context-driven large-scale garden scene point cloud semantic segmentation method. There are many methods and ways to implement this technical solution. The above description is only a specific implementation of the present application. It should be noted that for ordinary technical personnel in this technical field, without departing from the principles of the present application, several improvements and refinements can be made, which should also be considered as the protection scope of the present application. The components not explicitly described in the embodiments can be implemented using existing technology.​

Claims

1. A context-driven large-scale garden scene point cloud semantic segmentation method, characterized in that, The method comprises the following steps: Step 1, multi-scale spatial context information extraction: extracting point cloud position encoding point by point, and aggregating local spatial information according to the near neighbor points of the point cloud, finally extracting point-by-point multi-scale spatial context information through feature splicing; Step 2, global-local context extraction: using a local Transformer encoding network to extract local features of the point cloud, and using a global Transformer encoding network to extract global features of the point cloud; Step 3, global-local context fusion and Transformer decoding: fusing the global and local features of the point cloud according to the correlation of the global features and the local features, and using a Transformer decoder to decode the features to obtain point-by-point class labels, completing point cloud semantic segmentation; Step 1 comprises the following steps: Step 1-1, position encoding is performed on the input garden point cloud to obtain a point-by-point position encoding feature, denoted as , p , where N represents the number of input garden point clouds, and the position encoding feature is , c represents the dimension of the position encoding feature of the point cloud. Step 1-2, the near neighbor points include 8 near neighbors, 16 near neighbors and 32 near neighbors, and point cloud multi-scale neighborhood queries are used to obtain point cloud 8 near neighbor encoding , 16 near neighbor encoding and 32 near neighbor encoding ; Step 1-3, local information aggregation, respectively encode the point cloud 8 neighbors , 16 neighbor encoding , 32 neighbor encoding Average pooling operation is performed to obtain 8 neighbor information ; Step 1-4, feature stitching, point-wise position encoding features , 8-neighbor information , 16-neighbor information , and 32-neighbor information Stitching in the second dimension, getting point cloud multi-scale spatial context information ; Step 2 comprises the following steps: Step 2-1, building local Transformer encoding network ; Step 2-2, building a global Transformer encoding network , the global Transformer encoding network comprising a key network , a value network , and a query network , the key network F k , the value network F v , and the query network F q each comprising a fully connected layer, the input channel of the fully connected layer being , the output channel being , and c2 representing the dimension of the global feature; Step 2-3, extract the key points of the garden point cloud using the farthest point sampling algorithm and key point descriptors ; Step 2-4, obtaining the point cloud multi-scale spatial context information f mssca Input local Transformer encoding network Obtaining local features Wherein m1 represents the number of local features, and c1 represents the dimension of the local features; obtaining the key point descriptor f key Input global Transformer encoding network Obtaining global features And the key point index corresponding to the global features Wherein m2 represents the number of global features.

2. The context-driven semantic segmentation method for large-scale garden scene point cloud according to claim 1, characterized in that, The point cloud multi-scale neighborhood query in step 1-2 comprises the following steps: Step 1-2-1, constructing a KD tree for the input garden point cloud; Step 1-2-2, Querying 8-neighbor, 16-neighbor and 32-neighbor point indices of point cloud according to KD-tree ; Step 1-2-3, query location encoding features according to the index of the nearest points get 8 nearest neighbor encodings , 16 nearest neighbor encodings , and 32 nearest neighbor encodings .

3. The context-driven semantic segmentation method for large-scale garden scene point cloud according to claim 2, characterized in that, Step 2-3 comprises the following steps: Step 2-3-1, extract the point cloud key point index set using the farthest point sampling algorithm ; Step 2-3-2, extracting the key points of the garden point cloud according to the point cloud key point index set extracting the key points of the garden point cloud from the garden point cloud ; according to the point cloud key point index set from the point cloud multi-scale spatial context information extracting the key point descriptor of the garden point cloud .

4. The context-driven semantic segmentation method of large-scale garden scene point cloud according to claim 3, characterized in that, Step 2 - Global feature f glb The calculation of the step 2 - global feature f is as follows: Step 2 - 4 - 1, key - value - query pair generation, keypoint descriptor input key network get key keypoint descriptor input value network get value keypoint descriptor input query network generate query ; Step 2 - 4 - 2, attention computation, according to key K ey , query Q eury Compute key points Attention value matrix between each pair of key points in the middle ; Step 2-4-3, for each key point in the key points , perform feature weighting to obtain global features and key point indexes corresponding to the global features .

5. The context-driven semantic segmentation method of large-scale garden scene point cloud according to claim 4, characterized in that, Step 2-4-2 includes: based on key K ey Query Q eury Calculate key points Attention value matrix between pairs of key points ,in Indicates the first The key point and the first The correlation coefficient between key points For the first The key point and the first Calculate the relationship between the key points, and the first key point Querying key points and the Key points The relationship coefficients are obtained by performing a vector dot product. Similarly, for the first The key point and the first Calculate the relationship between the key points, and then... Querying key points and the Key points The relationship coefficients are obtained by performing a vector dot product. For key points The above calculation is performed on all pairwise points to obtain the attention value matrix A; Step 2-4-3 includes: for the first key point, its global feature is calculated as follows: , On the key points The above operation is performed on each key point to obtain global features .

6. The context-driven semantic segmentation method of large-scale garden scene point cloud according to claim 5, characterized in that, Step 3 comprises the following steps: Step 3-1, calculating the correlation of the local features and the relevant global features of the local features to obtain a correlation matrix Col; Step 3-2, building the weight network , the weight network contains 5 layers, the first four layers each consist of a fully connected layer, a batch normalization layer and a ReLU layer, wherein the input channels of the first layer are , the output channels are , wherein is a floor operation; the input channels and the output channels of the second layer are both ; the input channels of the third layer are , the output channels are ; the input and output channels of the fourth layer are both ; the last layer is a softmax layer; Step 3-3, input the correlation matrix Col into the weight network F w get the feature weights ; Step 3-4, build a feature fusion network , the feature fusion network sequentially comprises a full connection layer, a batch normalization layer and a ReLU layer, wherein the input feature channel number of the full connection layer is , and the output feature channel number is ; Step 3-5, global-local feature input feature fusion network fusion is performed to obtain a fused feature; Step 3-6, inputting the fused features into a Transformer decoding network to obtain final point-by-point class labels, completing point cloud semantic segmentation.

7. The context-driven semantic segmentation method of large-scale garden scene point cloud according to claim 6, characterized in that, The calculation of the feature correlation in step 3-1 comprises the following steps: Step 3-1-1, generating local points according to local point index I loc generating local points , generating global points according to global point index I glb generating global points ; Step 3-1-2, for each local point in , query the nearest global point from , get local-global point index ; Step 3-1-3, according to the local-global point index I l2g From the global feature f glb Extract the relevant global feature ; Step 3-1-4, calculate the correlation between the local features and the relevant global features to obtain a correlation matrix : , wherein .

8. The context-driven semantic segmentation method of large-scale garden scene point cloud according to claim 7, characterized in that, Step 3-5 comprises the following steps: Step 3-5-1, multiply the relevant global feature and the feature weight W element by element to obtain the weighted relevant global feature ; Step 3-5-2, concatenate the weighted relevant global features and local features along the channel to get the concatenated features ; Step 3-5-3, concatenating features input feature fusion network to obtain final fused features , c3 represents the number of output feature channels.

Citation Information

Patent Citations

  • A semantic segmentation method for scene point clouds

    CN109410307B

  • Large-scene point cloud semantic segmentation method

    CN112819833A

  • Lane line detection method based on key point regression and multi-scale feature fusion

    CN113627228A

  • Semantic segmentation method for three-dimensional point cloud data

    CN113989504A