A 3D point cloud segmentation method and system based on enhanced recurrent slicing network

By adopting enhanced cyclic slicing network, local spatial coding and attention pooling technologies in the three-dimensional point cloud segmentation technology, the problem of difficulty in extracting local features in the existing technology is solved, and the effect of three-dimensional point cloud segmentation is significantly improved.

CN114897912BActive Publication Date: 2025-05-13GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210435200.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-24
Publication Date
2025-05-13
Estimated Expiration
2042-04-24

AI Technical Summary

Technical Problem

The existing three-dimensional point cloud segmentation technology is difficult to effectively extract local features, especially in objects with more subtle features (such as windows, bookshelfs, etc.), and the segmentation effect is poor.

Method used

Using a method based on enhanced circular slice network, the significant local spatial features are extracted through local spatial coding technology, and the importance scores of different local characteristics are trained to learn, the local characteristics are weighted and combined, and the fused local characteristics are output.

Benefits of technology

The effect of three-dimensional point cloud segmentation is improved, especially when processing objects with multiple subtle features, segmentation can be performed more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114897912B_ABST
    Figure CN114897912B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional point cloud segmentation method and system based on an enhanced recurrent slicing network, the method comprising: grouping and slicing a point cloud set according to the size of the x, y, and z axis coordinate values; performing local spatial encoding on each slice; obtaining the ordered aggregation features of each slice; inputting the ordered aggregation features of each slice into an RNN network for training, modeling the neighboring relationship, and obtaining the interactive features; performing an anti-pooling decoding operation on the interactive features, and mapping the interactive features to each point through a convolutional layer and local spatial decoding to obtain the segmentation results of all slice branches; calculating the loss function of all slice branches, and continuously updating the weight parameters of the local spatial encoding network, the hidden layer of the RNN, and the attention score of the attention pooling network until the network training is stopped; after the network training of all slice branches is stopped, the segmentation results of each slice branch are aggregated into the final segmentation result. The present invention can enhance the three-dimensional point cloud segmentation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional point cloud segmentation, and in particular to a three-dimensional point cloud segmentation method and system based on an enhanced recurrent slicing network. Background Art

[0002] Point cloud is a collection of points that comes from various 3D sensors, such as 3D scanners, LiDAR, etc. Compared with 2D images, 3D point clouds can provide unique information about spatial distribution and have greater application value. With the rapid development of 3D point cloud acquisition technology, 3D point cloud segmentation technology has become a hot topic in current point cloud research. The task of 3D point cloud segmentation is to classify point clouds in different spatial areas into various object categories.

[0003] The current mainstream 3D point cloud segmentation technology is based on deep learning networks. For example, the input point cloud information (such as coordinates, colors, etc.) is input into the deep neural network, the spatial distribution information of the point cloud is trained and learned, and a category label is attached to each point to complete the point cloud segmentation task. However, the existing 3D point cloud segmentation method uses a multi-layer perceptron (MLP) to train and learn point cloud features, so it is impossible to effectively extract local features. In recent years, some scholars have proposed a point cloud segmentation technology based on a recurrent slicing network. Using the point cloud slicing pooling technology, the features of unordered points are mapped to an ordered feature vector sequence, the local features of the point cloud are extracted, and then input into the recurrent neural network (RNN) to model the neighboring dependency relationship to complete the point cloud segmentation task. However, the existing point cloud segmentation technology based on the recurrent slicing network only uses a multi-layer perceptron to extract local features from the slice when extracting local features from the slice, and does not further encode its neighbor information. Therefore, the extraction of local feature information may not be sufficient. In addition, when the existing technology extracts multiple local features from a slice, it uses a maximum pooling algorithm to obtain only the most significant local features of the slice, without using less significant local features. This results in poor results when performing point cloud segmentation on objects with more subtle features (such as windows, bookshelves, etc.). Summary of the invention

[0004] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a 3D point cloud segmentation method based on an enhanced recurrent slicing network, which uses slicing local space encoding technology to extract sub-significant local space features and enhance the 3D point cloud segmentation effect. In addition, the attention pooling technology is used to train and learn the importance scores of different local features, weightedly combine the importance scores of local features, and output fused local features to further improve the segmentation effect.

[0005] To achieve the above purpose, the technical solution provided by the present invention is:

[0006] A 3D point cloud segmentation method based on an enhanced recurrent slicing network comprises the following steps:

[0007] S1. Input a point cloud set, and group and slice the point cloud set according to the x-, y-, and z-axis coordinate values ​​to generate multiple point sets, one point set being one slice;

[0008] S2, performing local spatial encoding on each slice to obtain local spatial features of each slice;

[0009] S3, input the local features obtained from each slice into the attention pooling module to obtain the ordered aggregation features of each slice;

[0010] S4, input the ordered aggregation features of each slice into the RNN network for training, model the neighbor relationship, and obtain the interaction features;

[0011] S5, performing an unpooling decoding operation on the interactive features, and mapping the interactive features to each point through a convolutional layer and local space decoding to obtain the segmentation results of all slice branches of the x, y, and z groups;

[0012] S6, calculate the loss function of the slice branch of the x, y, z group, and determine whether to stop training the network according to the threshold. If the training is not stopped, update the weight parameters of the local spatial encoding network, the hidden layer of the RNN, and the attention score of the attention pooling network; if the value of the loss function is less than or equal to the threshold, stop training the network;

[0013] S7. After the network training of all slice branches stops, the segmentation results of each slice branch are aggregated into the final segmentation result.

[0014] Furthermore, in step S1, the process of grouping slices includes:

[0015] S1-1, let the point cloud set P = {p1,p2,…,p i ,…,p M The z-axis coordinate value of} is in [z min ,z max ], where M represents the total number of elements in set P, and the slice is divided into , symbol represents the upward integer function, r is the slice resolution;

[0016] S1-2, point p i Assign to slice g, where z i is the z coordinate of the ith point, symbol Represents the downward integer function, where i = 1, 2, …, M.

[0017] Furthermore, the step S2 comprises:

[0018] S2-1. Select a point p in the slice j , find point p j The nearest K neighbors The K value is the set value. Represents point p j The kth neighbor point, k = 1, 2, ..., K;

[0019] S2-2, point p j The K nearest neighbor points of the point are encoded relative to the point position; the relative point position is encoded as a feature vector

[0020]

[0021] where p j and The xyz space coordinate vector representing the point, symbol ||·|| calculates the Euclidean distance between adjacent points and the center point, symbol Indicates feature connection operation, which means and The vector elements of are concatenated in order into the vector p j Behind, vector p j The dimension of becomes longer, MLP(·) means feature extraction using MLP;

[0022] S2-3, Neighbor Point RGB features Refers to the RGB color value, and the features encoded with the relative point position Perform feature connection operation, that is

[0023]

[0024] The symbol Represents the feature connection operation. The feature connection operation here is the same as above, which means that the vector The elements are concatenated in a vector Behind, vector The dimension of becomes longer, and a new vector is obtained

[0025] S2-4. Select other points p in the slice j , repeat steps S2-1 to S2-3 to obtain the local feature representation of the slice Where j = 1, 2, ..., Q, Q is the number of points in the slice.

[0026] Furthermore, the step S3 comprises:

[0027] S3-1. The output slice local feature F is input into the attention pooling module, and the attention score is learned through MLP. Where j = 1, 2, ..., Q, Q represents the number of points in the slice, Represents the feature vector The attention score vector of ;

[0028] S3-2, group the local feature vectors of the slice Perform the dot product of the corresponding vector with the attention score vector group, that is, Obtain a weighted feature vector group;

[0029] S3-3. Sum the weighted feature vectors to obtain the ordered aggregate features of the slices Right now

[0030]

[0031] Furthermore, the step S5 includes

[0032] S5-1, map the interaction features to each local feature vector through the convolution layer;

[0033] S5-2. Input the local feature vector obtained in step S5-1 into the decoding MLP, and output the point cloud segmentation results of all slice branches.

[0034] Further, the step S6 comprises:

[0035] S6-1, calculate the loss function;

[0036] The loss function uses cross entropy, and the loss functions of the x-axis grouping and slicing branch, the y-axis grouping and slicing branch, and the z-axis grouping and slicing branch are defined as:

[0037]

[0038]

[0039]

[0040] in, They represent the predicted probability that point i belongs to category j when grouped and sliced ​​by x-axis, y-axis, and z-axis, respectively; M is the total number of point clouds, C is the total number of categories of point cloud segmentation, and u ij The function is defined as follows:

[0041]

[0042] S6-2. When the loss function Loss x 、Loss y and Loss zWhen the value of is greater than the threshold l, continue to train the corresponding slice branch network, update the weight parameters of the local spatial encoding network, the hidden layer of the RNN, and the attention score of the attention pooling network; when the loss function Loss x 、Loss y and Loss z When the value of is less than or equal to the threshold l, stop training the corresponding axis slice branch network.

[0043] To achieve the above object, the present invention further provides a three-dimensional point cloud segmentation system based on an enhanced recurrent slicing network, which includes a grouping slicing layer, a local space encoding layer, an attention pooling layer, an RNN layer, a de-pooling decoding layer, and an aggregation layer;

[0044] The grouping and slicing layer is used to group and slice the point cloud set according to the size of the x-, y-, and z-axis coordinate values;

[0045] The local space coding layer is used to perform local space coding on each slice to obtain local space features of each slice;

[0046] The attention pooling layer learns the importance scores of different local features, and uses the importance scores to weight and combine local features to obtain ordered aggregate features of each slice;

[0047] The RNN layer is used to train the ordered aggregation features of each input slice, model the neighbor relationship, and obtain the interactive features;

[0048] The de-pooling decoding layer performs de-pooling decoding operations on the interactive features. The interactive features are mapped to each point through the convolution layer and local space decoding to obtain the segmentation results of all slice branches of the x, y, and z groups;

[0049] The loss function calculation layer is used to calculate the loss function of the slice branch grouped by x, y, and z;

[0050] The judgment layer is used to judge whether to stop training the network;

[0051] The aggregation layer is used to aggregate the segmentation results of each slice branch into a final segmentation result.

[0052] Compared with the existing technology, the principle and advantages of this solution are as follows:

[0053] On the basis of the existing slice-loop slice segmentation technology, the local spatial feature encoding technology is added to extract more detailed local spatial features of the point cloud. The attention pooling technology is used to effectively retain different detailed local features, which helps to improve the 3D point cloud segmentation performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the services required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0055] Figure 1 This is a principle flow chart of a three-dimensional point cloud segmentation method based on an enhanced recurrent slicing network of the present invention;

[0056] Figure 2 It is a structural schematic diagram of a three-dimensional point cloud segmentation system based on an enhanced recurrent slicing network of the present invention;

[0057] Figure 3 It is a structural schematic diagram of a local space coding layer in a three-dimensional point cloud segmentation system based on an enhanced recurrent slicing network according to the present invention;

[0058] Figure 4 This is a structural schematic diagram of an attention pooling layer in a three-dimensional point cloud segmentation system based on an enhanced recurrent slicing network of the present invention. DETAILED DESCRIPTION

[0059] The present invention will be further described below in conjunction with specific embodiments:

[0060] The three-dimensional point cloud segmentation method based on the enhanced recurrent slicing network described in this embodiment adopts local space coding technology and attention pooling technology. The original point cloud is separated according to the coordinate values ​​of the x-axis, y-axis, and z-axis respectively. Taking the z-axis as an example, the point cloud is separated according to the coordinate value of the z-axis. After separation, many point sets will be generated, and a point set is called a slice. There is a certain amount of spatial adjacent information between adjacent slices in the same axial direction, which lays the foundation for the subsequent input into the RNN to model the neighboring relationship. Inside the point cloud set of each slice, there are some important local features, such as the Euclidean distance between adjacent points, the coordinates of adjacent points, and the color.

[0061] First, this embodiment performs encoding operations on the local features of the slice through the local space coding layer to extract more detailed local features. Because there are many specific geometric features of the points and there are many points in the slice, the total number of internal features of the slice is large. Secondly, before inputting the local features into the RNN layer for training and learning, the attention pooling technology designed in this embodiment performs pooling operations on the internal features of the slice, extracts the feature representation of each slice, and forms an ordered feature vector sequence. The attention pooling layer of this embodiment can train the importance of different local features inside the slice and weight them for output. The more important the feature, the higher the weight. After passing through the RNN network, a set of predicted feature vectors with unchanged dimensions are output, and finally the predicted features are mapped back to each point through the inverse pooling decoding layer operation to achieve point cloud segmentation.

[0062] like Figure 1 As shown, the principle flow of the specific method is as follows:

[0063] S1. Input point cloud data and slice and separate the point cloud according to the x, y, and z axis coordinate values. Taking the z axis direction as an example, the point cloud set P = {p1, p2, ..., p i ,…,p M}(p i represents a point, and the same applies to others) is sliced ​​into N groups of point sets S in the z-axis direction z ={S1,S2,…,S i ,…,S N},S i It is called the i-th slice, and from step S2 to the end, the slice S in the z-axis direction i As an example (i can be 1 to N, and other slices go through the same steps). Slice separation along the x-axis and y-axis directions to obtain point sets S x and S y , S x and S y There are N groups of slices. The specific process of this step includes:

[0064] S1-1. Point cloud data includes the spatial coordinates (x, y, z) and RGB colors of the points. Taking the z-axis direction as an example, let the point cloud set P = {p1, p2, ..., p i ,…,p M The z-axis coordinate value of} is in [z min ,z max ], where M represents the total number of elements in set P, and the slice is divided into , symbol represents the upward integer function, r is the slice resolution (used to control the number of slices);

[0065] S1-2, point p iAssign to slice g, where z i is the z coordinate of the ith point, symbol represents the downward integer function, where i = 1, 2, ..., M;

[0066] S2, set slice S i There are Q points in total, for slice S i Point p inside j (j=1,2,…,Q) using Figure 3 The local spatial coding layer shown performs spatial coding to obtain local spatial features. This step includes:

[0067] S2-1. Select a point p in Si j , find point p j The nearest K neighbors The K value is set based on experience or requirements. Represents point p j The kth neighbor point (k = 1, 2, ..., K);

[0068] S2-2, point p j The K nearest neighbor points of are encoded relative to the point position. The relative point position is encoded as a feature vector:

[0069]

[0070] where p j and The xyz space coordinate vector representing the point, symbol ||·|| calculates the Euclidean distance between adjacent points and the center point, symbol Indicates feature connection operation, which means and The vector elements of are concatenated in order into the vector p j Behind, vector p j The dimension of becomes longer, MLP(·) means feature extraction using MLP;

[0071] S2-3, Neighbor Point RGB features Refers to the RGB color value, and the features encoded with the relative point position Perform feature connection operation, that is

[0072]

[0073] The symbol Represents the feature connection operation. The feature connection operation here is the same as above, which means that the vector The elements are concatenated in a vector Behind, vector The dimension of becomes longer, and a new vector is obtained

[0074] S2-4, select slice S i Other points p in j (j=1,2,…,Q), repeat steps S2-1 to S2-3 to obtain slice S i Local feature representation Where j = 1, 2, ..., Q;

[0075] S3, the above local features obtained by each slice Input the attention pooling layer to obtain the ordered aggregation features of the slices; this step specifically includes:

[0076] S3-1, attention pooling layer structure diagram as shown Figure 4 As shown. The slice features output from step S2 Input to the attention pooling layer and learn the attention score through MLP where j = 1, 2, ..., Q, Represents the feature vector The attention score vector of

[0077] S3-2, slice S i The local eigenvector group of Perform the dot product of the corresponding vector with the attention score vector group, that is, Obtain a weighted feature vector group;

[0078] S3-3: Sum the weighted eigenvectors to get slice S i Aggregation features Right now

[0079]

[0080] S4, repeat the above step S3 to obtain the ordered aggregation features of different slices And input them into the RNN layer for training, model the neighboring relationship, and obtain the interactive features in Slice S n Aggregation features The corresponding interaction features.

[0081] S5, the interactive features are subjected to an unpooling decoding operation, and the interactive features are mapped to each point through a convolutional layer and local space decoding to obtain the segmentation results of all slice branches of the x, y, and z groups; this step includes:

[0082] S5-1, map the interaction features to each local feature vector through the convolution layer;

[0083] S5-2. Input the local feature vector obtained in step S5-1 into the decoding MLP, and output the point cloud segmentation results of all slice branches.

[0084] S6, calculate the loss function of the slice branch of the x, y, z group, and determine whether to stop training the network according to the threshold. If the training is not stopped, update the weight parameters of the local spatial encoding network, the hidden layer of the RNN, and the attention score of the attention pooling network; if the value of the loss function is less than or equal to the threshold, stop training the network;

[0085] This step includes:

[0086] S6-1, calculate the loss function;

[0087] The loss function uses cross entropy, and the loss functions of the x-axis grouping and slicing branch, the y-axis grouping and slicing branch, and the z-axis grouping and slicing branch are defined as:

[0088]

[0089]

[0090]

[0091] in, They represent the predicted probability that point i belongs to category j when grouped and sliced ​​by x-axis, y-axis, and z-axis, respectively; M is the total number of point clouds, C is the total number of categories of point cloud segmentation, and u ij The function is defined as follows:

[0092]

[0093] S6-2. When the loss function Loss x 、Loss y and Loss z When the value of is greater than the threshold l, continue to train the corresponding slice branch network, update the weight parameters of the local spatial encoding network, the hidden layer of the RNN, and the attention score of the attention pooling network; when the loss function Loss x 、Loss y and Loss z When the value of is less than or equal to the threshold l, stop training the corresponding axis slice branch network.

[0094] S7. After the network training of all slice branches stops, the segmentation results of each slice branch are aggregated into the final segmentation result.

[0095] In addition, this embodiment also provides a 3D point cloud segmentation system based on an enhanced recurrent slicing network, such as Figure 2As shown, it includes a grouping slicing layer, a local space encoding layer, an attention pooling layer, an RNN layer, an anti-pooling decoding layer, and an aggregation layer;

[0096] The grouping and slicing layer is used to group and slice the point cloud set according to the size of the x-, y-, and z-axis coordinate values;

[0097] The local space coding layer is used to perform local space coding on each slice to obtain local space features of each slice;

[0098] The attention pooling layer learns the importance scores of different local features, and uses the importance scores to weight and combine local features to obtain ordered aggregate features of each slice;

[0099] The RNN layer is used to train the ordered aggregation features of each input slice, model the neighbor relationship, and obtain the interactive features;

[0100] The de-pooling decoding layer performs de-pooling decoding operations on the interactive features. The interactive features are mapped to each point through the convolution layer and local space decoding to obtain the segmentation results of all slice branches of the x, y, and z groups;

[0101] The loss function calculation layer is used to calculate the loss function of the slice branch grouped by x, y, and z;

[0102] The judgment layer is used to judge whether to stop training the network;

[0103] The aggregation layer is used to aggregate the segmentation results of each slice branch into a final segmentation result.

[0104] Based on the existing slicing cycle slicing segmentation technology, this embodiment adds local spatial feature encoding technology, which can extract more detailed local spatial features of point clouds. It uses attention pooling technology to effectively retain different detailed local features, which helps to improve the three-dimensional point cloud segmentation performance.

[0105] The embodiments described above are only preferred embodiments of the present invention and are not intended to limit the scope of implementation of the present invention. Therefore, all changes made according to the shape and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A 3D point cloud segmentation method based on enhanced recurrent slicing network, characterized in that: The steps include: S1. Input a point cloud set, and group and slice the point cloud set according to the x-, y-, and z-axis coordinate values ​​to generate multiple point sets, one point set being one slice; S2, perform local spatial encoding on each slice to obtain local spatial features of each slice; the process includes: selecting a point in the slice, finding the K nearest neighbor points of the point, and performing relative point position encoding on the K nearest neighbor points of the point; S3, input the local spatial features obtained for each slice into the attention pooling module to obtain the ordered aggregation features of each slice; S4, input the ordered aggregation features of each slice into the RNN network for training, model the neighbor relationship, and obtain the interaction features; S5, performing an unpooling decoding operation on the interactive features, and mapping the interactive features to each point through a convolutional layer and local space decoding to obtain the segmentation results of all slice branches of the x, y, and z groups; S6, calculate the loss function of the slice branch of the x, y, z group, and determine whether to stop training the network according to the threshold. If the training is not stopped, update the weight parameters of the local spatial encoding network, the hidden layer of the RNN, and the attention score of the attention pooling network; if the value of the loss function is less than or equal to the threshold, stop training the network; S7. After the network training of all slice branches stops, the segmentation results of each slice branch are aggregated into the final segmentation result.

2. The three-dimensional point cloud segmentation method based on enhanced recurrent slicing network according to claim 1 is characterized in that: In step S1, the process of grouping slices includes: S1-1, let the point cloud set P = {p1,p2,…,p i ,…,p M The z-axis coordinate value of} is in [z min ,z max ], where M represents the total number of elements in set P, and the slice is divided into , symbol represents the upward integer function, r is the slice resolution; S1-2, point p i Assign to slice g, where z i is the z coordinate of the ith point, symbol Represents the downward integer function, where i = 1, 2, …, M.

3. The three-dimensional point cloud segmentation method based on enhanced recurrent slicing network according to claim 1 is characterized in that: The step S2 comprises: S2-1. Select a point p in the slice j , find point p j The nearest K neighbors The K value is the set value. Represents point p j The kth neighbor point, k = 1, 2, ..., K; S2-2, point p j The K nearest neighbor points of the point are encoded relative to the point position; the relative point position is encoded as a feature vector where p j and The xyz space coordinate vector representing the point, symbol ||·|| calculates the Euclidean distance between adjacent points and the center point, symbol Indicates feature connection operation, which means and The vector elements of are concatenated in order into the vector p j Behind, vector p j The dimension of becomes longer, MLP(·) means feature extraction using MLP; S2-3, Neighbor Point RGB features Refers to the RGB color value, and the features encoded with the relative point position Perform feature connection operation, that is The symbol Represents the feature connection operation. The feature connection operation here is the same as above, which means that the vector The elements are concatenated in a vector Behind, vector The dimension of becomes longer, and a new vector is obtained S2-4. Select other points p in the slice j , repeat steps S2-1 to S2-3 to obtain the local spatial feature representation of the slice Where j = 1, 2, ..., Q, Q is the number of points in the slice.

4. The three-dimensional point cloud segmentation method based on enhanced recurrent slicing network according to claim 1 is characterized in that: The step S3 comprises: S3-1. The output slice local spatial feature F is input into the attention pooling module, and the attention score is learned through MLP. Where j = 1, 2, ..., Q, Q represents the number of points in the slice, Represents the feature vector The attention score vector of ; S3-2, the local spatial feature vector group of the slice Perform the dot product of the corresponding vector with the attention score vector group, that is, Obtain a weighted feature vector group; S3-3. Sum the weighted feature vectors to obtain the ordered aggregate features of the slices Right now 5. The three-dimensional point cloud segmentation method based on enhanced recurrent slicing network according to claim 1, characterized in that: The step S5 includes S5-1, map the interaction features to each local space feature vector through the convolution layer; S5-2. Input the local spatial feature vector obtained in step S5-1 into the decoding MLP, and output the point cloud segmentation results of all slice branches.

6. The three-dimensional point cloud segmentation method based on enhanced recurrent slicing network according to claim 1, characterized in that: The step S6 comprises: S6-1, calculate the loss function; The loss function uses cross entropy, and the loss functions of the x-axis grouping and slicing branch, the y-axis grouping and slicing branch, and the z-axis grouping and slicing branch are defined as: in, They represent the predicted probability that point i belongs to category j when grouped and sliced ​​by x-axis, y-axis, and z-axis, respectively; M is the total number of point clouds, C is the total number of categories of point cloud segmentation, and u ij The function is defined as follows: S6-2. When the loss function Loss x 、Loss y and Loss z When the value of is greater than the threshold l, continue to train the corresponding slice branch network, update the weight parameters of the local spatial encoding network, the hidden layer of the RNN, and the attention score of the attention pooling network; when the loss function Loss x 、Loss y and Loss z When the value of is less than or equal to the threshold l, stop training the corresponding axis slice branch network.

7. A 3D point cloud segmentation system based on enhanced recurrent slicing network, characterized in that: It includes grouping and slicing layer, local space encoding layer, attention pooling layer, RNN layer, anti-pooling decoding layer, loss function calculation layer, judgment layer, and aggregation layer; The grouping and slicing layer is used to group and slice the point cloud set according to the size of the x-, y-, and z-axis coordinate values; The local space coding layer is used to perform local space coding on each slice to obtain local space features of each slice. The process includes: selecting a point in the slice, finding the K nearest neighbor points of the point, and performing relative point position coding on the K nearest neighbor points of the point; The attention pooling layer is used to learn the importance scores of different local spatial features, and use the importance scores to weight and combine the local spatial features to obtain the ordered aggregation features of each slice; The RNN layer is used to train the ordered aggregation features of each input slice, model the neighbor relationship, and obtain the interactive features; The de-pooling decoding layer performs de-pooling decoding operations on the interactive features. The interactive features are mapped to each point through the convolution layer and local space decoding to obtain the segmentation results of all slice branches of the x, y, and z groups; The loss function calculation layer is used to calculate the loss function of the slice branch grouped by x, y, and z; The judgment layer is used to judge whether to stop training the network; The aggregation layer is used to aggregate the segmentation results of each slice branch into a final segmentation result.

Citation Information

Patent Citations

  • Semantic segmentation method for point cloud data

    CN112257597A

  • Systems and Methods for Semantic Segmentation of 3D Point Clouds

    US20190108639A1