Point cloud sequence target identification method based on PointNet and LSTM
Through the combination of GCN sampling and PointNet-LSTM, the problem of neglecting time sequence information between point clouds is solved, and efficient and accurate point cloud target recognition is achieved, which is suitable for autonomous driving systems.
Patent Information
- Application Number
- CN202510434279.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, the inter-frame timing information of point clouds is ignored, and the random sampling or FPS sampling methods are complex and uneven in calculations, resulting in low point cloud processing efficiency and difficult to accurately identify the target.
The improved graph convolution network GCN sampling method is adopted, combined with PointNet and LSTM, and the local and global features of point clouds are extracted through graph structure modeling and feature aggregation, the timing continuity between point clouds is captured, and the deep learning network model is constructed for target recognition.
It improves the accuracy and efficiency of point cloud sequence target recognition, can identify surrounding targets in real time in complex scenarios, and improves the safety and decision-making capabilities of the autonomous driving system.
Smart Images

Figure CN120355991A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of millimeter-wave radar, and particularly to a method for identifying target point cloud sequences based on PointNet and LSTM. Background Art
[0002] Compared with traditional radar systems, 4D millimeter-wave radar can provide height information (z) of the measured target in addition to three-dimensional information (x, y, v), that is, the motion information of the fourth dimension; through the information of this additional dimension, 4D radar can detect and identify surrounding targets more accurately; 4D millimeter-wave radar usually performs target recognition by processing point cloud data; with the development of deep learning technology, more and more deep learning-based models are applied to 4D millimeter-wave radar target recognition because they can automatically learn features from radar data and improve the recognition accuracy.
[0003] However, traditional point cloud processing methods usually regard point clouds as unordered sets, ignoring the temporal information between point cloud frames; and due to the random sampling or FPS sampling (Farthest Point Sampling) of point clouds, random sampling may lead to uneven spatial distribution of the selected points and cannot fully represent the overall structure of the point cloud, while the computational complexity of FPS sampling is relatively high because the distance between all unselected points and the selected point set needs to be calculated each time a point is selected, especially when dealing with large-scale point cloud data, resulting in performance bottlenecks. Summary of the Invention
[0004] In view of this, the present invention provides a method for identifying target point cloud sequences based on PointNet and LSTM, which effectively solves the problems in the prior art that point clouds are regarded as unordered sets, ignoring the temporal information between point cloud frames and the limitations of point cloud sampling methods, and effectively improves the accuracy and efficiency of target recognition of point cloud sequences.
[0005] To achieve the above object, a method for identifying target point cloud sequences based on PointNet and LSTM of the present invention includes the following steps:
[0006] S1. The millimeter-wave radar collects point cloud data to obtain a point cloud data set P, and preprocesses the point cloud data set P to obtain a new point cloud data set P', and divides the preprocessed point cloud data set into a training set, a validation set, and a test set according to a ratio of 8:1:1;
[0007] S2. Use an improved graph convolutional network GCN sampling module to sample the point cloud data p in the training set i to obtain local features, and obtain a new point cloud data set P'';
[0008] S201. Construct a graph structure, taking each point cloud p in the training set i as a node in the graph, connecting it to neighbor nodes through edges, calculating the mean and standard deviation of the distances of each node in the local K-nearest neighbors to obtain a dynamic distance threshold, and constructing an adjacency matrix A using the dynamic threshold. The expression of the adjacency matrix A is:
[0009]
[0010] where A ij represents the adjacency weight between point cloud p i and point cloud p j , d(p i , p j ) represents the Euclidean distance between point cloud p i and point cloud p j , τ i represents the dynamic distance threshold. When d(p i , p j ) is much smaller than τ i , A ij is close to 1. When d(p i , p j ) is much larger than τ i , A ij approaches 0;
[0011] S202. Perform graph convolution processing to aggregate the information of neighbor nodes and extract local features through the graph convolution layer;
[0012] S203. Select the point with the largest amount of information from the obtained local features as the sampling point and perform sampling to obtain a new point cloud data set P”,
[0013] perform weighted summation on the feature vectors of each point cloud to obtain a feature importance score, and select the point with the highest feature importance score as the sampling point;
[0014] S3. Obtain the global feature g of the point cloud data p i through the PointNet feature extraction module;
[0015] S4. Capture the temporal continuity between point cloud frames through the long short-term memory network LSTM temporal modeling module;
[0016] S5. Use the preprocessed point cloud data training set to train a deep learning network model composed of a sampling module, a feature extraction module, and a temporal modeling module until the model converges, then input the data to be measured, and output the final recognition result through the recognition module.
[0017] Preferably, the point cloud data set P has the expression:
[0018] P = {p1, p2, …, pN}
[0019] Among them, p i represents the i-th point cloud data, and N represents the number of original point clouds;
[0020] The point cloud data p i has the following expression:
[0021] p i =(x i , y i , z i , I i , v i ), i = 1, 2,..., N
[0022] Among them, (x i , y i , z i ) represents the 3D coordinates of the i-th point cloud, x i represents the abscissa of the i-th point cloud, y i represents the ordinate of the i-th point cloud, z i represents the height information of the i-th point cloud, I i represents the intensity information of the i-th point cloud, and v i represents the velocity information of the i-th point cloud.
[0023] Preferably, the preprocessing includes the following steps:
[0024] S101. Filter the point clouds with weak energy through a preset intensity threshold δ;
[0025] S102. Use the DBSCN clustering method to filter the discrete point clouds to obtain a new point cloud data set P′, and the expression is:
[0026] P′={p1, p2,..., p M}
[0027] Among them, M represents the number of point clouds after preprocessing.
[0028] Preferably, the expression of the average distance μ i of each point cloud p i as a node in the local K-nearest neighbors is:
[0029]
[0030] The expression of the standard deviation σ i is:
[0031]
[0032] Among them, μ i represents the point cloud p iMean distance among its K nearest neighbors, σ i represents the point cloud p i Standard deviation of distances among its K nearest neighbors;
[0033] Dynamic threshold τ i The expression of is:
[0034] τ i = μ i + ασ i
[0035] where τ i represents the dynamic distance threshold of the point cloud, and α represents an adjustment parameter and is a positive number.
[0036] Preferably, the expression of the graph convolution is:
[0037]
[0038] where represents the adjacency matrix with self-loops added, A represents the degree matrix, E represents the identity matrix, and H (l) represents the node features of the l-th layer, and W (l) represents the learnable weight matrix, represents the normalized degree matrix, and σ represents the non-linear activation function;
[0039] The expression of the new point cloud dataset P” is:
[0040] P” = {p1, p2, …, p K}
[0041] where K represents the number of point clouds after GCN sampling.
[0042] Preferably, the process of the PointNet neural network obtaining the global features of the point cloud data p i includes the following steps:
[0043] S301. Input feature transformation, using the T-Net neural network to perform a spatial transformation on the input point cloud, and the expression is:
[0044] T = f(P”)
[0045] where T represents the transformation matrix, and f represents the spatial transformation function;
[0046] S302. Point-by-point feature extraction, using a multi-layer perceptron MLP to extract features from each point cloud data, and the expression is:
[0047] q i = MLP(p i ) = σ(W3σ(W2σ(W1p i+b1)+b2)+b3), i = 1, 2, …, K
[0048] Among them, q i represents the feature of the i-th point cloud, K represents the number of point clouds after GCN sampling, W1, W2, W3, b1, b2, and b3 represent learning parameters, and σ represents a non-linear activation function;
[0049] S303. Global feature extraction is performed using global max pooling, and the expression is:
[0050] g = max{q1, q2, …, q K}
[0051] Among them, g represents the global feature of the point cloud.
[0052] Preferably, the recognition module includes an MLP layer, a dropout layer, and a Softmax layer;
[0053] The fully connected layer of the MLP layer maps the final hidden state h of the LSTM module t to the category space and then connects to the Softmax layer to calculate the probability of each category The expression is:
[0054]
[0055] Among them, n represents the number of categories, exp(·) represents the exponential function, and z represents the category.
[0056] Preferably, the deep learning network model is trained and verified using a training set and a validation set, and the model uses a joint loss function as the classification loss function L total , and the expression is:
[0057] L total = L CE + λ1L reg + λ2L seg + λ3L spatial
[0058]
[0059] Among them, L CE represents the cross-entropy loss, C represents the total number of categories, y c represents the true label, taking the value 1 for the true category and 0 for other categories, represents the predicted probability output by the model for category c; L reg represents the SNT regularization loss, T represents the affine transformation matrix predicted by the SNT module, E represents the identity matrix, and ||·|| Fdenotes the Frobenius norm; L reg denotes the temporal consistency loss, g t denotes the global spatial feature extracted from the t-th frame by PointNet, h t denotes the temporal feature extracted from the t-th frame by LSTM, f(·) denotes the mapping function, ||·||2 denotes the L2 norm; L spatial denotes the spatial reconstruction loss, x i denotes the coordinate of the i-th point cloud in the original point cloud data, g denotes the global spatial feature extracted by PointNet, G(·) denotes the inverse mapping function.
[0060] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0061] The present invention adopts the GCN graph convolution sampling method. Through graph structure modeling and feature aggregation, the local geometric structure information of the point cloud data is retained. PointNet is used to extract the global geometric features of the point cloud, and LSTM is used to capture the temporal continuity between point cloud frames. Through the synergistic effect of these modules, the geometric information of the point cloud and the temporal information between point cloud frames are considered. When processing large-scale and dynamically changing point cloud data, it has higher computational efficiency and stronger global information capture ability, improves the sampling efficiency and quality, can more accurately identify targets, and enhances the accuracy and efficiency of identification; especially in complex autonomous driving scenarios, the present invention can help vehicles to identify various surrounding targets in real time and make corresponding decisions to ensure driving safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is a schematic diagram of the module structure of the present invention;
[0063] Figure 2 is a flow chart of the improved dynamic distance threshold adjustment GCN sampling of the present invention;
[0064] Figure 3 is a schematic diagram of the actual application scenario in the embodiment of the present invention;
[0065] Figure 4 is a schematic diagram of the LSTM temporal modeling module structure of the present invention;
[0066] Figure 5 is a flow chart of the calculation of the temporal-spatial joint loss function of the present invention;
[0067] Figure 6 is a schematic diagram of the connection and data flow of each module for extracting n frames of point clouds in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0068] To further illustrate the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, describe in detail the specific implementation manners, structures, features and their effects of the present invention as follows.
[0069] As Figure 3 shown in the scenario, a 4D radar is installed on the front side of the leftmost vehicle in the figure to achieve autonomous driving. The autonomous driving vehicle needs to identify various surrounding targets (including other vehicles, pedestrians, obstacles, etc.) in real time and make corresponding decisions; target recognition is of great significance in the autonomous driving system: (1) ensuring safety, accurate target recognition, real-time judgment of the position and speed of obstacles, pedestrians or other traffic participants ahead can avoid vehicle collisions; (2) real-time decision support, the autonomous driving vehicle needs to make decisions according to the real-time environment (such as target position, speed, etc.). Through target recognition, the vehicle can perceive the intentions of other traffic participants and take measures such as avoidance and deceleration; (3) adapting to complex scenarios, in complex urban environments (such as multi-lane roads, highways, narrow roads, etc.), target recognition can help autonomous driving vehicles identify various traffic objects (such as motor vehicles, non-motor vehicles, pedestrians, buildings, trees, etc.), enabling the vehicle to adapt to different traffic scenarios.
[0070] In order to improve the accuracy and efficiency of point cloud sequence target recognition, the present invention provides the following technical solution: a method for point cloud sequence target recognition based on PointNet and LSTM, including the following steps:
[0071] S1. The millimeter-wave radar collects point cloud data to obtain a point cloud data set P, and preprocesses the point cloud data set P to obtain a new point cloud data set P'. The preprocessed point cloud data set is divided into a training set, a validation set and a test set according to the ratio of 8:1:1;
[0072] The expression of the point cloud data set P is:
[0073] P = {p1, p2,..., p N}
[0074] where pi i represents the i-th point cloud data, and N represents the number of original point clouds;
[0075] The expression of the point cloud data pi i is:
[0076] pi i = (xi i , yi i , zi i , Ii i , vi i ), i = 1, 2,..., N
[0077] Among them, (x i , y i , z i ) represents the 3D coordinates of the i-th point cloud, where x i represents the abscissa of the i-th point cloud, y i represents the ordinate of the i-th point cloud, and z i represents the height information of the i-th point cloud. I i represents the intensity information of the i-th point cloud, and v i represents the velocity information of the i-th point cloud;
[0078] The preprocessing includes the following steps:
[0079] S101. Filter the point clouds with weak energy through a preset intensity threshold δ;
[0080] S102. Use the DBSCN clustering method to cluster and filter the discrete point clouds of the point cloud to obtain a new point cloud data set P', and the expression is:
[0081] P’ = {p1, p2, …, p M}
[0082] where M represents the number of point clouds after preprocessing.
[0083] In the processing of point cloud data, the selection of the sampling method has an important impact on feature extraction and model performance. Although traditional sampling methods (such as random sampling and Farthest Point Sampling (FPS)) are simple and easy to use, they have limitations in capturing local geometric structures and retaining key information. In contrast, the Graph Convolutional Network (GCN) sampling method can more effectively extract local features of point clouds through graph structure modeling and feature aggregation. The Graph Convolutional Network (GCN) is a deep learning model based on graph-structured data that can effectively process non-Euclidean data (such as point clouds). In the sampling task of point clouds, GCN can adaptively select key points and retain the local geometric structure information of point clouds through graph structure modeling and feature aggregation;
[0084] The core idea of GCN sampling is to aggregate the neighbor information of each point in the point cloud through graph convolution operations, extract local features, and select sampling points according to the importance of the features;
[0085] S2. Use an improved Graph Convolutional Network (GCN) sampling module to sample the point cloud data p i in the training set to obtain local features, and a new point cloud data set P” is obtained. The steps for the Graph Convolutional Network (GCN) to perform sampling processing include:
[0086] S201. Construct a graph structure. Treat each point cloud p in the training set as a node in the graph, connect it to neighboring nodes through edges, calculate the mean and standard deviation of the distances in the local K-nearest neighbors of each node, and obtain a dynamic distance threshold. The expression for the mean distance μ is as follows: i (The specific expressions for μ are not shown here as they are represented by the given tags -
[0088] .) i The expression for the standard deviation σ is as follows:
[0087]
[0088] (The specific expressions for σ are not shown here as they are represented by the given tags -
[0090] .) i The expression for the standard deviation σ is as follows:
[0089]
[0090] Where μ i represents the mean distance of the point cloud p i in its K-nearest neighbors, and σ i represents the standard deviation of the distances of the point cloud p i in its K-nearest neighbors;
[0091] The expression for the dynamic threshold τ i is as follows:
[0092] τ i = μ i + ασ i
[0093] Where τ i represents the dynamic distance threshold of the point cloud, and α represents a tuning parameter and is a positive number;
[0094] Construct an adjacency matrix A using the dynamic threshold. The expression for the adjacency matrix A is as follows:
[0095]
[0096] Where A ij represents the adjacency weight between the point cloud p i and the point cloud p j d(p i , p j ) represents the Euclidean distance between the point cloud p i and the point cloud p j τ i represents the dynamic distance threshold. When d(p i , p j ) is much smaller than τ i , A ij is close to 1. When d(p i , p j ) is much larger than τ i , A ij tends to 0;
[0097] S202. Graph convolution processing. Information of neighboring nodes is aggregated through a graph convolution layer to extract local features. The expression of graph convolution is as follows:
[0098]
[0099] Among them, represents the adjacency matrix with self-loops added, A represents the identity matrix, E represents the degree matrix, and H (l) represents the node features of the l-th layer, and W (l) represents the learnable weight matrix, represents the normalized degree matrix, which is used to balance the node degree differences and optimize the graph convolution process. σ represents the non-linear activation function (such as the ReLU function);
[0100] S203. Select the point with the largest amount of information in the obtained local features as the sampling point and perform sampling to obtain a new point cloud data set P”, and the expression is:
[0101] P” = {p1, p2, …, p K}
[0102] Among them, K represents the number of point clouds after GCN sampling;
[0103] Perform weighted summation on the feature vectors of each point cloud to obtain the feature importance score, and select the point with the highest feature importance score as the sampling point.
[0104] PointNet directly learns point cloud data through a Multiplayer Perception (MLP) and MaxPooling without explicitly constructing a grid or voxelization, so it is suitable for processing millimeter-wave radar data. PointNet consists of parts such as an Input Transform module, feature extraction (MLP), and global feature extraction (Max Pooling);
[0105] S3. Obtain the global feature g of the point cloud data p i through the PointNet feature extraction module, including the following steps:
[0106] S301. Input feature transformation (Input Transform). Use the T-Net neural network to perform a spatial transformation on the input point cloud to ensure that the input data has rotational invariance. The expression is:
[0107] T = f(P”)
[0108] Among them, T represents a 3×3 or 6×6 transformation matrix, and f represents the spatial transformation function;
[0109] S302. Point - by - point feature extraction (MLP layer). Use a multi - layer perceptron MLP to extract features from each point cloud data. The expression is:
[0110] q i = MLP(p i ) = σ(W3σ(W2σ(W1p i + b1)+ b2)+ b3), i = 1, 2, …, K
[0111] where q i represents the feature of the i - th point cloud, K represents the number of point clouds after GCN sampling, W1, W2, W3, b1, b2, b3 represent learning parameters, and σ represents a non - linear activation function;
[0112] S303. Global feature extraction (Max Pooling layer). Use global max - pooling for global feature extraction. The expression is:
[0113] g = max{q1, q2, …, q K}
[0114] where g represents the global feature of the point cloud, containing information about the entire point cloud.
[0115] Long Short - Term Memory (LSTM) is a special type of Recurrent Neural Network (RNN), specifically designed to address the vanishing gradient and exploding gradient problems that traditional RNNs encounter when dealing with long - sequence data. LSTM can effectively capture long - term dependencies in sequence data by introducing "memory cells" and "gating mechanisms"; the core idea of LSTM is to store and update information through "memory cells" and control the flow of information through "gating mechanisms";
[0116] S4. Capture the temporal continuity between point cloud frames through the LSTM temporal modeling module of the long short - term memory network.
[0117] The LSTM neural network includes the following three gates:
[0118] Forgetting gate: Determines the information to be discarded in the memory cell;
[0119] Input gate: Determines the new information to be stored in the memory cell;
[0120] Output gate: Determines the information output from the memory cell;
[0121] The structure diagram of LSTM is as shown in Figure 4 and the core formula is as follows:
[0122] ft = σ(W f [h t-1 , x t + b f )
[0123] i t = σ(W i [h t-1 , x t + b i )
[0124] o t = σ(W o [h t-1 , x t + b o )
[0125]
[0126] h t = o t ⊙ tanh(C t )
[0127] Among them, x t represents the input at the current time t, that is, the global feature g of the PointNet feature extraction module; f t , i t , o t respectively represent the forget gate, input gate, and output gate; W f , W i , W o and W C respectively represent the weight matrices of the forget gate, input gate, output gate, and candidate memory cell state; b f , b i , b o and b C respectively represent the bias terms of the forget gate, input gate, output gate, and candidate memory cell state; C t , respectively represent the memory cell state and candidate memory cell state; h t represents the hidden state; ⊙ represents element-wise multiplication; σ represents the Sigmoid activation function, and tanh represents the hyperbolic tangent activation function;
[0128] Through the LSTM time series modeling module, the long-term dependencies between point cloud frames can be captured, and the last hidden state h t of the LSTM is input to the fully connected layer.
[0129] S5. Use the preprocessed point cloud data training set to train a deep learning network model composed of a sampling module, a feature extraction module, and a temporal modeling module until the model converges. Then input the data to be measured, and pass it through the recognition module to output the final recognition result;
[0130] As Figure 6 Shown in the schematic diagram of the model structure. Assume that n frames of point clouds are selected at a time, and the LSTM neural network is used for temporal modeling. Preprocessing, GCN sampling, and PointNet feature extraction are performed on the point clouds from the k-th frame to the k + n-th frame respectively. The features of these n frames are extracted and input into the LSTM temporal modeling module together, connected to the MLP and Softmax layers to obtain the recognition result. Since there are many model parameters when training a deep neural network, it is easy to overfit in the case of insufficient data volume; Dropout (random inactivation) can play a role in alleviating overfitting in the neural network. Therefore, a Dropout layer is connected after the MLP layer;
[0131] The fully connected layer of the MLP layer maps the final hidden state h of the LSTM module t to the category space and then connects to the Softmax layer to calculate the probability of each category The expression is:
[0132]
[0133] where n represents the number of categories, exp(·) represents the exponential function, and z represents the category.
[0134] The model provided in this embodiment is trained using the Adam optimizer. The training set and the validation set are used to train and validate the deep learning network model. The model uses the joint loss function as the classification loss function L total ; Based on the cross-entropy loss, the temporal consistency loss and the spatial reconstruction loss are added to form a joint loss function, so that the model can optimize the inter-frame temporal continuity and the single-frame spatial feature expression simultaneously during training. The classification loss function L total The expression is:
[0135] L total = L CE + λ1L reg + λ2L seg + λ3L spatial
[0136]
[0137] where L CE represents the cross-entropy loss, C represents the total number of categories (for example, 3, namely motorcycle, car, and person), y cDenotes the true label, taking the value 1 for the true class and 0 for other classes (using one-hot encoding). Denotes the predicted probability of the model for class c; L reg Denotes the SNT regularization loss, T denotes the affine transformation matrix predicted by the SNT module, E denotes the identity matrix, and ||·|| F Denotes the Frobenius norm, which is used to measure the "size" of a matrix; L seg Denotes the temporal consistency loss, g t Denotes the global spatial feature extracted by PointNet for the t-th frame, h t Denotes the temporal feature extracted by LSTM for the t-th frame, and f(·) denotes the mapping function (fully connected layer) used to map the temporal feature h t To the same dimension as the spatial feature g t+1 ||·||2 denotes the L2 norm, which calculates the Euclidean distance between vectors; L spatial Denotes the spatial reconstruction loss, x i Denotes the coordinates of the i-th point cloud in the original point cloud data, g denotes the global spatial feature extracted by PointNet, and G(·) denotes the inverse mapping function used to reconstruct the global feature g into an estimate of the original point cloud data. The spatial reconstruction loss function L spatial Ensures that the global feature F can fully reflect the spatial structure of the original point cloud.
[0138] As described above, it is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed as above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the above-disclosed technical content to obtain equivalent embodiments with equivalent changes, but as long as the technical content of the present invention is not departed from, any brief modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A method for target recognition of point cloud sequences based on PointNet and LSTM, characterized in that It includes the following steps: S1. The millimeter-wave radar collects point cloud data to obtain a point cloud data set P, and preprocesses the point cloud data set P to obtain a new point cloud data set P'. The preprocessed point cloud data set is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1; S2. Use the improved graph convolutional network (GCN) sampling module to sample the point cloud data p in the training set i to obtain local features, resulting in a new point cloud data set P''; S201. Construct a graph structure, and regard each point cloud p in the training set as a node in the graph. Connect it with neighbor nodes through edges, calculate the mean and standard deviation of the distances of each node in the local K-nearest neighbors to obtain a dynamic distance threshold, and construct an adjacency matrix A using the dynamic threshold. The expression of the adjacency matrix A is as follows: i When regarded as a node in the graph, it is connected to neighbor nodes through edges. Calculate the mean and standard deviation of the distances of each node in the local K-nearest neighbors to obtain a dynamic distance threshold, and construct an adjacency matrix A using the dynamic threshold. The expression of the adjacency matrix A is: Among them, A ij represents the adjacency weight between point cloud p i and point cloud p j , and d(p i , p j ) represents the Euclidean distance between point cloud p i and point cloud p j . τ i represents the dynamic distance threshold. When d(p i , p j ) is much smaller than τ i , A ij is close to 1. When d(p i , p j ) is much larger than τ i , A ij approaches 0; S202. Graph convolution processing, aggregating the information of neighbor nodes through a graph convolution layer and extracting local features; S203. The point with the largest amount of information in the obtained local features is used as a sampling point and sampled to obtain a new point cloud data set P''; The feature vectors of each point cloud are weighted and summed to obtain a feature importance score, and the point with the highest feature importance score is selected as the sampling point; S3. Obtain the global feature g of the point cloud data p through the PointNet feature extraction module i ; S4. The long short-term memory network LSTM time series modeling module is used to capture the time series continuity between point cloud frames; S5. Use the preprocessed point cloud data training set to train a deep learning network model composed of a sampling module, a feature extraction module, and a time series modeling module until the model converges, then input the data to be measured, and output the final recognition result through the recognition module.
2. The method for identifying target of point cloud sequence based on PointNet and LSTM according to claim 1, wherein The point cloud data set P, the expression is: P = {p1, p2, …, p N} where p i represents the i-th point cloud data, and N represents the number of the original point clouds; The point cloud data p i has the following expression: p i =(x i ,y i ,z i ,I i ,v i ), i = 1, 2, …, N Among them, (x i , y i , z i ) represents the 3D coordinates of the i-th point cloud, x i represents the abscissa of the i-th point cloud, y i represents the ordinate of the i-th point cloud, z i represents the height information of the i-th point cloud, I i represents the intensity information of the i-th point cloud, v i represents the velocity information of the i-th point cloud.
3. The method for identifying point cloud sequence targets based on PointNet and LSTM according to claim 2, wherein, The preprocessing includes the following steps: S101. Filter the point cloud with weak energy through a preset intensity threshold δ; S102. Use the DBSCN clustering method to filter the discrete point cloud to obtain a new point cloud data set P', and the expression is: P′ = {p1, p2, …, p M} Where M represents the number of point clouds after preprocessing.
4. A method for target recognition of point cloud sequences based on PointNet and LSTM according to claim 1, characterized in that Each point cloud p i As the average distance μ of the node in the local K-nearest neighbors i The expression is as follows: Standard deviation σ i The expression for Among them, μ i represents the average distance of the point cloud p i in its K-nearest neighbors, and σ i represents the standard deviation of the distance of the point cloud p i in its K-nearest neighbors; Dynamic threshold τ i The expression is as follows: τ i = μ i + ασ i Among them, τ i represents the dynamic distance threshold of the point cloud, and α represents the adjustment parameter and is a positive number.
5. A method for recognizing target of point cloud sequence based on PointNet and LSTM according to claim 4, characterized in that, The expression of the graph convolution is: Among them, represents the adjacency matrix with self-loops added, A represents the degree matrix, E represents the identity matrix, and H (l) represents the node features of the l-th layer, and W (l) represents the learnable weight matrix, represents the normalized degree matrix, and σ represents the non-linear activation function; The expression of the new point cloud data set P'' is: P″ = {p1, p2, …, p K} Where K represents the number of point clouds after GCN sampling.
6. The method for identifying point cloud sequence targets based on PointNet and LSTM according to claim 1, characterized in that The process of obtaining the global features of the point cloud data p by the PointNet neural network includes the following steps: i S301. Input feature transformation, use the T-Net neural network to perform a spatial transformation on the input point cloud, and the expression is: T = f(P'') Where T represents the transformation matrix, and f represents the spatial change function; S302. Point-by-point feature extraction, use a multi-layer perceptron MLP to extract features from each point cloud data, and the expression is: q i = MLP(p i ) = σ(W3σ(W2σ(W1p i + b1)+ b2)+ b3), i = 1, 2, …, K where q i represents the feature of the i-th point cloud, K represents the number of point clouds after GCN sampling, W1, W2, W3, b1, b2, b3 represent learning parameters, and σ represents a non-linear activation function; S303. Use global max pooling for global feature extraction, and the expression is: g = max{q1, q2, …, q K} Where g represents the global feature of the point cloud.
7. A method for recognizing target of point cloud sequence based on PointNet and LSTM according to claim 1, characterized in that, The recognition module includes an MLP layer, a dropout layer, and a Softmax layer; The fully connected layer of the MLP layer maps the final hidden state h of the LSTM module t to the class space and then connects to the Softmax layer to calculate the probability of each class The expression is: Where n represents the number of categories, exp(·) represents the exponential function, and z represents the category.
8. A method for target recognition of point cloud sequences based on PointNet and LSTM according to claim 7, characterized in that, The deep learning network model is trained and verified using the training set and the validation set. The model uses the joint loss function as the classification loss function L total , and the expression is: L total = L CE + λ1L reg + λ2L seg + λ3L spatial Among them, L CE represents the cross-entropy loss, C represents the total number of categories, y c represents the true label, taking 1 for the true category and 0 for other categories, represents the predicted probability output by the model for category c; L reg represents the SNT regularization loss, T represents the affine transformation matrix predicted by the SNT module, E represents the identity matrix, ||·|| F represents the Frobenius norm; L seg represents the temporal consistency loss, g t represents the global spatial feature extracted by PointNet for the t-th frame, h t represents the temporal feature extracted by LSTM for the t-th frame, f(·) represents the mapping function, ||·||2 represents the L2 norm; L spatial represents the spatial reconstruction loss, x i represents the coordinates of the i-th point cloud in the original point cloud data, g represents the global spatial feature extracted by PointNet, and G(·) represents the inverse mapping function.