Continuous human motion recognition method based on point cloud data and distance Doppler spectrogram
By combining point cloud data and distance Doppler spectrum, a PVT-BiLSTM network is built, and multimodal data and dual-stream network are used to solve the problem of low accuracy of existing radar continuous human body movement recognition methods, achieving higher recognition accuracy and performance.
Patent Information
- Application Number
- CN202510111827.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The existing radar continuous human body movement recognition method has low recognition accuracy, especially the radar human body movement characterization based on single mode has resolution limitations and sparseness problems.
The continuous human motion recognition method based on point cloud data and distance Doppler spectrum is adopted. By building an FMCW radar data acquisition platform and PVT-BiLSTM network, combined with the Pointnet-BiLSTM module, ViT module and fusion module, signal processing and feature extraction are performed, and multimodal data and dual-stream network are used to improve the recognition accuracy.
The accuracy of continuous human movement recognition is significantly improved. Through the combination of multimodal data and the design of dual-stream network, the resolution and sparsity of single-modal data are solved, and higher recognition performance is achieved.
Smart Images

Figure CN120047998A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of radar target recognition, and particularly relates to a continuous human action recognition method based on point cloud data and range-Doppler spectrogram. Background Art
[0002] Human action recognition technology has important research significance and application prospects in the fields of smart home, security monitoring, human-computer interaction, etc. Currently, sensors that can be used in the field of human action recognition include cameras, wearable devices, radars, etc. Cameras are sensitive to light and environmental conditions, and there is a risk of individual privacy leakage. Wearable devices may cause discomfort and burden to individuals. Compared with other sensors, radar detects and senses targets by emitting electromagnetic waves, and has the advantages of working all-weather, non-contact, not being affected by light conditions, and having no risk of leaking user privacy. Therefore, the human action recognition method based on radar sensors has received more and more attention, and many radar-based human action recognition methods have been proposed one after another. In addition, human actions are usually continuous and dynamic. Compared with discrete human action recognition, continuous human action recognition is closer to the real scenario.
[0003] In recent years, deep learning technology has gradually emerged in the field of continuous human action recognition based on radar. However, most methods still rely on single-modal radar human action representations. For example, range-Doppler features and radar point clouds. Although these methods are practical, the range-Doppler features based on two-dimensional spectrograms have limitations in resolution, and the point cloud data of conventional multiple-input multiple-output radars is sparse, with limited ability to provide human action features. Therefore, there are certain defects in achieving high-precision radar continuous human action recognition.
[0004] Therefore, it has become an urgent problem to propose a new continuous human action recognition method to improve the accuracy of continuous human action recognition. Summary of the Invention
[0005] In view of this, the present invention provides a continuous human action recognition method based on point cloud data and range-Doppler spectrogram to at least solve the problem of low recognition accuracy of existing radar continuous human action recognition methods.
[0006] To achieve the above object, the present invention provides a continuous human action recognition method based on point cloud data and range-Doppler spectrogram, including:
[0007] Step 1: Build an FMCW radar data acquisition platform, and use the platform to collect a series of echo signals containing various continuous human actions to obtain sampled intermediate-frequency signals. Then, preprocess the intermediate-frequency signals to obtain an overall matrix of intermediate-frequency signals. Next, use a sliding window to segment the overall matrix of intermediate-frequency signals to obtain intermediate-frequency signal matrices corresponding to discrete actions. Finally, perform signal processing on the intermediate-frequency signal matrices to obtain corresponding range-Doppler spectrograms and point cloud data;
[0008] Step 2: Build a PVT-BiLSTM network. The PVT-BiLSTM network includes a Pointnet-BiLSTM module, a ViT module, and a fusion module. The Pointnet-BiLSTM module uses two T-net networks, five one-dimensional convolutional layers, two BiLSTM layers with 64 and 128 neurons respectively, one one-dimensional global max pooling layer, and two fully connected layers with 256 and 128 neurons respectively. Each convolutional layer and fully connected layer includes a BN layer and a ReLU activation function. The number of convolutional kernels in the first three one-dimensional convolutional layers is 32 and the size is 1, the number of convolutional kernels in the fourth convolutional layer is 64 and the size is 1, the number of convolutional kernels in the fifth convolutional layer is 512, and the size is 1. The T-net network includes three one-dimensional convolutional layers with 32, 64, and 512 convolutional kernels respectively and a size of 1, one one-dimensional global max pooling layer, two fully connected layers with 256 and 128 neurons respectively, a fully connected layer containing r 2 neurons, and a transformed matrix after reshaping. The size of the transformed matrix of the T-net network is r×r, where r is the dimension value of the transformed matrix. For the two T-net networks, the dimension values of the transformed matrices are 4 and 32 respectively; The ViT module is divided into five processing steps: Patch segmentation, linear projection, class token, positional embedding, Transformer encoder, and multi-layer perceptron;
[0009] Step 3: Divide the range-Doppler spectrograms and point cloud data into a training set and a validation set, and input them into the PVT-BiLSTM network for training and validation to obtain a trained PVT-BiLSTM network model. Then, use the trained PVT-BiLSTM network model to identify continuous human actions.
[0010] Preferably, Step 1 includes the following steps:
[0011] (1) Use the FMCW radar data acquisition platform to measure the echo of a complete continuous human action, and sample the echo of the continuous human action to obtain sampled intermediate-frequency signals. Then, stack the sampled intermediate-frequency signals column by column to form an overall matrix S of intermediate-frequency signals IF of the intermediate-frequency signals, and the overall matrix S of the intermediate-frequency signalsIF Each element is represented as s IF (m, n), where m = 1, 2, ..., M and n = 1, 2, ..., N total , m is the fast time index, n is the slow time index, M is the total number of fast time samplings corresponding to a single Chirp, and N total is the total number of transmitted Chirps;
[0012] (2) Use a fixed window length and a sliding step size to perform sliding window segmentation on the overall matrix S of the intermediate frequency signals IF along the slow time dimension to obtain the intermediate frequency signal matrix X corresponding to each single discrete action i , i = 1, 2, ..., I, where I is the number of discrete action samples generated by the sliding window segmentation. Each element in the i-th intermediate frequency signal matrix is represented as X i (m, n sub ), n sub = 1, 2, ..., N, where N is the total number of transmitted Chirps corresponding to each intermediate frequency signal matrix;
[0013] (3) Perform a fast Fourier transform on the intermediate frequency signal matrix X i in the fast time dimension to obtain the range-time matrix Y i . Then, perform a fast Fourier transform on the range-time matrix Y i in the slow time dimension to obtain the range-Doppler matrix Z i . Next, map the range-Doppler matrix Z i into a three-channel RGB color image to obtain the range-Doppler spectrogram. Finally, scale the size of the range-Doppler spectrogram to 30×30×3;
[0014] (4) Further segment the range-time matrix Y i along the slow time dimension into L Z range-time submatrices of the same time where is the index of the range-time submatrix;
[0015] (5) Perform a fast Fourier transform on each range-time submatrix separately in the slow time dimension to obtain L Z range-Doppler submatrices
[0016] (6) Apply a static threshold filter to each range-Doppler submatrix to screen out useful human action features. The filtering formula is as follows:
[0017]
[0018] where μ is the threshold ratio, and its value range is [0, 1], and n a = 1, 2,..., N a , n a is the Doppler frequency index of each range-Doppler sub-matrix, and N a is the number of samples of the Doppler frequency, is the dB value of the elements in is the maximum dB value of the elements in is the range-Doppler sub-matrix after being processed by the static threshold filter;
[0019] (7) Stack the range-Doppler sub-matrix after being processed by the static threshold filter along the slow time dimension to form a three-dimensional tensor of range-Doppler-time. Store the coordinates of the points with intensities greater than 0 in a list, and assign four basic variables: range, Doppler, time, and intensity, so as to obtain a four-dimensional point cloud, which represents the evolution of human motion characteristics within a certain window time. Then, screen the number of points in the point cloud stored in the list according to the maximum intensity order, and select the top 1024 points with the maximum intensity.
[0020] Further preferably, in step 2, the loss function of the PVT-BiLSTM network is as follows:
[0021] L(p t ) = α·CE(p t ) + β·FL(p t ) (3)
[0022] where p t represents the probability of the predicted true positive, t = 1, 2,..., Q, and Q is the number of samples in a batch. C(p t ) represents the cross-entropy loss value of sample t, FL(p t ) represents the focal loss value of sample t, L(p t ) represents the combined loss value of sample t, and both α and β are weight coefficients used to adjust the weights of the two loss functions;
[0023] Among them, the focal loss formula is as follows:
[0024] FL(p t ) = -α t (1 - p t ) γ log(p t ) (2)
[0025] where p tRepresents the probability of a predicted true positive example, t = 1, 2, ..., Q, where Q is the number of samples in a batch, and α t is the balance factor used to adjust the ratio between positive and negative sample losses, and γ is the focusing parameter used to adjust the weights of easy-to-classify samples. When γ is 0, the focal loss degenerates into the standard cross-entropy loss.
[0026] Further preferably, α = β = 0.5, α t = 0.25, and γ = 2.
[0027] Further preferably, in step 3, the complete network model training process is as follows:
[0028] For each input point cloud data sample with a dimension of 1024x4, the Pointnet-BiLSTM module multiplies it by a 4×4 transformation matrix obtained through feature extraction learning by the T-net network to obtain a spatially corrected 1024×4 point cloud, and then inputs it into the first two convolutional layers for feature extraction. Then, it multiplies with a 32×32 transformation matrix of the second T-net network, and then inputs it into the third and fourth convolutional layers to map the features to 1024×64. Through two layers of BiLSTM, the features are further dimensionally elevated to 1024×256. The point cloud features extracted by the one-dimensional convolutional layer meet the input requirements of BiLSTM. BiLSTM regards the point cloud features as a sequence and obtains more comprehensive context information by running two LSTMs simultaneously in the forward and reverse directions of the sequence, complementing the convolutional layer to perform dimensional elevation mapping on the point cloud features. Subsequently, it is input into the last convolutional layer to map the features to 1024×512, and then the features are retained to a 1×512 vector through a one-dimensional global max pooling operation. After dimensional reduction through two fully connected layers, the optimal feature vector is obtained
[0029] The ViT module first divides each input 30×30×3 RGB range-Doppler spectrogram into a series of flattened squares where H and W are the height and width of the spectrogram, H = 30, W = 30, the Patch size P = 6, and a total of N z = HW / P 2The number of patches. Through patch segmentation, the entire range-Doppler spectrogram is reshaped into a one-dimensional sequence. The patches obtained by segmentation and flattening are then linearly projected through a fully connected layer without an activation function to generate patch embeddings, mapping the feature vectors of each patch to a new feature space. After patch embedding, class tokens are added to the sequence and passed through all layers of the model to capture the global information of the entire input sequence. Finally, position embeddings are introduced by adding an additional vector that associates each patch with its position in the original range-Doppler map;
[0030] The Transformer encoder consists of two key layers: the MSA layer and the MLP layer. Both layers are designed with residual connections, and an LN layer is applied before each layer;
[0031] The MSA layer contains multiple independent and parallel self-attention modules. The MSA layer assigns different weights to each feature vector, and the formula is as follows:
[0032]
[0033] MSA(Q, K, V) = Concatenate(head 1 , head 2 ,..., head θ )W O (5)
[0034]
[0035] In the formula, Q, K, and V represent the query, key, and value matrices respectively, QK T represents the similarity score of the query and key pair, d represents the dimension of the feature vector processed by each self-attention head, represents the scaling factor, and W O represent learnable weight matrices, is the th self-attention head, and θ is the number of self-attention heads;
[0036] After passing through the Transformer encoder layer, an LN layer is applied. The MLP layer in the last layer of the network consists of two fully connected layers with 512 and 256 neurons, and is accompanied by a dropout layer with a dropout rate of 0.5. The GELU function is used as the activation function, and its formula is expressed as:
[0037]
[0038] In the formula, Φ(x) represents the cumulative distribution function of the standard Gaussian distribution;
[0039] The entire computational process of the ViT module is expressed as:
[0040]
[0041] z′ l = MSA(LN(z l-1 )) + z l-1 (9)
[0042] z l = MLP(LN(z′ l )) + z′ l (10)
[0043]
[0044] Wherein, W represents the weight of the linear projection layer, D represents the projection dimension, D = 64, x class represents the learnable class token, W pos represents the positional embedding, L represents the total number of stacked Transformer encoders, L = 8, l represents the serial number of the Transformer encoder, l = 1, 2,..., L, z l represents the output feature of the l-th Transformer encoder, and the class token output from the last Transformer encoder is represented as Applying the final MLP and LN respectively to obtains the best feature vector y that can represent the entire range-Doppler spectrogram,
[0045] The fusion module concatenates the vector x best with y, and the formula is as follows:
[0046] g = Concatenate(x best , y) (12)
[0047] Wherein, g represents the fused feature vector;
[0048] After two fully-connected layers with 512 and 384 neurons respectively using the GELU function as the activation function and two dropout layers with a dropout rate of 0.5, a residual connection is made. The last layer is a fully-connected layer with 5 neurons using the Softmax function as the activation function, and the predicted probability of each class is output.
[0049] Further preferably, in step 3, during the process of training the network model, the AdamW optimizer is adopted, the weight decay coefficient is 0.0001, the batch size is 32, the learning rate is 0.001, and the number of iterations is 60 rounds.
[0050] The continuous human action recognition method based on point cloud data and range-Doppler spectrogram provided by the present invention can accurately realize continuous human action recognition. Using radar four-dimensional point cloud (range-Doppler-time-intensity) and range-Doppler spectrogram multimodal data, it has sufficient representation of human actions and is superior to traditional single-modal data; the adopted PVT-BiLSTM dual-stream network combines Pointnet and BiLSTM, and has the advantages of both local feature extraction and time series feature extraction. Combining Pointnet-BiLSTM with the ViT network can complement the features of point cloud and spectrogram; the adopted loss function combining cross-entropy loss and focal loss can solve the problem of sample imbalance under continuous human actions and can significantly improve the recognition performance of continuous human actions. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a flowchart of the continuous human action recognition method based on point cloud data and range-Doppler spectrogram provided by the present invention;
[0052] Figure 2 is the overall structure diagram of the PVT-BiLSTM network;
[0053] Figure 3 is the network structure diagram of the T-net module;
[0054] Figure 4 is the network structure diagram of the ViT module;
[0055] Figure 5 is a curve graph showing the change of the action recognition accuracy of the training set and the validation set with the increase of the number of iterations;
[0056] Figure 6 is the confusion matrix graph of the test set;
[0057] Figure 7 is a bar graph of the accuracy of the test set and related evaluation indicators. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] The following further elaborates in detail the specific embodiments of the present invention in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the present invention, but not to limit the scope of the present invention. As Figure 1 shown, the present invention provides a continuous human action recognition method based on point cloud data and range-Doppler spectrogram, including the following steps:
[0059] Step 1: Build an FMCW radar data acquisition platform, and use the platform to collect a series of echo signals containing various continuous human actions to obtain sampled intermediate-frequency signals. Then, preprocess the intermediate-frequency signals to obtain an overall matrix of intermediate-frequency signals. Next, perform sliding window segmentation on the overall matrix of intermediate-frequency signals to obtain intermediate-frequency signal matrices corresponding to respective discrete actions. Finally, perform signal processing on the intermediate-frequency signal matrices to obtain corresponding range-Doppler spectrograms and point cloud data, which specifically include the following steps:
[0060] (1) Use the FMCW radar data acquisition platform to measure the echo of a continuous action of a complete human body, and sample the echo of the continuous human action to obtain sampled intermediate-frequency signals. Then, stack the sampled intermediate-frequency signals column by column to form an overall matrix S of intermediate-frequency signals IF , and represent each element of the overall matrix s of intermediate-frequency signals IF as S IF (m, n), where m = 1, 2,..., M, n = 1, 2,..., N total , m is the fast time index, n is the slow time index, M is the total number of fast time samples corresponding to a single Chirp, and N total is the total number of transmitted Chirps;
[0061] (2) Use a fixed window length and a sliding step to perform sliding window segmentation on the overall matrix s of intermediate-frequency signals IF along the slow time dimension to obtain intermediate-frequency signal matrices X corresponding to each single discrete action i , where i = 1, 2,..., I, and I is the number of discrete action samples generated by sliding window segmentation. Represent each element in the i-th intermediate-frequency signal matrix as X i (m, n sub ), where n sub = 1, 2,..., N, and N is the total number of transmitted Chirps corresponding to each intermediate-frequency signal matrix;
[0062] (3) Perform fast Fourier transform (FFT) on the intermediate-frequency signal matrix X i in the fast time dimension to obtain a range-time matrix Y i . Then, perform FFT on the range-time matrix Y i in the slow time dimension to obtain a range-Doppler matrix Z i . Next, map the range-Doppler matrix Z i into a three-channel RGB color image to obtain a range-Doppler spectrogram. Finally, scale the size of the range-Doppler spectrogram to 30×30×3;
[0063] (4) Further segment the range-time matrix Y i along the slow time dimension into LZ Distance-time sub-matrix at the same time wherein is the index of the distance-time sub-matrix;
[0064] (5) For each distance-time sub-matrix perform FFT on the slow-time dimension respectively to obtain L Z distance-Doppler sub-matrices
[0065] (6) Apply a static threshold filter to each distance-Doppler sub-matrix to screen out useful human motion features. The filtering formula is as follows:
[0066]
[0067] where μ is the threshold ratio, and its value range is [0, 1]. Preferably, μ = 0.5, n a = 1, 2,..., N a , n a is the Doppler frequency index of each distance-Doppler sub-matrix, N a is the number of samples of the Doppler frequency, is the dB value of the elements in is the maximum dB value of the elements in is the distance-Doppler sub-matrix after being processed by the static threshold filter;
[0068] (7) Stack the distance-Doppler sub-matrices after being processed by the static threshold filter along the slow-time dimension to form a three-dimensional distance-Doppler-time tensor. Store the coordinates of the points with intensities greater than 0 in the tensor into a list, and assign four basic variables: distance, Doppler, time, and intensity, so as to obtain a four-dimensional point cloud, which represents the evolution of human motion features within a certain window time. At this time, the stored list contains the point cloud of all points. Fix the number of points in the point cloud uniformly during the network input stage. Since the method of generating the four-dimensional point cloud does not need to worry about the problem of insufficient quantity, screen the number of points in the point cloud in the order of the maximum intensity, and select the top 1024 points with the maximum intensity.
[0069] Step 2: Build a PVT-BiLSTM network, as Figure 2 shown. The PVT-BiLSTM network includes a Pointnet-BiLSTM module, a ViT module, and a fusion module. Among them, the Pointnet-BiLSTM module is designed as Figure 2As shown in the figure, the network adopts two T-net networks, five one-dimensional convolutional (Conv1D) layers, two BiLSTM layers with 64 and 128 neurons respectively, one one-dimensional global max pooling (GlobalMaxPooling1D) layer, and two fully connected (Dense) layers with 256 and 128 neurons respectively. Among them, each convolutional layer and fully connected layer contains a BatchNorm (BN) layer and a ReLU activation function. The number of convolutional kernels in the first three one-dimensional convolutional layers is 32 and the size is 1. The number of convolutional kernels in the fourth convolutional layer is 64 and the size is 1. The number of convolutional kernels in the fifth convolutional layer is 512 and the size is 1. The T-net network is designed as shown in Figure 3 As shown in the figure, the network contains three one-dimensional convolutional layers with 32, 64, and 512 convolutional kernels respectively and a size of 1, one one-dimensional global max pooling layer, two fully connected layers with 256 and 128 neurons respectively, a fully connected layer containing r 2 neurons, and a reshaped transformation matrix. The size of the transformation matrix of the T-net network is r×r, where r is the dimensional value of the transformation matrix. For the two T-net networks, the dimensional values of the transformation matrix are 4 and 32 respectively;
[0070] The ViT module is designed as shown in Figure 4 As shown in the figure. The module is divided into five processing steps: Patch segmentation, linear projection, class token, positional embedding, Transformer encoder, and multi-layer perceptron (MLP);
[0071] Set the loss function of the PVT-BiLSTM network. The cross-entropy function is usually used as the loss function in traditional classification tasks, but there is a problem of unbalanced action samples of different categories in continuous human action recognition. Focal loss is proposed for the problem of unbalanced sample categories. It makes the model pay more attention to difficult-to-classify samples by reducing the weights of easy-to-classify samples and increasing the weights of difficult-to-classify samples. Therefore, the present invention designs a loss function that combines focal loss and cross-entropy loss. Among them, the focal loss formula is as follows:
[0072] FL(p t )=-α t (1-p t ) γ log(p t ) (2)
[0073] In the formula, p t represents the probability of the true positive example predicted, t = 1, 2,..., Q, Q is the number of samples in a batch, and α t is the balance factor, which is used to adjust the ratio between the losses of positive and negative samples. Preferably, αt = 0.25, where γ is the focusing parameter used to adjust the weight of easily classified samples. Preferably, γ = 2. When γ is 0, the focal loss degenerates into the standard cross-entropy loss. FL(p t ) represents the focal loss value for sample t. The loss function formula combining the focal loss and the cross-entropy loss is defined as follows:
[0074] L(p t ) = α·CE(p t ) + β·FL(p t ) (3)
[0075] In the formula, CE(p t ) represents the cross-entropy loss value for sample t, and L(p t ) represents the combined loss value for sample t. Both α and β are weight coefficients used to adjust the weights of the two loss functions. Preferably, α = β = 0.5;
[0076] Step 3: Divide the range-Doppler spectrogram and the point cloud data into a training set and a validation set, and input them into the PVT-BiLSTM network for training and validation to obtain a trained PVT-BiLSTM network model. Then, use the trained PVT-BiLSTM network model to identify continuous human actions. It should be particularly emphasized that since this network is a two-stream input framework, the sizes and label correspondence relationships of the two data sets are required to be consistent. The specific process of training the complete network model includes the following steps:
[0077] For each input of the point cloud data sample with a dimension of 1024×4, the Pointnet-BiLSTM module multiplies it by the 4×4 transformation matrix obtained by feature extraction learning through the T-net network to obtain the spatially corrected 1024×4 point cloud, and then inputs it into the first two convolutional layers for feature extraction. Then, it multiplies by the 32×32 transformation matrix of the second T-net network, and then inputs it into the third and fourth convolutional layers to map the features to 1024×64. Through two layers of BiLSTM, the features are further dimensionally enhanced to be mapped to 1024×256. The point cloud features extracted by the one-dimensional convolutional layer meet the input requirements of BiLSTM. BiLSTM can regard the point cloud features as a sequence, and by running two LSTMs simultaneously in the forward and reverse directions of the sequence, more comprehensive context information can be obtained, which can effectively perform dimensional mapping of the point cloud features in complement with the convolutional layer. Subsequently, it is input into the last convolutional layer to map the features to 1024×512, and then through the one-dimensional global max pooling operation, the features are retained to a 1×512 vector, and the optimal feature vector is obtained after dimensional reduction through two fully connected layers To prepare for subsequent fusion.
[0078] The Vision Transformer (ViT) module first takes each input RGB range-Doppler spectrogram of 30×30×3 and divides it into a series of flattened squares where H and W are the height and width of the spectrogram, H = 30, W = 30, the patch size P = 6, and a total of N z = HW / P 2 patch numbers are generated. Through patch division, the entire range-Doppler spectrogram is reshaped into a one-dimensional sequence. The patches obtained by dividing and flattening are then linearly projected through a fully connected layer without an activation function to generate patch embeddings. This linear projection is a direct mathematical transformation that maps the feature vector of each patch to a new feature space without introducing any non-linear changes. After patch embedding, a class token is added to the sequence. The class token is a special vector, usually located at the beginning of the sequence and passed through all layers of the model to capture the global information of the entire input sequence. This token is crucial for classification tasks because it allows the model to distinguish different classes and plays a role in the final classification decision. Finally, to preserve the position information of the patches in the original range-Doppler spectrogram, position embeddings are introduced. Position embeddings are achieved by adding an additional vector that associates each patch with its position in the original range-Doppler map to the patch embeddings. This enables the model to identify the position of each individual patch, which is a crucial step because the multi-head self-attention (MSA) layer in the Transformer encoder does not consider the order or position of the patches and only focuses on their mutual relationships. The position embedding of the class token is usually a learnable parameter that allows the model to adjust its value during training to better perform the classification task. In this way, the ViT model can not only understand the mutual relationships between patches but also utilize position information to improve the accuracy of classification.
[0079] The Transformer encoder consists of two key layers: the MSA layer and the MLP layer. Both of these layers are designed with residual connections, and a LayerNorm (LN) layer is applied before each layer. In the process of finding the most effective representation of the learnable class token, ViT operates in a manner similar to dictionary lookup.
[0080] The MSA layer contains multiple independent and parallel self-attention modules. The advantage of MSA is that it assigns different weights to each feature vector, which helps to extract useful features. Its formula is as follows:
[0081]
[0082] MSA(Q, K, V) = Concatenate(head 1 , head 2 ,..., head θ )W O (5)
[0083]
[0084] where Q, K, and v represent the query, key, and value matrices respectively, QK T represents the similarity score of the query and key pair, d represents the dimension of the feature vector processed by each self-attention head, represents the scaling factor, and W O represent learnable weight matrices, is the th self-attention head, and θ = 4 is the number of self-attention heads.
[0085] After passing through the Transformer encoder layer, an LN layer is applied. The MLP layer in the last layer of the network consists of two fully connected layers with 512 and 256 neurons respectively, and is accompanied by a Dropout layer with a dropout rate of 0.5 to avoid overfitting. The GELU function is used as the activation function, and its formula can be expressed as:
[0086]
[0087] where Φ(x) represents the cumulative distribution function of the standard Gaussian distribution.
[0088] Finally, the entire computational process of ViT can be expressed as:
[0089]
[0090] z' l = MSA(LN(z l-1 )) + z l-1 (9)
[0091] z l = MLP(LN(z' l )) + z' l (10)
[0092]
[0093] where W represents the weight of the linear projection layer, D represents the projection dimension, D = 64, x class represents the learnable class token, W pos represents the positional embedding, Let \(L\) denote the total number of stacked Transformer encoders, where \(L = 8\), and let \(l\) denote the serial number of the Transformer encoder, \(l=1,2,\cdots,L\), \(z\) l represents the output feature of the \(l\)-th Transformer encoder, and the class token output from the last Transformer encoder is denoted as Apply the final MLP and LN to obtain the optimal feature vector \(y\) that can represent the entire range-Doppler spectrogram,
[0094] The fusion module concatenates the vector \(x\) best with \(y\), and the formula is as follows:
[0095] \(g=\text{Concatenate}(x\) best ,y)(12)
[0096] In the formula, represents the fused feature vector;
[0097] After two fully connected layers with 512 and 384 neurons respectively using the GELU function as the activation function and two dropout layers with a dropout rate of 0.5, residual connection is performed. The last layer is a fully connected layer with 5 neurons using the Softmax function as the activation function to output the prediction probability of each class.
[0098] The network model is trained from scratch using the training set data and the AdamW optimizer. The weight decay coefficient is 0.0001, which helps to avoid the influence of weight decay on the bias parameters. The batch size is 32, the learning rate is 0.001, and the number of iterations is 60 rounds. The validation set data is used to monitor the performance of the network model until the network converges, and the trained network model is obtained.
[0099] The following uses specific embodiments to elaborate in detail on the above method for continuous human action recognition of FMCW radar based on the fusion of point cloud data and range-Doppler spectrogram:
[0100] In this embodiment, the FMCW radar system is placed on a table 1.2 meters high, and a continuous action sequence including five actions is measured for 5 volunteers in an indoor environment, including (a) marching in place, (b) standing still, (c) walking, (d) jumping, and (e) bending down. The starting frequency is 77 GHz, the bandwidth is 2 GHz, the ADC sampling frequency is 32 MHz, the number of fast-time sampling points is 128, the number of slow-time sampling points is 64000, the frame period is 100 ms, the number of frames is 500, and the number of chirp signals per frame is 128. Four sets of action sequences are collected for each participant, totaling 20 sets of action sequences. The measurement duration of each continuous action sequence is 50 seconds and is composed of five different action combinations.
[0101] The overall matrix of the intermediate-frequency signal is segmented by a sliding window and signal processing is performed to obtain point cloud data and range-Doppler spectrograms. The sliding window segmentation has a window length of 2 seconds and a sliding step of 0.1 second. When the point cloud data is generated, the corresponding duration of the range-time submatrix is 0.25 seconds, and the sliding step of the range-time submatrix is 0.125 seconds. Each point cloud data contains 15 range-Doppler submatrices, and the static filter threshold ratio is 0.5. After preprocessing, 7680 four-dimensional point cloud data and 7680 three-channel RGB range-Doppler spectrograms are obtained.
[0102] The point cloud data and range-Doppler spectrograms are used to divide the dataset into a training set and a test set according to a ratio of 4:1 in the same random order. Then, 20% of the training set is further divided into a validation set. Taking the range-Doppler spectrogram as an example, a total of 4916 training set images, 1228 validation set images, and 1536 test set images are generated.
[0103] The number of point cloud data is selected as 1024 points, and the size of the range-Doppler spectrogram is 30×30×3. Training starts from scratch using the AdamW optimizer, with a weight decay coefficient of 0.0001, which helps to avoid the influence of weight decay on the bias parameters. The batch size is 32, the learning rate is 0.001, and the number of iterations is 60 rounds. The network model of the present invention uses the TensorFlow 2.8 deep learning framework, and all model training, validation, and testing experiments are carried out on an NVIDIA GTX 1050Ti GPU.
[0104] The experimental results of the curve graph showing the change of the action recognition accuracy of the training set and the validation set with the increase of the number of iterations are as Figure 5 shown. In the first round, the accuracy of the validation set is as high as 83%. As the number of iterations increases, the accuracy of the validation set stabilizes at about 98%. Figure 6 The confusion matrix of the test set of the five actions is shown. The abscissa represents the predicted label, and the ordinate represents the true label. Due to continuous action transition segments and human marking errors, there are inevitably some prediction errors. However, the proposed network model still maintains high recognition performance.Figure 7 The test accuracy, precision, recall, and F1 score of the network model are given. The test set accuracy reaches 98.24%.
[0105] The continuous human action recognition method based on point cloud data and range-Doppler spectrogram provided by the present invention uses a Pointnet-BiLSTM network to learn four-dimensional radar point cloud data in the PVT-BiLSTM dual-stream network framework. At the same time, it makes full use of the ViT network with self-attention mechanism to learn the features of the range-Doppler spectrogram, and finally classifies by fusing the features of the two single-stream modalities. The cross-entropy loss and focal loss functions are combined to solve the problem of sample classification imbalance under continuous human actions, effectively improving the accuracy of continuous human action recognition.
[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.
Claims
1. A method for continuous human motion recognition based on point cloud data and range Doppler spectrogram, characterized in that: include: Step 1: Build an FMCW radar data acquisition platform, and use the platform to collect a series of echoes containing multiple continuous human motions to obtain sampled intermediate frequency signals. Then, preprocess the intermediate frequency signals to obtain an overall intermediate frequency signal matrix. Then, use a sliding window to segment the overall intermediate frequency signal matrix to obtain an intermediate frequency signal matrix of corresponding discrete motions. Finally, perform signal processing on the intermediate frequency signal matrix to obtain corresponding range Doppler spectra and point cloud data. Step 2: Build a PVT-BiLSTM network, wherein the PVT-BiLSTM network includes a Pointnet-BiLSTM module, a ViT module and a fusion module. The Pointnet-BiLSTM module uses two T-net networks, five one-dimensional convolutional layers, two BiLSTM layers with 64 and 128 neurons respectively, a one-dimensional global maximum pooling layer and two fully connected layers with 256 and 128 neurons respectively, wherein each convolutional layer and fully connected layer contains a BN layer and a ReLU activation function, the first three one-dimensional convolutional layers have 32 convolution kernels and a size of 1, the fourth convolutional layer has 64 convolution kernels and a size of 1, the fifth convolutional layer has 512 convolution kernels and a size of 1, the T-net network contains three one-dimensional convolutional layers with three convolution kernels of 32, 64, 512 and a size of 1, a one-dimensional global maximum pooling layer, two fully connected layers with 256 and 128 neurons respectively, and a one-dimensional global maximum pooling layer. 2 The fully connected layer of neurons and the reshaped transformation matrix, the transformation matrix size of the T-net network is r×r, where r is the dimension value of the transformation matrix. For the two T-net networks, the dimension values of the transformation matrix are 4 and 32 respectively; the ViT module is divided into five processing steps: patch segmentation, linear projection, category labeling, position embedding, Transformer encoder and multi-layer perceptron; Step 3: Divide the range Doppler spectrum and point cloud data into a training set and a validation set, and input them into the PVT-BiLSTM network for training and validation to obtain a trained PVT-BiLSTM network model. Then, use the trained PVT-BiLSTM network model to recognize continuous human body movements.
2. The method for continuous human motion recognition based on point cloud data and range Doppler spectrogram according to claim 1, characterized in that: Step 1 includes the following steps: (1) An FMCW radar data acquisition platform is used to measure the echo of a complete human body continuous motion, and the human body continuous motion echo is sampled to obtain a sampled intermediate frequency signal. Then, the sampled intermediate frequency signal is stacked by column into an intermediate frequency signal overall matrix s IF , the intermediate frequency signal overall matrix S IF Each element is represented by S IF (m, n), where m = 1, 2, ..., M, n = 1, 2, ..., N total , m is the fast time index, n is the slow time index, M is the total number of fast time samples corresponding to a single Chirp, N total is the total number of Chirps emitted; (2) Using a fixed window length and a sliding step size to calculate the intermediate frequency signal overall matrix S IF Sliding window segmentation is performed along the slow time dimension to obtain the intermediate frequency signal matrix X corresponding to each single discrete action i , i = 1, 2, ..., I, I is the number of discrete action samples generated by sliding window segmentation, and each element in the i-th intermediate frequency signal matrix is represented as X i (m,n sub ), n sub =1, 2, ..., N, where N is the total number of transmitted Chirps corresponding to each intermediate frequency signal matrix; (3) The intermediate frequency signal matrix X i Perform a fast Fourier transform on the fast time dimension to obtain the distance-time matrix Y i , then, the distance time matrix Y i Perform a fast Fourier transform in the slow time dimension to obtain the range Doppler matrix Z i , then, the range Doppler matrix Z i Mapping into a three-channel RGB color image to obtain a range Doppler spectrum, and finally scaling the size of the range Doppler spectrum to 30×30×3; (4) For the distance time matrix Y i Further divided into L along the slow time dimension Z The distance-time submatrices of the same time in, is the index of the distance-time submatrix; (5) For each distance-time submatrix Perform fast Fourier transform of the slow time dimension to obtain L Z Range Doppler submatrix (6) For each range-Doppler submatrix A static threshold filter is applied to filter useful human motion features, where the filtering formula is as follows: Where μ is the threshold ratio, ranging from [0, 1], n a =1,2,...,N a , n a is the Doppler frequency index of each range-Doppler submatrix, N a is the number of samples of Doppler frequency, for The dB value of the elements in, for The maximum dB value of the elements in , is the range-Doppler submatrix after being processed by the static threshold filter; (7) The range-Doppler submatrix after being processed by the static threshold filter The slow-time dimensions are stacked to form a three-dimensional tensor of distance-Doppler-time. The coordinates of the points in the tensor with intensity greater than 0 are stored in a list and assigned with four basic variables: distance, Doppler, time, and intensity, thereby obtaining a four-dimensional point cloud, which represents the evolution of human motion characteristics within a certain window time. Afterwards, the point cloud points in the stored list are filtered in order of maximum intensity, and the top 1024 points with the largest intensity are filtered out.
3. The method for continuous human motion recognition based on point cloud data and range Doppler spectrogram according to claim 1, characterized in that: In step 2, the loss function of the PVT-BiLSTM network is as follows: L(p t )=α·CE(p t )+β·FL(p t ) (3) In the formula, p t represents the probability of the predicted true positive, t = 1, 2, ..., Q, Q is the number of samples in a batch, CE(p t ) represents the cross entropy loss value for sample t, FL(p t ) represents the focal loss value for sample t, L(p t ) represents the combined loss value for sample t, α and β are weight coefficients used to adjust the weights of the two loss functions; Among them, the focal loss formula is as follows: FL(p t )D-α t (1-px) γ log ( p ) t ) (2) In the formula, p t represents the probability of the predicted true positive, t = 1, 2, ..., Q, Q is the number of samples in a batch, α t is a balancing factor used to adjust the ratio between the positive and negative sample losses, γ is a focusing parameter used to adjust the weight of easy-to-classify samples. When γ is 0, the focus loss degenerates into the standard cross entropy loss.
4. The method for continuous human motion recognition based on point cloud data and range Doppler spectrogram according to claim 3, characterized in that: α=β=0.5,α t =0.25, γ=2.
5. The method for continuous human motion recognition based on point cloud data and range Doppler spectrogram according to claim 1, characterized in that: In step 3, the complete network model training process is as follows: For each point cloud data sample with a dimension of 1024×4, the Pointnet-BiLSTM module multiplies it with the 4×4 transformation matrix obtained by the T-net network through feature extraction to obtain a spatially corrected 1024×4 point cloud. It is then input into the first two convolutional layers for feature extraction, and then multiplied with the 32×32 transformation matrix of the second T-net network. It is then input into the third and fourth convolutional layers to map the features to 1024×64, and further feature dimension is increased through two layers of BiLSTM. The feature is mapped to 1024x256. The point cloud features extracted by the one-dimensional convolution layer meet the input requirements of BiLSTM. BiLSTM regards the point cloud features as a sequence. By running two LSTMs simultaneously in the forward and reverse directions of the sequence, more comprehensive context information is obtained, which complements the convolution layer to map the point cloud features to a higher dimension. The feature is then input to the last convolution layer to map the feature to 1024×512, and then the feature is retained to a 1×512 vector through a one-dimensional global maximum pooling operation. After the dimension is reduced by two fully connected layers, the optimal feature vector is obtained. The ViT module first converts each input 30×30×3 RGB range Doppler spectrum Split into a series of flattened squares Among them, H and W are the height and width of the spectrum, H = 30, W = 30, Patch size P = 6, and a total of N z =HW / P 2 The number of patches is , through patch segmentation, the entire range-Doppler spectrum is reshaped into a one-dimensional sequence, and the patch obtained by segmentation and flattening is then linearly projected through a fully connected layer without an activation function to produce a patch embedding, mapping the feature vector of each patch to a new feature space. After patch embedding, the category label is added to the sequence and passed through all layers of the model to capture the global information of the entire input sequence. Finally, position embedding is introduced by adding an additional vector that associates each patch with its position in the original range-Doppler map to the patch embedding; The Transformer encoder consists of two key layers: the MSA layer and the MLP layer. Both layers are designed with residual connections, and the LN layer is applied before each layer. The MSA layer contains multiple independent and parallel self-attention modules. The MSA layer assigns different weights to each feature vector. The formula is as follows: MSA(Q,K,V)=Concatenate(head1,head2,...,head θ )W O (5) Where Q, K, and V represent query, key, and value matrices respectively. QK T represents the similarity score between the query and the key pair, d represents the dimension of the feature vector processed by each self-attention head, represents the scale factor, and W O represents the learnable weight matrix, It is self-attention heads, θ is the number of self-attention heads; After the Transformer encoder layer, a LN layer is applied. The last MLP layer of the network consists of two fully connected layers with 512 and 256 neurons, accompanied by a dropout layer with a dropout rate of 0.
5. The GELU function is used as the activation function, which is expressed as: Where Φ(x) represents the cumulative distribution function of the standard Gaussian distribution; The entire calculation process of the ViT module is expressed as: z′ l =MSA(LN(z l-1 ))+z l-1 (9) With l =MLP(LN(z′ l ))+z′ l (10) Where W represents the weight of the linear projection layer, D represents the projection dimension, D = 64, x class represents a learnable category label, W pos represents position embedding, L represents the total number of stacked Transformer encoders, L=8, and l represents the sequence number of the Transformer encoder, l=1, 2, ..., L, z l represents the output feature of the lth Transformer encoder, and the category label output from the last Transformer encoder is represented as Apply the final MLP and LN to Get the best feature vector y that can represent the entire range Doppler spectrum, The fusion module transforms the vector x best Concatenate with y, the formula is as follows: g=Concatenate(x best ,y) (12) In the formula, g represents the fused feature vector; After two fully connected layers with 512 and 384 neurons using the GELU function as the activation function and two dropout layers with a dropout rate of 0.5, a residual connection is performed. The last layer is a fully connected layer with 5 neurons using Softmax as the activation function to output the predicted probability of each category.
6. The method for continuous human motion recognition based on point cloud data and range Doppler spectrogram according to claim 5, characterized in that: In step 3, during the training of the network model, the AdamW optimizer is used, the weight decay coefficient is 0.0001, the batch size is 32, the learning rate is 0.001, and the number of iterations is 60.
Citation Information
Patent Citations
Frequency modulation continuous wave radar human body action recognition method based on data enhancement
CN113296087A
Shower head adjusting method, shower head, storage medium and electronic equipment
CN114332610A
Cross-modal supervised pedestrian gait recognition method based on millimeter wave radar
CN117746496A
Through-the-wall radar human body behavior identification method based on micro-Doppler angular point features and dynamic graph neural network
CN118068320A
Three-dimensional lidar point cloud semantic segmentation method and apparatus based on deep learning
WO2024130776A1
Cited By
Target identification method, meteorological radar equipment and storage medium
CN120831644A