Lightweight human body action recognition method based on millimeter wave radar three-dimensional point cloud

Through the combination of mPCT module and LSTM module, the problem of high model complexity in millimeter wave radar sparse point cloud recognition is solved, and the balance between high recognition accuracy and low model complexity is achieved, which is suitable for embedded devices.

CN120340116APending Publication Date: 2025-07-18JIANGSU UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510324740.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing millimeter-wave radar human body movement recognition methods are difficult to reduce model complexity while ensuring high recognition accuracy. Especially the processing of sparse point clouds leads to excessive computational complexity and cannot be applied to embedded terminal devices with limited storage space.

Method used

Using the combination of mPCT module and LSTM module, the mPCT module extracts the spatial characteristics of sparse point clouds through the embedding layer and bias attention mechanism. The LSTM module captures the timing relationship between multi-frame point clouds and combines the full connection layer to output the recognition results.

Benefits of technology

On the premise of ensuring high recognition accuracy, the number of convolutional layers is significantly reduced, the complexity of the model is reduced, and it is suitable for embedded device applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340116A_ABST
    Figure CN120340116A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight human body action recognition method based on millimeter-wave radar three-dimensional point cloud, and the method comprises the following steps: carrying out the processing of a millimeter-wave radar sampling signal, and obtaining the three-dimensional point cloud information of human body actions; constructing a lightweight network, wherein the lightweight network comprises an mPCT module and an LSTM module; for the three-dimensional point cloud information, performing three-dimensional point cloud spatial feature extraction by using an mPCT module; according to the spatial features of the three-dimensional point clouds, performing time sequence feature extraction among multiple frames of point clouds by using an LSTM module; and mapping the extracted time sequence features among the multiple frames of point clouds to a label set through nonlinear transformation by a full connection layer, and outputting an identification result. According to the method, the number of convolutional layers is greatly reduced on the premise of ensuring the accuracy, and the point cloud spatial features are extracted to the greatest extent through a bias attention mechanism under the condition of not increasing the number of features, so that the method can give consideration to high recognition accuracy and low model complexity at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of millimeter-wave radar human motion recognition, relates to human motion recognition, and specifically relates to a lightweight human motion recognition method based on millimeter-wave radar three-dimensional point cloud. Background Art

[0002] Non-contact human activity recognition (HAR) is widely used in scenarios such as smart homes, autonomous driving, augmented reality, virtual reality, and activity monitoring for the disabled and the elderly, as described in the literature "Q. Wan, Y. Li, C. Li, et al., 'Gesture recognition for smart home applications using portable radar sensors,' in 201436th Annual International Conference of the IEEE Engineering in Medicine and Biology Society, Chicago, IL, USA, 2014, pp. 6414 - 6417." And accurate human activity recognition is the key to realizing the above applications. Traditional non-contact human activity recognition is achieved through vision sensors, ultrasonic sensors, lidar, etc. However, vision sensors may invade personal privacy and cannot be used in dark environments. The sensing distance of ultrasonic sensors is limited, and lidar is expensive and cannot be used when blocked. Compared with the above three sensors, millimeter-wave radar is low-cost, has a relatively long sensing distance, can work normally in environments without light such as at night, and has no risk of invading personal privacy. Therefore, in recent years, human activity recognition based on millimeter-wave radar has received extensive attention, as described in the literature "B. Jin, X. Ma, Z. Zhang, Z. Lian, and B. Wang, 'Interference-Robust Millimeter-Wave Radar-Based Dynamic Hand Gesture Recognition Using 2-DCNN-Transformer Networks,' IEEE Internet Things J., vol. 11, no. 2, pp. 2741–2752, Jan. 2024. DOI: 10.1109 / JIOT.2023.3293092".

[0003] The millimeter-wave radar processes the human echo signal, extracts relevant features, and then uses a classifier to identify human activities. Among them, feature extraction and recognition classification are the two most crucial steps. Currently, most human activity recognition methods using millimeter-wave radar first construct feature spectrograms of human postures, such as Doppler-time spectrograms, range-Doppler spectrograms, and angle-Doppler spectrograms, and then input the two-dimensional spectrogram features into corresponding neural networks for classification. For example, the literature "K. Alirezazad and L. Maurer, 'FMCW radar-based hand gesture recognition using dual-stream CNN-GRU model.' in 2022 24th International Microwave and Radar Conference (MIKON), Gdansk, Poland, 2022, pp. 1-5. DOI: 10.23919 / MIKON54314.2022.9924984" extracts range-Doppler maps and range-angle maps and inputs them into a multi-input single-output network of a 2D convolutional neural network-gated recurrent unit (2DCNN-GRU), with an identification accuracy rate of 92.50%. The Gan team designed a feature extraction method of range-Doppler matrix focusing (RDMF) and used a 3D CNN + LSTM encoder classification framework to classify the features, achieving an average accuracy rate of 95.9% for 8 human activities. The literature "B. Jin, Y. Peng, X. Kuang, Z. Zhang, Z. Lian, and B. Wang, 'Robust dynamic hand gesture recognition based on millimeter wave radar using atten-tsnn,' IEEE Sens. J., vol. 22, no. 11, pp. 10861–10869, Jun. 2022. DOI: 10.1109 / JSEN.2022.3170311" constructs range-time and Doppler-time feature images of dynamic gestures and learns local and global features in the feature images based on a neural network of self-attention time series, achieving an accuracy rate of 98.43% on a dataset of 5 types of activities with interference.The literature "Y. Song, L. Wu, Y. Zhao, et al., 'High-accuracy gesture recognition using mm-wave radar based on convolutional block attention module.' in 2023 IEEE International Conference on Image Processing (ICIP), Kuala Lumpur, Malaysia, 2023, pp. 1485-1489. DOI: 10.1109 / ICIP49359.2023.10222362" extracts a hybrid feature spectrogram composed of range-time spectrogram, Doppler-time spectrogram, azimuth-time spectrogram, and elevation-time spectrogram of human activity targets, and uses a neural network model based on DenseNet and convolutional attention module, with an identification accuracy reaching 99.03%. However, the number of parameters of this model is 8.0M, making it difficult to be applied to actual embedded terminal devices.

[0004] Although existing feature extraction methods have achieved good recognition effects. However, the features extracted by the above methods, such as two-dimensional feature spectrograms, have limited interpretability and are difficult to distinguish the movements of different body parts. Moreover, there are two major problems with this form of input data: First, there is a large amount of redundant information in the feature spectrograms, and the neural network will learn a large amount of information unrelated to human behavior postures from these feature spectrograms, which to a certain extent affects the recognition accuracy. Second, most current recognition classifiers are based on deep neural networks. The performance of deep neural networks is closely related to the type of features extracted (i.e., the form of data input into the network). Since the feature spectrograms are in two-dimensional or three-dimensional data forms, it will lead to an excessive number of parameters in the backend deep neural network, and a network with overly bloated parameters is not suitable for edge devices with limited storage space.

[0005] To address the above two problems, relevant scholars convert the original echo data of millimeter-wave radar into point cloud features as the input of the model. Compared with traditional methods, the millimeter-wave radar point cloud contains rich and refined geometric structure information of human postures, which can reduce the redundancy of input data. Moreover, the scale of point cloud data is smaller, which can greatly reduce the complexity of the neural network, thereby reducing the data transmission, preprocessing, and inference time required for deployment on real-time terminal devices.

[0006] However, most of the existing point cloud data processing methods are designed for the dense point clouds generated by lidar or depth sensors. For example, the literature "C.R. Qi, H. Su, K. Mo, et al., "Pointnet: Deep learning on point sets for 3d classification and segmentation." in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 652-660." proposed the PointNet architecture, which uses point-wise MLP (Multilayer Perceptron) layers to extract point cloud features and then uses a pooling layer with permutation invariance to aggregate features. The literature "Y. Min, Y. Zhang, X. Chai, et al., "An efficient pointlstm for point clouds based gesture recognition." in 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2020, pp. 5760-5769. DOI: 10.1109 / CVPR42600.2020.00580" proposed Point-LSTM to capture the long-term spatial correlations between dense point cloud sequences. The literature "M.-H. Guo, J.-X. Cai, Z.-N. Liu, T.-J. Mu, R.R. Martin, and S.-M. Hu, "Pct: Point cloud transformer," Comput. Vis. Media (Beijing), vol. 7, no. 2, pp. 187–199, Jun. 2021. DOI: 10.1007 / s41095-021-0229-5" proposed the PCT (Point cloud transformer) model, which uses the permutation invariance inherent in the Transformer when processing sequences to process unordered point clouds.

[0007] Although the above network models have achieved good results in processing dense point clouds, the point clouds generated by millimeter-wave radars are very sparse. If these point cloud processing methods are directly applied to millimeter-wave radar point clouds, it may lead to model overfitting and cause unnecessary computational burdens. The literature "S. Palipana, D. Salami, L. A. Leiva, and S. Sigg, "Pantomime: Mid-air gesture recognition with sparse millimeter-wave radar point clouds," Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., vol. 5, no. 1, pp. 1–27, Mar. 2021. DOI: 10.1145 / 3448110" uses PointNet++ and LSTM networks to process dynamic point clouds. Although it achieves an accuracy of 95.0% on its self-built 21-class Pantomime dataset, it directly applies the dense point cloud processing network PointNet++ to sparse millimeter-wave radars, and its network computational complexity is as high as 15.16 G FLOPS. The literature "D. Salami, R. Hasibi, S. Palipana, P. Popovski, T. Michoel, and S. Sigg, "Tesla-Rapture: A Lightweight Gesture Recognition System From mmWave Radar Sparse Point Clouds," IEEE Trans. Mobile Comput., vol. 22, no. 8, pp. 4946–4960, Aug. 2023. DOI: 10.1109 / TMC.2022.3153717" recently proposed a graph convolution method based on message passing neural networks (MPNNs): Tesla, and its version for embedded devices: Tesla-V. The number of parameters of Tesla-V is 1.56 M, and the computational complexity is 0.40 G FLOPS, achieving an accuracy of 96.6% on the Pantomime dataset. Although these above-mentioned networks have achieved initial success, there is still much room for improvement in terms of accuracy, model parameters, computational complexity, etc. Summary of the Invention

[0008] Object of the Invention: Aiming at the problem that it is impossible to balance high recognition accuracy and low model complexity when using millimeter-wave radars for human action recognition, a lightweight human action recognition method based on three-dimensional point clouds of millimeter-wave radars is provided.

[0009] Technical solution: To achieve the above object, the present invention provides a lightweight human action recognition method based on millimeter-wave radar three-dimensional point cloud, including the following steps:

[0010] S1: Process the millimeter-wave radar sampling signal to obtain three-dimensional point cloud information of human activities;

[0011] S2: Construct a lightweight network, including an mPCT module and an LSTM module;

[0012] S3: For the three-dimensional point cloud information, use the mPCT module to extract three-dimensional point cloud spatial features;

[0013] S4: According to the three-dimensional point cloud spatial features, use the LSTM module to extract temporal features between multiple frames of point clouds;

[0014] S5: Through the fully connected layer, map the temporal features between multiple frames of point clouds extracted in step S4 to the label set through non-linear transformation and output the recognition result.

[0015] Further, the step S1 specifically includes:

[0016] A1: Distance FFT (one-dimensional FFT): Sample the intermediate frequency signal of the millimeter-wave radar and perform FFT operation in the fast time dimension to calculate the corresponding frequency and obtain the distance information R of the human point cloud i , where i is the serial number of the point cloud;

[0017] A2: Velocity FFT (two-dimensional FFT): For the same target point, perform FFT operations on the linear frequency modulation signals of two adjacent periods respectively, that is, perform FFT operation in the slow time dimension to obtain a two-dimensional range-Doppler map (RD map);

[0018] A3: Constant false alarm rate (CFAR) detection: The CFAR algorithm is an adaptive target detection technology used to separate the reflected signal from the background noise. Sum multiple two-dimensional FFT matrices to generate a pre-detection matrix, and then use the CFAR algorithm to perform threshold decision on the pre-detection matrix to obtain multiple corresponding peaks;

[0019] A4: Angle FFT (three-dimensional FFT): Perform FFT operations on the multiple peak points detected by the CFAR algorithm in the angle dimension (including azimuth dimension and elevation dimension), and calculate the azimuth angle θ of each point using the phase difference between antennas i and elevation angle φ i ;

[0020] A5: According to the distance information R i , azimuth angle θ i and elevation angle φ i, convert the point cloud position from the polar coordinate system of the radar to the Cartesian coordinate system to obtain the three-dimensional coordinates (x i , y i , z i ) of the i-th point cloud.

[0021] Further, the three-dimensional coordinates (x i , y i , z i ) of the i-th point cloud in step A5 are expressed as:

[0022]

[0023] Further, in step S2, the mPCT module includes an embedding layer module and a bias attention module. The embedding layer module includes two LBR layers, and four bias attention layers (OAlayer) are connected in series in the bias attention module. After verification, connecting four bias attention layers can maximize the spatial feature extraction ability of the model while controlling the computational complexity from becoming too large.

[0024] Further, the process of using the mPCT module to extract the three-dimensional point cloud spatial features in step S3 includes:

[0025] B1: Use an LBR layer (linear layer + batch normalization + ReLU) to increase the number of feature channels of the input point cloud from 3 to d in ;

[0026] B2: Use the Furthest Points Sample (FPS) algorithm to select N out center points, and obtain k nearest neighbor points for each center point through the K Nearest Neighbors (KNN) algorithm to obtain the features of the nearest neighbor points of each center point, and its data dimension is N out ×k×d in ; The operation of repeating the features of the center point k times is performed so that its data dimension is also N out ×k×d in ;

[0027] B3: Subtract the features of the nearest neighbor points from the features of the center points to obtain the difference features, and then splice these difference features with the features of the center points;

[0028] B4: Through an LBR layer and the operation of taking Max, the embedding layer outputs a tensor with a sampling point number of N out , and the number of feature channels is d out ;

[0029] B5: The bias attention module concatenates the input with the outputs of all bias attention layers in the feature dimension. After sending the concatenated features into the LBR layer, the output of the bias attention module is finally obtained, which is also the output of the mPCT module.

[0030] Further, the operation of the bias attention layer in step B5 includes:

[0031] Considering the disordered characteristics of point clouds, the Query (Q), Key (K), and Value (V) matrices in the module are all generated from the Input through a linear layer with shared weights for all point clouds. The specific calculation method is as follows:

[0032] (Q, K, V) = F in ·(W q , W k , W v )

[0033] Q, K ∈ R N×d / 4 , V ∈ R N×d

[0034] W q , W k ∈ R d×d / 4 , W v ∈ R d×d

[0035] where F in represents the input feature, and W q , W k , W v represent learnable linear transformations shared for all point clouds. Among them, W q , W k outputs features with a channel number of d / 4, and W v outputs features with a channel number of d;

[0036] After obtaining the attention features through the self-attention operation, the offset between the self-attention (SA) features and the input features is calculated by element-wise subtraction. This offset feature is added to the input features element-wise after the LBR layer. In the present invention, the offset attention layer is represented by OA (Offset-Attention), and LBR represents Linear layer + BatchNorm + ReLU. The output F out of the offset attention layer is:

[0037] F out = OA(F in ) = LBR(F in - F sa ) + F in

[0038] where Fsa Denotes the attention feature, and its calculation process is as follows:

[0039] F sa = SL(QK T )·V

[0040] Among them, SL represents the softmax + l1norm operation.

[0041] Furthermore, for the entire bias attention module in step B5, assuming that the tensor shape output by the Embedding layer to the bias attention module is N×d, then the tensor shape after splicing in the four concatenated bias attention layers is N×5d, and the output tensor shape of the final mPCT module after passing through the LBR layer is N×4d.

[0042] Furthermore, the LSTM module in step S4 includes a forget gate F t , an input gate I t , an output gate O t , and a candidate memory cell

[0043] Suppose there are h hidden units, the batch size is n, the input vector dimension is d, the input is X t ∈R d×h , the hidden layer state is H t-1 ∈R n×h , then the specific operation process of the LSTM module is as follows:

[0044]

[0045] Among them, W xf , W xi , W xc W xo ∈R d×h and W hf , W hi , W hc , W ho ∈R h×h are weight parameters, and b f , b i , b c , b o ∈R 1×h are bias parameters.

[0046] Based on the above content, the innovation points of the present invention are summarized as follows:

[0047] Embedding layer module: In the present invention, the features of neighboring points are subtracted from the features of the central point to obtain difference features, and then these difference features are concatenated with the features of the central point. In this way, the obtained features take into account both the features of the central point and the difference features between the central point and its neighbors.

[0048] Self-attention algorithm: The Query (Q), Key (K), and Value (V) matrices in the module are all generated from the Input through a linear layer with shared weights for all point clouds. Therefore, the generated Q, K, and V are permutation-invariant to the point clouds.

[0049] Bias attention method: After obtaining the attention features through the self-attention operation, it calculates the offset between the self-attention (SA) features and the input features through element-wise subtraction. This offset feature is added to the input features element-wise after the LBR layer. In this way, without changing the number of feature channels, both the original input features and the offset features of the self-attention are retained.

[0050] Advantageous effects: Compared with the prior art, the present invention specifically invents the mPCT module for millimeter-wave radar sparse point clouds. The module includes a neighborhood embedding method and a bias attention mechanism, enabling it to directly and efficiently extract the spatial geometric features of millimeter-wave radar sparse point clouds; using the mPCT module and the LSTM module to extract the spatial features of single-frame point clouds and the temporal relationship between multi-frame point clouds respectively, so as to comprehensively utilize the spatio-temporal features of point clouds to capture the deep features of human complex activities; the present invention greatly reduces the number of convolutional layers while ensuring the accuracy rate, and through the bias attention mechanism, it maximally extracts the spatial features of point clouds without increasing the number of features. Therefore, the present invention can take into account both high recognition accuracy and low model complexity. Description of the drawings

[0051] Figure 1 is the basic flowchart of the method of the present invention

[0052] Figure 2 is a schematic diagram of the generation process of millimeter-wave radar point cloud data;

[0053] Figure 3 is the mPCT-LSTM network architecture diagram;

[0054] Figure 4 is the structure diagram of the embedding layer module;

[0055] Figure 5 is the structure diagram of the bias attention module and the bias attention layer;

[0056] Figure 6 is the structure diagram of the LSTM network;

[0057] Figure 7It is a confusion matrix diagram of the mPCT-LSTM model tested on three datasets. Detailed implementation manners

[0058] The present invention will be further clarified below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications made by those skilled in the art fall within the scope defined by the appended claims of this application.

[0059] Embodiment 1:

[0060] This embodiment provides a lightweight human action recognition method based on millimeter-wave radar three-dimensional point cloud, as Figure 1 shown, which includes the following steps:

[0061] S1: Process the millimeter-wave radar sampling signal to obtain three-dimensional point cloud information of human activities;

[0062] S2: Construct a lightweight network, namely mPCT-LSTM, including an mPCT module and an LSTM module, and the overall architecture is as Figure 3 shown. The mPCT module includes an embedding layer module and a bias attention module;

[0063] S3: For the three-dimensional point cloud information, use the mPCT module to extract three-dimensional point cloud spatial features;

[0064] S4: According to the three-dimensional point cloud spatial features, use the LSTM module to extract temporal features between multiple frames of point clouds;

[0065] S5: Through a fully connected layer, map the temporal features between multiple frames of point clouds extracted in step S4 to the label set through a non-linear transformation and output the recognition result.

[0066] In this embodiment, the signal after sampling by the FMCW millimeter-wave radar is processed according to the signal processing flow as Figure 2 shown, and three-dimensional point cloud information of human activities can be obtained. Therefore, step S1 specifically includes:

[0067] A1: Range FFT (one-dimensional FFT): Sample the intermediate frequency signal of the millimeter-wave radar and perform an FFT operation in the fast time dimension to calculate the corresponding frequency, and obtain the range information R of the human point cloud i , where i is the serial number of the point cloud;

[0068] A2: Velocity FFT (two-dimensional FFT): For the same target point, perform FFT operations on the linear frequency modulation signals of two adjacent periods respectively, that is, perform an FFT operation in the slow time dimension to obtain a two-dimensional range-Doppler map (RD map);

[0069] A3: Constant False Alarm Rate (CFAR) Detection: The CFAR algorithm is an adaptive target detection technology used to separate reflected signals from background noise. Multiple two-dimensional FFT matrices are summed to generate a pre-detection matrix, and then the CFAR algorithm is used to perform threshold decision on the pre-detection matrix to obtain multiple corresponding peaks;

[0070] A4: Angular FFT (Three-dimensional FFT): Perform FFT operations on the multiple peak points detected by the CFAR algorithm in the angular dimension (including azimuth dimension and elevation dimension), and calculate the azimuth angle θ of each point using the phase difference between antennas i and the elevation angle φ i ;

[0071] A5: According to the distance information R i , azimuth angle θ i and elevation angle φ i , convert the point cloud position from the polar coordinate system of the radar to the Cartesian coordinate system to obtain the three-dimensional coordinates (x i , y i , z i ) of the i-th point cloud:

[0072]

[0073] The embedded layer module structure in step S3 is as Figure 4 shown, which includes two LBR layers. In the figure, N in represents the number of input point clouds; The process of extracting three-dimensional point cloud spatial features using the mPCT module includes:

[0074] B1: Use an LBR layer (linear layer + batch normalization + ReLU) to increase the number of feature channels of the input point cloud from 3 to d in ;

[0075] B2: Use the Furthest Points Sample (FPS) algorithm to select N out central points, and obtain k nearest neighbor points of each central point through the K Nearest Neighbors (KNN) algorithm to obtain the features of the nearest neighbor points of each central point, and its data dimension is N out ×k×d in ; The operation of repeating the features of the central points k times is performed so that its data dimension is also N out ×k×d in ;

[0076] B3: Subtract the features of the nearest neighbor points from the features of the central points to obtain difference features, and then concatenate these difference features with the features of the central points;

[0077] B4: Through an LBR layer and a Max operation, the embedding layer outputs a tensor with a sampling point number of N out , and a feature channel number of d out ;

[0078] B5: As Figure 5 shown, 4 offset attention layers (OA layers) are connected in series in the offset attention module, and the input is concatenated with the outputs of all offset attention layers in the feature dimension. After the concatenated features are sent into the LBR layer, the output of the offset attention module is finally obtained, which is also the output of the mPCT module.

[0079] Referring to Figure 5 the lower part of

[0080] Considering the disordered characteristics of the point cloud, the Query (Q), Key (K), and Value (V) matrices in the module are all generated from the Input through a linear layer that shares weights for all point clouds. The specific calculation method is as follows:

[0081] (Q, K, V) = F in ·(W q , W k , W v )

[0082] Q, K ∈ R N×d / 4 , V ∈ R N×d

[0083] W q , W k ∈ R d×d / 4 , W v ∈ R d×d

[0084] Among them, F in represents the input feature, and W q , W k , W v represent learnable linear transformations that are shared for all point clouds. Among them, the output feature channel number of W q , W k is d / 4, and the output feature channel number of W v is d;

[0085] After obtaining the attention feature through the self-attention operation, the offset between the self-attention (SA) feature and the input feature is calculated through element-wise subtraction, and this offset feature is element-wise added to the input feature after the LBR layer; in the present invention, the offset attention layer is represented by OA (Offset-Attention), and LBR represents Linear layer + BatchNorm + ReLU. The output of the offset attention layer is Fout is:

[0086] F out = OA(F in ) = LBR(F in -F sa ) + F in

[0087] where, F sa represents the attention feature, and its calculation process is:

[0088] F sa = SL(QK T )·V

[0089] where, SL represents the softmax + l1norm operation.

[0090] For the entire bias attention module, assuming that the tensor shape output by the Embedding layer to the bias attention module is N×d, then the tensor shape after concatenation in the four concatenated bias attention layers is N×5d, and the output tensor shape of the final mPCT module after passing through the LBR layer is N×4d.

[0091] In step S4, the spatial features of each frame of point cloud aggregated by the mPCT module are sequentially input into the LSTM layer in chronological order to further extract the temporal features between multiple frames of point clouds. The uniqueness of LSTM lies in the introduction of "memory units" and "forget gates", which can selectively maintain the flow of information in long sequences. Therefore, LSTM can capture and understand complex dependencies in long sequences. The specific structure of LSTM is as Figure 6 shown. In the figure, Ft is the forget gate, It is the input gate, Ot is the output gate, is the candidate memory cell;

[0092] Assuming there are h hidden units, the batch size is n, the input vector dimension is d, the input X t ∈R d×h , the hidden layer state H t-1 ∈R n×h , then the specific operation process of the LSTM module is:

[0093]

[0094] where, W xf , W xi , W xc W xo ∈R d×h and W hf , W hi , W hc , Who ∈R h×h is a weight parameter, b f , b i , b c , b o ∈R 1×h is a bias parameter.

[0095] Example 2:

[0096] In this example, the method of the present invention is applied and analyzed based on three datasets, namely MMactivity, Pantomime, and mHomeGes. The confusion matrices of the mPCT-LSTM model tested on the three datasets are as Figure 7 shown, and the specific process is as follows:

[0097] Step 1: Point cloud data preprocessing

[0098] To make the point cloud data better adapt to the neural network classifier, the data format corresponding to each activity is uniformly adjusted. Different from the other two datasets, the MMActivity dataset has multiple repeated identical activities in the same piece of data. Therefore, each activity needs to be separated from the original data first. Using sliding window sampling, a time window of two seconds (60 frames) is taken, and the original data is intercepted with a step of 1 / 3 second (10 frames). The intercepted 2-second data is a complete activity. Then, all the point clouds of each activity in the three datasets are evenly divided into multiple aggregated frames, and the point clouds are randomly resampled within each aggregated frame. We will determine the number of aggregated frames and resampled points in the hyperparameter tuning experiment.

[0099] Step 2: Initialize the mPCT-LSTM model

[0100] The parameters in the embedding layer module are crucial for the model performance, which determines the amount of information of the features that the entire network can extract. Table 1 shows the key parameters in the model: the number of input features d in , the number of output features d out , the number of center points N out , the number of neighboring points k of each center point, the input size of the lstm layer, the lstm hidden layer size, and the size of the output fully connected layer.

[0101] Table 1 Model parameters

[0102]

[0103]

[0104] Step 3: Model training, validation, and testing

[0105] The mPCT-LSTM model is trained, validated, and tested on a server configured with an AMD 5600x processor and an NVIDIA RTX 2060 graphics card based on the Pytorch deep learning framework.

[0106] The dataset is divided into a training set, a validation set, and a test set in the ratio of 8:1:1. The number of iteration cycles is set to 1000, the learning rate is 0.001, and the Adam optimization algorithm and cross-entropy loss function are used. As can be seen from Table 2, mPCT-LSTM outperforms other baseline models in terms of the three metrics of accuracy, AUC, and AP on the three datasets. The accuracies reach 99%, 99.26%, and 94.14% respectively, which is on average 3.32% higher than the second-best model.

[0107] Table 2 Comparison with other models on the MMActivity, Pantomime, and HomeGes datasets

[0108]

[0109] To test the robustness of mPCT-LSTM, this embodiment also uses the Pantomime dataset to test the performance of mPCT-LSTM in an unfamiliar environment. The training set and the validation set are divided in the ratio of 8:2 in the Office and Open environments, and then the performance of the mPCT-LSTM model is tested in the Office&Open, Factory, Restaurant, and Multi people environments. It should be noted that in the Multi people environment, in the Open environment, there are three additional people interfering with the activity executor in the environment.

[0110] The performance of each model is shown in Table 3. The accuracy of mPCT-LSTM proposed in the present invention is better than that of other models in the three scenarios except Factory. The accuracies in Office&Open and Multi people reach 99.28% and 95.24% respectively, which are 3.53% and 1.91% higher than the second-best model respectively. However, the accuracies of mPCT-LSTM in the Factory and Restaurant environments are relatively low, 91.43% and 84.35% respectively. The reason is that this embodiment is trained in the Office and Open environments, so the model has not been trained in these two new environments. In addition, both of these environments have serious multipath interference problems, which pose higher requirements for the generalization of the model.

[0111] Table 3 Comparison of the accuracies of each model in the Factory, Restaurant, and Multi people environments

[0112]

[0113] Table 4 compares the spatial complexity and time complexity of the model proposed in the present invention and various baseline models. Since the time complexity of the model is related to the size of the n_chunks parameter, the models with the best performance on three datasets are taken, and the average value of their time complexities is calculated to be used as the computational complexity of the final model. The number of parameters of mPCT-LSTM is 0.34M, among which the number of parameters of the point cloud spatial feature extraction module mPCT is 0.14M, and the number of parameters of the temporal feature extraction module LSTM is 0.20M. Compared with the current lightest Tesla-V model, the mPCT-LSTM model proposed in the present invention reduces by 83%, 78%, and 78% respectively in terms of model size, time complexity, and number of parameters. Therefore, on the premise of ensuring the recognition accuracy, the model proposed in the present invention greatly reduces the memory requirement of the recognition network and is more suitable for the application of embedded devices.

[0114] Comparison of complexity indicators of each model in Table 4 with other models

[0115]

Claims

1. A lightweight human action recognition method based on millimeter-wave radar three-dimensional point cloud, characterized in that It includes the following steps: S1: Process the millimeter-wave radar sampling signal to obtain the three-dimensional point cloud information of human activities; S2: Construct a lightweight network, including an mPCT module and an LSTM module; S3: For the three-dimensional point cloud information, use the mPCT module to extract the three-dimensional point cloud spatial features; S4: According to the three-dimensional point cloud spatial features, use the LSTM module to extract the temporal features between multiple frames of point clouds; S5: Through the fully connected layer, map the temporal features between multiple frames of point clouds extracted in step S4 to the label set through non-linear transformation and output the recognition result.

2. The lightweight human action recognition method based on millimeter-wave radar three-dimensional point cloud according to claim 1, wherein The specific steps of step S1 include: A1: Distance FFT: Sample the intermediate frequency signal of the millimeter-wave radar, perform an FFT operation in the fast time dimension, calculate the corresponding frequency, and obtain the distance information R of the human body point cloud i , where i is the serial number of the point cloud; A2: Velocity FFT: For the same target point, perform FFT operations on the chirp signals of two adjacent periods respectively, that is, perform FFT operations in the slow time dimension to obtain a two-dimensional range-Doppler map; A3: Constant false alarm rate detection: Sum multiple two-dimensional FFT matrices to generate a pre-detection matrix, and then use the CFAR algorithm to perform threshold decision on the pre-detection matrix to obtain multiple corresponding peaks; A4: Angle FFT: Perform angle - dimension FFT operations on multiple peak points detected by the CFAR algorithm, and calculate the azimuth angle θ of each point using the phase difference between antennas i and the elevation angle φ i ; A5: According to the distance information R i , azimuth angle θ i and elevation angle φ i , convert the point cloud position from the polar coordinate system of the radar to the Cartesian coordinate system to obtain the three-dimensional coordinates (x i , y i , z i ) of the i-th point cloud.

3. A lightweight human action recognition method based on millimeter-wave radar three-dimensional point cloud according to claim 2, characterized in that, The three-dimensional coordinates (x i , y i , z i ) of the i-th point cloud in the step A5 are expressed as:

4. A lightweight human action recognition method based on millimeter-wave radar three-dimensional point cloud according to claim 1, characterized in that, In step S2, the mPCT module includes an embedding layer module and a bias attention module. The embedding layer module includes two LBR layers, and 4 bias attention layers (OA layer) are connected in series in the bias attention module.

5. A lightweight human action recognition method based on millimeter-wave radar three-dimensional point cloud according to claim 4, characterized in that, The process of using the mPCT module to extract the three-dimensional point cloud spatial features in step S3 includes: B1: Use an LBR layer to increase the number of feature channels of the input point cloud from 3 to d in ; B2: Select N center points using the farthest point sampling algorithm, obtain k nearest neighbor points for each center point through the k-nearest neighbor algorithm, and obtain the features of the nearest neighbor points of each center point. The data dimension is N out × k × d out ; Repeat the operation on the features of the center points k times, so that the data dimension is also N in × k × d out ; in ​ B3: Subtract the features of the neighboring points from the features of the central point to obtain the difference features, and then splice these difference features with the features of the central point; B4: Through an LBR layer and a Max operation, the embedding layer outputs a tensor with a sampling point number of N out , and a feature channel number of d out ; B5: The bias attention module splices the input with the outputs of all bias attention layers in the feature dimension. After sending the spliced features into the LBR layer, the output of the bias attention module is finally obtained, which is also the output of the mPCT module.

6. A lightweight human action recognition method based on millimeter-wave radar three-dimensional point cloud according to claim 5, characterized in that The operation of the bias attention layer in step B5 includes: Considering the disordered characteristics of the point cloud, the Query (Q), Key (K), and Value (V) matrices in the module are all generated from the Input through a linear layer that shares weights for all point clouds. The specific calculation method is as follows: (Q, K, V) = F in ·(W q , W k , W v ) Q, K ∈ R N×d / 4 , V ∈ R N×d W q ,W k ∈R d×d / 4 ,W v ∈R d×d Among them, F in represents the input feature, W q , W k , W v represents a learnable linear transformation shared for all point clouds, where W q , W k outputs feature channels with a number of d / 4, and W v outputs feature channels with a number of d; After obtaining the attention features through the self-attention operation, the offset between the self-attention features and the input features is calculated by element-wise subtraction, and this offset feature is element-wise added to the input features after the LBR layer; let OA denote the offset attention layer, LBR denote Linear layer + BatchNorm + ReLU, and the output F of the offset attention layer out is as follows: F out = OA(F in ) = LBR(F in -F sa ) + F in Among them, F sa represents the attention feature, and its calculation process is as follows: F sa = SL(QK T )·V where SL represents the softmax + l1norm operation.

7. A lightweight human action recognition method based on millimeter-wave radar three-dimensional point cloud according to claim 5, characterized in that For the entire bias attention module in step B5, assuming that the tensor shape output by the Embedding layer to the bias attention module is N×d, then the tensor shape after splicing in the four series-connected bias attention layers is N×5d, and the final output tensor shape of the mPCT module after passing through the LBR layer is N×4d.

8. A lightweight human action recognition method based on millimeter-wave radar three-dimensional point cloud according to claim 1, characterized in that, In the step S4, the LSTM module includes a forgetting gate F t , an input gate I t , an output gate O t , a candidate memory cell Suppose there are h hidden units, the batch size is n, the dimension of the input vector is d, and the input is X t ∈R d×h , the hidden layer state H t-1 ∈R n×h , then the specific operation process of the LSTM module is as follows: Among them, W xf ,W xi ,W xc W xo ∈R d×h and W hf ,W hi ,W hc ,W ho ∈R h×h are weight parameters, b f ,b i ,b c ,b o ∈R 1×h is the bias parameter.