A method for human behavior recognition using millimeter-wave radar
By generating micro Doppler spectrum and using Vision Transformer network model and combining the multi-head attention mechanism, the problem of time-consuming and labor-intensive feature extraction and insufficient generalization ability in the existing technology is solved, and efficient and accurate radar human behavior recognition is achieved.
Patent Information
- Application Number
- CN202211461976.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-11-21
AI Technical Summary
The existing radar-based human behavior recognition technology has the problem of time-consuming and laborious feature extraction and insufficient generalization ability. Deep CNN ignores the timing characteristics of radar signals, resulting in low recognition accuracy.
Short-time Fourier transform is used to generate micro Doppler spectrograms (MDM), combined with Vision Transformer (VIT) network model, micro Doppler features are extracted through multi-head attention mechanism, 1-dimensional position coding and time slice segmentation are used to focus on time features, and training data sets are constructed and human behavior recognition is performed.
It improves the accuracy of human behavior recognition and the generalization ability of the network, and achieves efficient and accurate real-time recognition.
Smart Images

Figure CN115902878B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human behavior recognition, and specifically to a millimeter-wave radar human behavior recognition method. Background Art
[0002] Human Activity Recognition (HAR) has very broad application prospects in many fields such as security warning and anomaly detection, which require rapid and accurate recognition of human behaviors. At the same time, with the popularization and development of artificial intelligence, the flexibility and accuracy of this technology have been further improved. HAR based on cameras has privacy leakage problems and is not suitable for dark or other low-light scenarios. While HAR based on radar has attracted wide attention due to its advantages such as all-day and all-weather operation, strong privacy, and adaptability to any lighting conditions. In particular, Frequency Modulated Continuous Wave (FMCW) radar has more prominent performance in HAR due to its advantages of high resolution, high detection accuracy, small size, and low cost.
[0003] The use of micro-Doppler features (MDF) generated by target motion for target recognition in radar has been widely applied. In HAR based on radar micro-Doppler characteristics, the Doppler provided by the movement of the human torso has very small characteristic differences, and our more distinguishable features mainly come from the micro-Doppler generated by the limb swings. Therefore, MDF is the basis for identifying different human behaviors. MDF is a feature in which the micro-Doppler frequency shift generated by human motion changes over time and can be observed in the joint time-frequency domain. How to extract effective MDF is a core issue in the HAR task. In the early stage, researchers needed to rely on prior knowledge to manually extract MDF, or extract MDF through Empirical Mode Decomposition (EMD), and then combine traditional Machine Learning (ML) to achieve HAR. However, the features extracted in this way are not only time-consuming and laborious, but also only suitable for a specific task, which limits the generalization ability of the network.
[0004] In recent years, the rapid development of deep learning (DL) has brought new research ideas to radar-based HAR. Compared with traditional ML, DL does not rely on manual experience to extract features. It can automatically learn the inherent features of data, with excellent performance and strong generalization ability. The typical DL model convolutional neural network (CNN) and its various deep variant forms have been widely applied in HAR tasks. References first applied deep CNN to HAR and achieved good classification results. And in HAR, multi-dimensional CNN is used to extract different features for fusion to improve the recognition accuracy. In addition, shallow CNN networks can reduce the network complexity and improve the network operation efficiency. For example, a two-layer CNN is only designed to detect human movements. This network has few parameters and low requirements for computing power, but the recognition effect for some behaviors is poor. The models mentioned above regard the representation map of human behaviors as a kind of visual image and extract the spatial correlation features between pixels from it through deep networks. However, the original human echo signal collected by radar is a kind of time-series data, and its MDF has spatial and temporal correlations. The deep multi-dimensional CNN ignores the temporal dependence relationship between time series, which limits the recognition accuracy. Therefore, we propose a method for human behavior recognition using millimeter-wave radar. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] In view of the deficiencies of the prior art, the present invention provides a method for human behavior recognition using millimeter-wave radar, which solves the above problems.
[0007] (2) Technical Solutions
[0008] To achieve the above object, the present invention provides the following technical solutions: A method for human behavior recognition using millimeter-wave radar, comprising the following steps:
[0009] The first step: Preprocess the collected radar human echo signal to generate an MDM map, and extract
[0010] The second step: Process through short-time Fourier transform to generate MDM, and use the generated MDM to construct a human behavior data set, including a training data set Otrain and a test data set Otest;
[0011] The third step: Use the VIT network model based on time-slice segmentation to extract MDF features from the constructed MDM data set to achieve HAR.
[0012] Preferably, the preprocessing in the first step includes the following steps:
[0013] S1: The radar receives the signals reflected by the human body and the surrounding environment, mixes them with the transmitted signal, and the output result is an intermediate-frequency signal;
[0014] S2: Perform range compression and angle compression using 2D FFT, concentrate the energy of each frame of signal on the position and angle grids of the target, and generate several two-dimensional matrices in the slow-time dimension.
[0015] S3: Background clutter suppression. Use the phasor mean cancellation algorithm to remove most of the strong stationary clutter without affecting the human motion signal. The sequence of two-dimensional matrices after clutter removal is
[0016] S4: Use the Moving Average - Ordered Statistic Constant False Alarm Rate (MAOS - CFAR) algorithm to lock the human target motion grid. Extract the vector of the grid where the target is located along the slow-time dimension.
[0017] Preferably, the intermediate frequency signal is:
[0018]
[0019] where f c is the radar carrier frequency, B is the signal bandwidth, T is the signal duration, R is the distance between the radar and the target, V is the radial velocity of the target, and c represents the speed of light. n = 0, 1, 2,..., N - 1 is the number of sampling points on a single chirp, l = 0, 1, 2,..., L - 1 is the number of chirps in a single frame, and t represents the current t-th frame.
[0020] Preferably, the specific steps of the third step are as follows:
[0021] S1: Divide the MDM into many tiles with the same time step in chronological order, and there is no overlap between each tile. Each tile can be regarded as a one-dimensional micro - Doppler time series. Vectorize these time series respectively, extract the position relationship of adjacent time series in the time dimension, and retain the complete time information of the MDM;
[0022] S2: First, use a one - dimensional CNN to generate the feature sequence of each slice sequence, and then, together with the token sequence for classification, add position information to the slice sequence using one - dimensional position encoding;
[0023] S3: Input the feature sequence into the normalization module to generate three trainable variables Q, K, V, and input these three parameters into the multi - head attention mechanism for training;
[0024] S4: Train the VIT model with all training samples according to S1 - S3, and use the trained model for the Otest sample data to test the final recognition effect of HAR.
[0025] Preferably, the specific steps of S3 are as follows:
[0026]
[0027] MultiheadAttention(Q, K, V) = Concat(head1...head i )W O ;
[0028] head i = Attention(Q, K, V);
[0029] Multi-head attention maps Q and K to multiple different subspaces in the original high-dimensional space to calculate similarities, which requires Q-K-v scaling. This scaling step is completed by multiple linear modules on the input of the encoding block. The number of linear modules is determined by the number of heads. The scaled Q-K-V is grouped and fed into multiple scaled dot attention modules to extract the time series features of the Doppler sequence respectively. Finally, the feature information output by different attention networks is synthesized using Concat. The essence of multi-head attention is to discover the feature correlation between Doppler sequences at different scales by fusing attention information from different distributed subspaces. This spatial decomposition and re-synthesis can reduce the vector dimension during each attention head calculation and prevent overfitting to a certain extent.
[0030] (III) Advantageous Effects
[0031] Compared with the prior art, the present invention provides a millimeter-wave radar human behavior recognition method, which has the following advantageous effects:
[0032] 1. This millimeter-wave radar human behavior recognition method divides the MDM image by column instead of by block, following the characteristic that the MDM picture reflects the time-varying change of the micro-Doppler component. The picture slices are given position encoding through 1D position encoding to fully extract the spatio-temporal features of the time series.
[0033] 2. This millimeter-wave radar human behavior recognition method uses the multi-head attention mechanism to focus on more important time features, ensuring the effective utilization of all times and effectively improving the recognition accuracy of human behaviors. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a schematic flow diagram of this method;
[0035] Figure 2 is a schematic diagram of extracting MDF by the Vision Transformer network model;
[0036] Figure 3 is a schematic flow diagram of the multi-head self-attention mechanism calculation;
[0037] Figure 4 For the micro-Doppler spectrogram and characteristic heat map during walking;
[0038] Figure 5 For the micro-Doppler spectrogram and characteristic heat map during running;
[0039] Figure 6 For the micro-Doppler spectrogram and characteristic heat map during squatting and standing up;
[0040] Figure 7 For the micro-Doppler spectrogram and characteristic heat map during bowing;
[0041] Figure 8 For the micro-Doppler spectrogram and characteristic heat map during turning around;
[0042] Figure 9 For the schematic diagram of the comparison results of different model performances;
[0043] Figure 10 For the schematic diagram of the comparison of the recognition effects of block segmentation and column segmentation. Detailed implementation manners
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0045] This solution combines the micro-Doppler characteristics in a frequency-modulated continuous-wave radar and the Vision Transformer (VIT) network model to propose an efficient HAR method. The specific flowchart is as Figure 1 shown. By preprocessing the original human body signals of the radar, a clear micro-Doppler spectrogram (MDM) that can represent MDF is generated. The MDM is used as the input of the VIT model. The MDM is divided into sequences with the same time step. 1D pos-embedding is used to add position encoding and extract the spatial features of the time series. The weight allocation mechanism of the Multi-head-Attention can focus on more important time features, ensure the effective utilization of all times, improve the efficiency of the network in processing information, and input the extracted MDF into the MLP classifier to complete the recognition of human behaviors. This solution has the advantages of stable network performance, high transferability, high recognition accuracy, etc., and can meet the requirements of accurate and real-time HAR.
[0046] The HAR method of this solution includes the following steps:
[0047] Step 1. Preprocess the collected radar human echo signals to generate an MDM map. The preprocessing process includes 2D FFT, background clutter suppression, and constant false alarm rate (CFAR) detection.
[0048] Step 1.1. The radar receives the signals reflected by the human body and the surrounding environment, mixes them with the transmitted signal, and the output result is an intermediate frequency signal. Considering the case where the radar transmits multiple frames of signals, the intermediate frequency signal can be expressed as:
[0049]
[0050] where f c is the radar carrier frequency, B is the signal bandwidth, T is the signal duration, R is the distance between the radar and the target, V is the radial velocity of the target, and c represents the speed of light. n = 0, 1, 2,..., N - 1 is the number of sampling points on a single chirp, l = 0, 1, 2,..., L - 1 is the number of chirps in a single frame, and t represents the current t-th frame.
[0051] Step 1.2. Use 2D FFT for range compression and angle compression to concentrate the energy of each frame of signals on the position and angle grids of the target, generating several two-dimensional matrices in the slow time dimension
[0052] Step 1.3. To avoid the human motion signal being masked in a clutter environment with strong interference such as the background, the phasor mean cancellation algorithm is used to remove most of the strong stationary clutter without affecting the human motion signal. The sequence of two-dimensional matrices after clutter removal is
[0053] Step 1.4. To lock the two-dimensional grid where the target is located, the moving average - ordered statistic constant false alarm rate (MAOS - CFAR) algorithm is used to lock the human target motion grid. Extract the vector of the grid where the target is located along the slow time dimension
[0054] Step 2. Process through short-time Fourier transform to generate MDM. Use the generated MDM to construct a human behavior dataset, including a training dataset Otrain and a test dataset Otest.
[0055] Step 3. Use the VIT network model to extract MDF features from the MDM dataset constructed in Step 2 to achieve HAR, and the whole process is as Figure 2 shown.
[0056] Step 3.1: Divide the MDM into many tiles with the same time step in chronological order, without overlap between each tile. Each tile can be regarded as a one-dimensional micro-Doppler time series. Vectorize these time series respectively, extract the positional relationship of adjacent time series in the time dimension, and retain the complete time information of the MDM.
[0057] Step 3.2: First, use a one-dimensional CNN to generate the feature sequence of each slice sequence, and then, together with the token sequence for classification, add positional information to the slice sequence using one-dimensional positional encoding.
[0058] Step 3.3: Pass the feature sequence into a normalization module to generate three trainable variables Q, K, and V. Pass these three parameters into the multi-head attention mechanism for training. The calculation process is as Figure 3 .
[0059]
[0060] Multi-head-Attention(Q,K,V)=Concat(head1...head i )W O ;
[0061] head i =Attention(Q,K,V);
[0062] The multi-head attention maps Q and K to multiple different subspaces in the original high-dimensional space to calculate the similarity, which requires Q-K-V scaling. This scaling step is completed by multiple linear modules on the input of the encoding block. The number of linear modules is determined by the number of heads. Group the ratios of Q-K-V and send them into multiple ratio point attention modules to extract the time series features of the Doppler sequence respectively. Finally, use Concat to synthesize the feature information output by different attention networks. The essence of multi-head attention is to discover the feature correlation between Doppler sequences at different scales by fusing the attention information from different distributed subspaces. This spatial decomposition and re-synthesis can reduce the vector dimension during each attention head calculation and prevent overfitting to a certain extent.
[0063] Step 3.4: Train the VIT model with all the training samples of Otrain according to Steps 3.1 - 3.4. The trained model is used for the Otest sample data to test the final recognition effect of HAR.
[0064] The method will be described in detail in combination with experiments. The experiments apply this method to actual HAR to verify the performance of this method. The specific implementation details are as follows:
[0065] The experiment uses the TI AWR1843 radar to collect human behavior data. The radar parameters are set as follows: carrier frequency 77 GHz, transmitting a sawtooth frequency-modulated continuous wave, number of sampling points 128, Doppler resolution 0.05 m / s, and the acquisition duration for a single behavior is 5 s. To achieve a sufficient detection area, the radar is installed at a height of 1.5 m, and the experimental personnel are 3 - 4 m away from the radar. Five common human behaviors are designed in the experiment, namely (1) walking, (2) running, (3) squatting and standing up, (4) bending over, and (5) turning around. To make the data generalizable, 7 males and 3 females are selected to participate in data collection, and they have different heights, weights, and ages. 10 experimental personnel repeat each behavior 20 times, totaling 1000 (5×10×20) groups of human behavior data. When collecting experimental data, only a single experimental personnel scenario is considered, and the standardization of behaviors is not strictly restricted and can be performed according to personal habits to ensure the diversity of the dataset. The processing of the radar raw data is implemented in MATLAB v2017. The output size of the 2D FFT is 134×256. One frame of MDM is generated every 50 ms, and the total accumulation time of 100 frames of MDM is 5 s. The MDMs of the five types of human behaviors are as Figure 9 shown. Figures 4 - 8 The results show that different human behaviors have unique MDFs, which are the basis for human behavior recognition. The MDM is converted into a grayscale image and the image size is reshaped to 112×112. After reshaping, the original detailed features of the MDM are not lost, and at the same time, the computational amount of the CLA hybrid multi-network model can be reduced. To ensure that there are sufficient samples in the dataset, data augmentation is also used. The total amount of data finally generated is 4950, which is divided into Otrain and Otest (3465:1485) in an 8:2 ratio. The VIT model is trained under the DL framework of pytorch v1.11.0 and python3.9. The network parameters are set as follows: time step 112, the MDM is divided into tiles with the same time step, and the size of each tile is 1×112. After going through all time steps, the input of the MDM is completed. The training model uses the Adam optimizer, and the learning rate, number of training epochs, and number of batch samples per iteration are 0.001, 40, and 200 respectively. At the same time, the experiment uses the Dropout method to turn off some neurons in the network to avoid the problem that the generalization ability of the model is reduced due to overfitting.
[0066] Figures 4 - 8 The feature heatmap shows that this method can focus the attention on the micro-Doppler region representing limb movements. This method is compared and analyzed with several common HAR models, and the experimental results are as Figure 9 shown. It can be seen from this that the average accuracy of this method can reach more than 99%, which is the best among all models.
[0067] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for human behavior recognition using millimeter-wave radar, characterized in that, Including the following steps: Step 1: Preprocess the collected radar human echo signal to generate an MDM map and extract it ; Step 2: Process through short-time Fourier transform to generate MDM, and use the generated MDM to construct a human behavior dataset, including a training dataset Otrain and a test dataset Otest; The third step: Use the VIT network model based on time-slice segmentation to extract MDF features from the constructed MDM dataset to achieve HAR, where HAR is human recognition behavior; The specific steps of the third step are as follows: S1: Divide the MDM into many tiles with the same time step in chronological order. There is no overlap between each tile, and each tile is regarded as a one-dimensional micro-Doppler time series. Vectorize these time series respectively, and extract the position relationship of adjacent time series in the time dimension, retaining the complete time information of the MDM; S2: First, use a one-dimensional CNN to generate the feature sequence of each slice sequence, and then, together with the token sequence for classification, use one-dimensional position encoding to add position information to the slice sequence; S3: Input the feature sequence into the normalization module to generate three trainable variables Q, K, and V, and input these three parameters into the multi-head attention mechanism for training; S4: Train the VIT model with all training samples according to S1-S3, and use the trained model for Otest sample data to test the final recognition effect of HAR; The specific steps of S3 are as follows: Attention maps Q and K to the same high-dimensional space to calculate similarity. Multi-head attention maps Q and K to different subspaces of the high-dimensional space to calculate similarity. Multi-head attention maps the same Q, K, and V to different subspaces of the original high-dimensional space to perform attention calculation while keeping the total number of parameters unchanged, and then combines the attention information in different subspaces in the last step.
2. The millimeter-wave radar human behavior recognition method according to claim 1, characterized in that: The preprocessing in the first step includes the following steps: S1: The radar receives the signals reflected by the human body and the surrounding environment, mixes them with the transmitted signal, and the output result is an intermediate-frequency signal; S2: Perform range compression and angle compression using 2D FFT, concentrate the energy of each frame of signal on the position and angle grids of the target, and generate several two-dimensional matrices in the slow time dimension ; S3: Background clutter suppression. The phasor mean cancellation algorithm is used to remove most of the strong stationary clutter without affecting the human motion signal. The two-dimensional matrix sequence after clutter removal is ; S4: The Moving Average Ordered Statistics Constant False Alarm Rate (MAOS-CFAR) algorithm is used to lock the motion grid of the human target, and the vector of the grid where the target is located along the slow time dimension is extracted. 。 3. The millimeter-wave radar human behavior recognition method according to claim 2, characterized in that: The intermediate-frequency signal is: Wherein, is the radar carrier frequency, B is the signal bandwidth, T is the signal duration, R is the distance between the radar and the target, V is the radial velocity of the target, c represents the speed of light, n = 0, 1, 2, ..., N-1 is the number of sampling points on a single chirp, l = 0, 1, 2, ..., L-1 is the number of chirps in a single frame, and t represents the current t-th frame.