Continuous human body action recognition method based on millimeter wave radar

The millimeter wave radar-based RA-TCN framework with attention mechanisms addresses the challenges of continuous human action recognition, providing accurate and reliable detection in real-world environments.

CN120314930APending Publication Date: 2025-07-15ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510371585.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify continuous human movements, especially in emergencies such as falls in the elderly, and existing equipment such as wearable sensors and cameras have convenience and privacy issues.

Method used

Millimeter wave radar signal processing and double-expanded one-dimensional time convolution network framework are used, combined with attention mechanism, continuous human movements are identified through sequence-to-sequence mode, spatial characteristics and short-term dependence information are extracted, and sequence prediction results are output.

Benefits of technology

It realizes efficient and accurate identification of continuous human movements without invading privacy, improves identification accuracy and computing efficiency, and is suitable for low-cost portable devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120314930A_ABST
    Figure CN120314930A_ABST
Patent Text Reader

Abstract

The invention discloses a continuous human body action recognition method based on a millimeter-wave radar. The method comprises the following steps: processing a millimeter-wave radar signal; the millimeter-wave radar collects a continuous human body action data set; providing a double-expansion one-dimensional time convolutional network framework, adding an attention mechanism, extracting spatial features and time long-term and short-term dependence information from continuous HAR sequence data through parallel processing, and outputting a sequence prediction result; and an evaluation index STT Score of radar continuous human body action recognition is established. The portable wearable collecting device has the advantages that the portable wearable collecting device is low in cost, and a subject can use the portable wearable collecting device anytime and anywhere. The assessment method facilitates assessment of the rehabilitation condition of the patient from the perspective of quantitative data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for continuous human action recognition, and more particularly to a method for continuous human action recognition based on millimeter-wave radar. Background Art

[0002] Nowadays, with the increasing aging of the global population and the continuous growth of the elderly population, the demand for indoor short-distance human action recognition technology is rising day by day. The ability to accurately recognize daily activities and abnormal events is crucial for timely emergency responses, such as when the elderly accidentally fall.

[0003] Wearable sensors monitor human movement by collecting acceleration data, but their effectiveness is limited by the user's willingness to wear continuously and the limited battery life, which undoubtedly increases the inconvenience of use. Therefore, non-contact sensors are gradually favored in the field of indoor human action recognition due to their convenience and practicality.

[0004] Due to the possible influence of environmental light changes and privacy protection issues on cameras, millimeter-wave radar sensors provide a promising solution. These sensors only capture key physical parameters - distance, speed, and angle, thus enabling accurate recognition of human actions without infringing on personal privacy.

[0005] In real-world scenarios, human actions are a continuous series of dynamic processes, with one action following another, and the exact duration and transition phases between them are often unpredictable. For this reason, the demand for accurately capturing and recognizing ongoing continuous human actions has become even more urgent. Summary of the Invention

[0006] The main technical problem solved by the present invention is the detection of indoor continuous human actions. A method for continuous human action recognition based on millimeter-wave radar is proposed. By using a non-contact FMCW radar to identify continuous human actions in a sequence-to-sequence manner, its main contents include:

[0007] A method for continuous human action recognition based on millimeter-wave radar, including: millimeter-wave radar signal processing; collecting a continuous human action dataset by millimeter-wave radar; proposing a dual dilated one-dimensional temporal convolutional network framework, adding an attention mechanism, and extracting spatial features and temporal long-term and short-term dependency information through parallel processing of the sequence data of continuous HAR, and outputting a sequence prediction result.

[0008] The millimeter-wave radar signal processing includes: configuring millimeter-wave radar signal acquisition parameters; preprocessing millimeter-wave radar signals.

[0009] A dual dilated one-dimensional time convolutional network framework is proposed, which includes three core modules: Dual Dilated Module, Attention Module, and Prediction Refinement Module, to process the time series data generated by radar. First, the Dual Dilated Module expands the receptive field through one-dimensional dilated convolutions with two different dilation factors in parallel, adjusts the channels using 1×1 convolutions, and utilizes a combination of exponentially increasing and decaying dilation factors to evenly expand features at different levels. Meanwhile, ReLU activation, Dropout, and residual connections are combined to stabilize the training and prevent gradient vanishing. Next, the Attention Module introduces the multi-head self-attention mechanism (MHSA), calculates the frame-level relationships through Query (Q), Key (K), and Value (V), and multiple attention heads capture global temporal information in parallel, enhancing the attention to key frames and improving the accuracy of frame-level classification. Finally, the Prediction Refinement Module further optimizes the preliminary prediction results, gradually refines the features using dilated convolutions, 1×1 convolutions, ReLU activation, and Dropout, and outputs the final frame-level classification results through Softmax. This framework combines the time series modeling ability of TCN and the global information capture ability of the self-attention mechanism, optimizes the modeling of long-term dependencies while ensuring computational efficiency, reduces the number of parameters and avoids overfitting, thereby improving the accuracy of human action recognition from radar data.

[0010] The added attention mechanism includes the self-attention mechanism (Self-Attention) and the multi-head self-attention mechanism (Multi-Head Self-Attention, MHSA), which are used to enhance the global information capture ability of frame-level features and improve the attention to key frames. First, in the self-attention mechanism, the input sequence generates Query (Q), Key (K), and Value (V) through three learnable weight transformation matrices, and obtains the global dependency information at each time step through the calculation of attention weights, thereby enhancing the modeling ability for long-term dependencies. Secondly, to further improve the modeling effect, the multi-head self-attention mechanism is introduced, and features in different subspaces are calculated in parallel through multiple attention heads (head1, head2, …, head i ). Each attention head independently learns different context information, and finally the final multi-head output is obtained through concatenation and linear transformation. This attention mechanism can effectively alleviate the problem of losing fine-grained information during deep feature extraction in the dilated convolutional layer, enhance the feature expression ability in time series data, make the model more accurate when processing similar frames, and improve the accuracy of frame-level prediction.

[0011] The sequence data of continuous HAR is processed in parallel to extract spatial features and temporal long - short - term dependence information, including: spatial feature information such as the target distance and speed collected by the FMCW radar, and temporal long - short - term dependence information such as the duration and transition time of activities, key frame information of action transitions, and global features. Among them, the spatial features reflect the relative position and dynamic changes between the target and the radar, while the temporal long - short - term dependence information helps the model understand the temporal structure and trend of activities, including short - term action transitions and long - term overall activity patterns. These information jointly support the accurate recognition of continuous human activities.

[0012] The output sequence prediction results include: the prediction of human activity types corresponding to each frame of data, which implies the duration information of activities and particularly focuses on capturing the key frames of activity transitions. These prediction results are given in the form of a sequence, maintaining the continuity and temporal order of activities, and achieving high - accuracy predictions through the proposed RA - TCN network.

[0013] Establish the evaluation index STT Score for radar - based continuous human action recognition;

[0014] The evaluation index STT Score for radar - based continuous human action recognition allows for a certain tolerance of the error in activity recognition within an extremely short time interval, thus more accurately reflecting the continuity and transitional ambiguity of human actions.

[0015] The segment - based F1 - score is generally used for action segmentation and detection tasks. This method evaluates the performance by calculating the IOU between each predicted action and the ground - truth action. The calculation of IOU can be expressed as:

[0016]

[0017] where P represents the predicted action time interval and T represents the ground - truth action time interval.

[0018] The segment - based method compares the IOU with the threshold K to determine true positives (TP), false positives (FP), and false negatives (FN). P represents the prediction result, GT represents the ground - truth label, and the calculation of F1 can be expressed as:

[0019]

[0020] If the segment - based prediction is correct, when the IOU is greater than or equal to the threshold K, it is labeled as TP, and less than K is FP. If the segment - based prediction is incorrect, it is labeled as FN. Additionally, when there are multiple correct predictions within the range of a single ground - truth action, only one will be classified as TP, and the others are FP.

[0021] Then calculate the precision and recall rates of the sum of TP, FP, and FN for all classes, and calculate the segment - based F1 - score.

[0022]

[0023] Based on F1, STT Score is proposed for continuous human action recognition by radar. len p Indicates the frame length of the prediction time. When counting TP, FP, and FN, if the comparison between the prediction result of the action transition frame and the ground truth is within 5 frames, it is not included in the calculation of FN. STT Score can be expressed as:

[0024]

[0025] Furthermore, the millimeter-wave radar collects a continuous human action dataset, including: radar data and synchronized video data of 5 volunteers (4 males and 1 female, aged 21 - 25 years) performing 8 daily activities in a real living room environment. The experimental site ranges approximately 3.5m × 7m, the radar system is installed at a height of 1m, and the subjects need to complete the preset activity combinations within the specified area. The activities are divided into two categories: Category A (main activities, drinking water, falling, picking up an object) and Category B (secondary activities, walking, staying in place, sitting on the sofa, standing up from the sofa, standing up from the ground), and 3 continuous human activity combination sequences are designed. Each subject completes the activities within 45 - 60s, and Category B activities can be freely interspersed before and after performing Category A activities, and there are no restrictions on the start, end, and transition methods of the actions. The same combination is repeated 4 times, a total of 60 groups of radar data are collected, the radar samples 25 frames per second, each group of data contains 1125 - 1500 frames, and a total of 78750 frames of radar data are obtained. This dataset provides radar time-series signals, synchronized video information, and continuous human activity sequences, which can be used for research such as behavior recognition and pattern analysis.

[0026] The advantages of the present invention are as follows:

[0027] 1. A low-cost portable wearable acquisition device that can be used by subjects anytime and anywhere.

[0028] 2. An evaluation method that facilitates the assessment of the rehabilitation status of patients from the perspective of quantitative data. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is an architecture diagram of a method for segmenting and detecting continuous human activities in an indoor environment using an FMCW radar;

[0030] Figure 2 is an architecture diagram of a method for continuous human action recognition based on a millimeter-wave radar;

[0031] Figure 3 is a parameter configuration diagram of an FMCW radar;

[0032] Figure 4 It is the flowchart of radar signal preprocessing;

[0033] Figure 5 It is the form diagram of a frame of FMCW radar data;

[0034] Figure 6 It is the diagram of 8 daily activities;

[0035] Figure 7 It is the diagram of activity sequence combination;

[0036] Figure 8 It is the diagram of the number of annotated frames for different activities;

[0037] Figure 9 It is the diagram of the sequence-to-sequence network framework;

[0038] Figure 10 It is the diagram of the double dilation model architecture;

[0039] Figure 11 The dilation factor is 2 l Overview diagram of the three-layer dilated convolutional network when;

[0040] Figure 12 The dilation factor is 2 L-l Overview diagram of the three-layer dilated convolutional network when;

[0041] Figure 13 It is the diagram of the attention model architecture;

[0042] Figure 14 It is the diagram of the multi-head self-attention mechanism architecture;

[0043] Figure 15 It is the diagram of the prediction optimization module architecture;

[0044] Figure 16 It is the diagram of the structure of the RA-TCN model and the output of each layer;

[0045] Figure 17 It is the example diagram of piecewise evaluation when k = 0.5;

[0046] Figure 18 It is the structure block diagram of (Bi-)LSTM and (Bi-)GRU networks;

[0047] Figure 19 It is the diagram of the size and attributes of (Bi-)LSTM and (Bi-)GRU networks;

[0048] Figure 20 It is the diagram of experimental results;

[0049] Figure 21 It is the example diagram of continuous activity recognition time for different networks. Detailed implementation manners

[0050] The present application will be further described in detail below with reference to the accompanying drawings.

[0051] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0052] Glossary:

[0053] Frequency Modulated Continuous Wave: Frequency Modulated Continuous Wave

[0054] Temporal Convolutional Network: Temporal Convolutional Network

[0055] DT: Doppler - Time Series

[0056] Dual Dilated Module: Dual Dilated Module

[0057] Attention Module: Attention Module

[0058] Predictio Refinement Module: Prediction Refinement Module

[0059] IOU: Intersection over Union

[0060] STT Score: Short - Term Tolerance Score

[0061] Acc: Frame - by - Frame Accuracy

[0062] Range FFT: Range Fast Fourier Transform

[0063] Doppler FFT: Distance Slow Fourier Transform

[0064] The main technical problem to be solved by the present invention is to provide a method for continuous human action recognition based on millimeter-wave radar. First, a dual dilated one-dimensional temporal convolutional network with attention mechanism, namely RA-TCN, is designed. This network uses dilated convolution to expand the receptive field and uses the attention mechanism to capture the information features of key frames. Each frame of data is processed in a sequence-to-sequence manner to achieve the segmentation and recognition of activities. The present invention also establishes a dataset of indoor human continuous activities, including 60 groups of data of 8 different activities, with a total duration of 52.5 minutes. Finally, different network models are used to conduct comparative experiments on this dataset. The RA-TCN network is evaluated through existing evaluation schemes and the STT method proposed by this method, verifying its feasibility in processing continuous human activities. The F1 score and STT score are 0.6872 and 0.8530 respectively, and it is very excellent in terms of inference time and training time.

[0065] Reference Figure 1 , the present invention is mainly divided into two parts: a method for continuous human action recognition based on millimeter-wave radar and the establishment of an evaluation index STT Score for radar continuous human action recognition.

[0066] The first part: A method for continuous human action recognition based on millimeter-wave radar, reference Figure 2 , including: millimeter-wave radar signal processing; collecting a continuous human action dataset by millimeter-wave radar; proposing a dual dilated one-dimensional temporal convolutional network framework; adding an attention mechanism; extracting spatial features and temporal long-term and short-term dependence information by parallel processing of continuous HAR sequence data; and outputting sequence prediction results.

[0067] Furthermore, the millimeter-wave radar signal processing includes: millimeter-wave radar signal acquisition parameter configuration; millimeter-wave radar signal preprocessing.

[0068] Furthermore, the extracting of spatial features and temporal long-term and short-term dependence information by parallel processing of continuous HAR sequence data includes: spatial feature information such as target distance and speed collected by FMCW radar, as well as temporal long-term and short-term dependence information, such as the duration and transition time of activities, key frame information of action transitions, and global features. Among them, the spatial features reflect the relative position and dynamic changes between the target and the radar, while the temporal long-term and short-term dependence information helps the model understand the temporal structure and trend of activities, including short-term action transitions and long-term overall activity patterns. These information jointly support the accurate recognition of continuous human activities.

[0069] Furthermore, the predicted output sequence includes: the prediction of the human activity type corresponding to each frame of data, which implies the duration information of the activity and particularly focuses on capturing the key frames of activity transitions. These prediction results are given in the form of a sequence, maintaining the continuity and chronological order of the activity, and achieving high-accuracy prediction through the proposed RA-TCN network.

[0070] Specifically, the millimeter-wave radar signal acquisition parameters refer to Figure 3 . Since collecting continuous human movements requires finer range resolution and higher signal-to-noise ratio, the design increases the chirps or chirp duration in one frame, increases the frequency slope, and designs a bandwidth of 1499 MHz, with 128 chirps configured per frame and the frame duration set to 40 ms to achieve high range resolution. Such parameter configuration can make the spectrum processed by the radar have the advantage of high pixels and is more advantageous for collecting continuous human movements.

[0071] Specifically, the millimeter-wave radar signal preprocessing process refers to Figure 4 , referring to Figure 5 Perform Range FFT on the radar data of each frame along the fast-time direction to obtain range information:

[0072]

[0073] where X(c, s) represents the frame data of the c-th Chirp and the s-th Samples, N S is the size of Samples, and w(l) represents a Hamming window of length l. X r (c, r) represents the range-processed data of the c-th Chirp and the r-th range bin, r ∈ [0, N S ).

[0074] After completing the Range FFT, use a fourth-order high-pass Butterworth filter with a cut-off frequency of 0.007 Hz to eliminate static clutter. Then, perform Doppler FFT along the Figure 3 slow-time direction in f ×N rx ×N d ×N r . where N d is the doppler bins, and N r is the range bins.

[0075]

[0076] where X rd (v, r) represents the velocity-range data, N cis the size of the Chirp, W(k) represents the Hamming window, c∈[0,N c ).

[0077] Accumulate the absolute value of RD Cube in the distance dimension to obtain Doppler time data DT0Cube per second, with a size of N t ×N rx ×N v ×N f Where N t Number of seconds of data, N v is the speed, N f is the number of frames per second.

[0078] In this embodiment, a one-transmit-four-receive mode is used, but because the receiving antenna distance of AWR1642 is very small and the target range is narrow, it is impossible to obtain sufficient angle information, so the four RX channels are averaged and regarded as one channel. In addition, when humans perform continuous activities indoors, they usually complete them at a lower speed, so the data is appropriately cropped to filter out the high-speed area and obtain the DT1 Cube.

[0079]

[0080] DT0(i) represents the DT cube of the i-th receiving channel. By concatenating them, we can obtain DT data as the input of the subsequent network.

[0081] Specifically, the millimeter wave radar collects a continuous human motion dataset and establishes a dataset of short-distance continuous human activities in real indoor scenes. The subjects perform free and continuous activities within the measurement range, performing 8 different types of activities to obtain samples. There are no restrictions on the duration and transition time of the above activities.

[0082] Specifically, the dataset of short-distance continuous human activities in real indoor scenes includes: recruiting volunteers for data collection, including males and females, aged between 21 and 25 years old; the scene is arranged in a real living room scene, with an activity range of approximately 3.5 meters × 7 meters, and the radar system is set at a height of 1 meter; the subjects are required to perform 8 different daily activities in the activity area in front of the radar sensor; and use a camera to synchronously record video during the collection process.

[0083] See also Figure 6 ,In order to make the data set more like real-life activities, the 8 different types of activities are divided into two categories, Category A: X5, X6, X8, Category B: X1, X2, X3, X4, X7, where Category A is the main activity and Category B is the secondary activity. The subjects will perform the designed activity combination within the specified activity range. Finally, a combination sequence of 3 continuous human activities is constructed.

[0084] The subject completes the activity according to the above combination sequence within 45 - 60 s. Activity B can be interspersed randomly before and after performing Activity A. There is no restriction on the start and end times of the activity during the whole process. The transition between two consecutive activities can be freely determined, and the same action will be repeated.

[0085] Reference Figure 7 , and each subject performs each activity combination 4 times. Reference Figure 8 , during the acquisition process, the radar generates 25 frames of data per second. Each group of samples contains 1125 - 1500 frames, and there are a total of 78750 frames of data. The radar data is annotated frame by frame according to the recorded video.

[0086] Furthermore, a dual - dilation one - dimensional temporal convolutional network framework with an attention mechanism is proposed. By parallel - processing the sequential data of continuous HAR, spatial features and temporal long - short - term dependency information are extracted, and the sequence prediction result is output.

[0087] Specifically, a dual - dilation one - dimensional temporal convolutional network framework is proposed. The TCN uses the method of dilated convolutional layers for time - series modeling and shows enhanced performance in the field of video action segmentation. The TCN model does not depend on the input, can perform parallel computing for sequential output, greatly reduces the training time, and improves the computing efficiency. At the same time, by stacking dilated convolutions, the receptive field can be quickly expanded without increasing the number of parameters, and it can process longer - time data. In this embodiment, a TCN architecture with a self - attention mechanism is proposed to process the DT data generated by the radar.

[0088] Reference Figure 9 , whose sequence - to - sequence neural network architecture consists of DDM, AM, and PRM. From left to right, the data is input into the dual - dilation module, the self - attention mechanism module, and three single - dilation modules, and finally the sequence prediction result is output through Softmax. L is the total number of layers in the module, l represents the current layer number, and Dilation represents the dilation factor.

[0089] When stacking dilated convolutions with the dilation factor increasing as the number of layers increases, it will result in a very large receptive field in the upper layer, and the final prediction information will be affected by very distant time steps, while the receptive field in the lower layer is still very small. To solve this problem, the method of a dual - dilation model is adopted, and two dilated convolutions with different dilation factors are combined to obtain a more abundant receptive field. The definition of the dilated convolution is as follows:

[0090]

[0091] d is the dilation factor, k is the size of the convolution kernel, x s-d·i represents the convolution of the past state.

[0092] Reference Figure 10 The double dilation module first adjusts the number of input feature channels through a 1×1 depth convolution. Each layer consists of two 1D dilation convolutions with dilation factors of 2 l and 2 L-l in parallel, followed by a 1×1 convolution, a RELU activation function, and a Dropout layer with a rate of 0.5. In the parallel structure, the dilation factor of the former increases exponentially with the number of layers, and the dilation factor of the latter decreases exponentially with the number of layers. The convolution kernel size is 3×3, and the dilation is only applied in the time dimension, while the Doppler dimension remains unchanged. At the same time, the network also uses residual connections to facilitate gradient flow throughout the deep network. Repeat L layers in sequence, and finally output to the next module through a 1×1 convolution. This method exponentially increases the receptive field size, captures long-term dependencies in sequence data, and reduces overfitting effects on training data without increasing model parameters. Reference Figure 11 and Figure 12 respectively show the receptive field change patterns of dilation convolutions with dilation factors of 2 l and 2 L-l .

[0093] Specifically, the addition of the attention mechanism, reference Figure 13 , includes: the Self-Attention mechanism and the Multi-Head Self-Attention (MHSA) mechanism, which are used to enhance the global information capture ability of frame-level features and improve the attention to key frames.

[0094] Parallel stacking of 1D dilation convolution layers can capture frame-level features with a larger temporal receptive field, but fine-grained information will be lost in the deep layer, resulting in prediction errors for similar frames. Attention can establish dependencies between elements when processing sequence data and has a strong global information capture ability.

[0095] The Self-Attention mechanism passes the input sequence through three transformation matrices W Q , W K and W V with learnable weights to obtain the corresponding Query (Q), Key (K), and Value (V). The calculation process of the original Self-Attention mechanism is as follows:[[]]

[0096]

[0097] where d DTf represents the dimension of the data features. Reference Figure 14, the multi-head self-attention mechanism is introduced, which uses multiple attentions to capture global features, considers the context information of the sequence at the same time, and strengthens the attention to the key frame information.

[0098] First, calculate head1, head2, …, head through Equation 6a i After that, through Equation 6b, the multi-head self-attention mechanism concatenates them and passes them into a linear layer to obtain the output MultiHead.

[0099] head i = Attention(QW i Q , KW i K , VW i V ) #(6a)

[0100] MultiHead = Concat(head1,..., head i )W O #(6b)

[0101] Where and are learnable training matrices in the i-th self-attention mechanism channel, and W O represents the linear transformation matrix.

[0102] After the frame-by-frame feature information passes through the DD Module and the Attention Module, the initial prediction result is generated as the input of the PRModule. Refer to Figure 15 , in the PRM, a one-dimensional dilated convolution with a dilation factor of 2 l is used, followed by a 1×1 convolution, RELU activation, and Dropout. In addition, the PRM combines 3 modules in sequence, and the input of each module is the output of the previous module, which can provide more context information and thus gradually refine the prediction result. Finally, the prediction result of each frame is output through Softmax. Refer to Figure 16 , summarizes the structural information of the RA-TCN network, and Output is the output size.

[0103] Part Two: The evaluation index STT Score for radar continuous human action recognition is established.

[0104] Specifically, the evaluation index STT Score for radar continuous human action recognition allows a certain tolerance for the error of activity recognition within a very short time interval, so as to more accurately reflect the continuity and transitional ambiguity of human actions.

[0105] In the field of continuous human activity recognition, traditional accuracy metrics provide a simple assessment of recognition accuracy while ignoring the problem of over-segmentation. Additionally, it is biased towards evaluating longer actions, thus weakening the importance of shorter actions. To address these drawbacks, this study introduces the segment-based F1 score as a supplementary evaluation metric to better assess model performance.

[0106] To evaluate a prediction model, the frame-by-frame accuracy (Acc) is generally used:

[0107]

[0108] Here, Correct represents the number of frames with correct predictions, and Total is the number of data frames.

[0109] However, from the perspective of continuous human activity recognition, Acc can only obtain a general recognition accuracy and cannot reflect the problem of over-segmentation; moreover, long-duration actions within the sample have a greater impact on Acc than short-duration actions. Therefore, this invention proposes the segment-based F1 score for further evaluation.

[0110] The segment-based F1 score is generally used in action segmentation and detection tasks. This method evaluates performance by calculating the IOU between each predicted action and the ground truth action.

[0111]

[0112] Among them, P represents the predicted action time interval, and T represents the ground truth action time interval.

[0113] As shown in the following formula, P represents the prediction result, and GT represents the ground truth label. The segmentation method compares the IOU with the threshold K to determine true positives (TP), false positives (FP), and false negatives (FN).

[0114]

[0115] Reference Figure 17 As shown, if the segment prediction is correct, when the IOU is greater than or equal to the threshold K, it is marked as TP, and if it is less than K, it is FP. If the segment prediction is incorrect, it is marked as FN. Additionally, when there are multiple correct predictions within the range of a single ground truth action, only one will be classified as TP, and the others will be FP.

[0116] Then, the precision and recall of the sum of TP, FP, and FN for all classes are calculated, and the segment-based F1 score is calculated.

[0117]

[0118] The segmented F1 evaluation can effectively penalize over-segmentation errors, and the score depends on the number of activities in the data rather than the number of frames, which is beneficial to providing a more detailed and accurate evaluation for continuous human activity recognition.

[0119] Since it is collected in a real scenario, the activities of the subjects occur one after another when performing some activities. When a person is walking (X1) and suddenly falls (X6), and when a person drinks water (X5) immediately after sitting on the sofa (X3) and then stands up from the sofa (X4), these activities occur continuously. This results in a very blurred transition between actions. Especially within a 0.2s time period, it is difficult to well represent an action, and it is almost okay to define it as the previous action or the next action.

[0120] Therefore, based on F1, we proposed the STT Score for radar continuous human action recognition. See Equation 11, len p represents the frame length of the prediction time. When counting TP, FP, and FN, if the comparison between the prediction result of the action transition frame and the ground truth is within 5 frames, it is not included in the calculation of FN. This evaluation method can tolerate the error of activity transition recognition within a very short time (0.2s), and can more accurately reflect the continuity of human actions while evaluating the action recognition performance.

[0121]

[0122] To verify the performance of the proposed model, four deep learning models, including LSTM, Bi-LSTM, GRU, and Bi-GRU, were selected for comparative experiments. The input and output forms of all models were kept consistent and were configured as frame sequences.

[0123] Reference Figure 18 , both (Bi-)LSTM and (Bi-)GRU consist of three parts: a 1×1 one-dimensional convolutional layer, three (Bi-)LSTM or (Bi-)GRU layers, and a connection layer. The convolutional layer adjusts the number of feature channels of the input data DT data to be the same, all 64. Then each (Bi-)LSTM or (Bi-)GRU layer contains 64 hidden units and is followed by a Dropout of 0.5 to prevent overfitting. Finally, the activity prediction sequence is output through the connection layer and the Softmax layer. Adam is used as the optimizer with a learning rate of 0.0005, and the loss function also uses a combination of classification loss and smoothing loss, where the cross-entropy loss is used as the classification loss and T-MSE is used as the smoothing loss. The training environment and platform are kept consistent. Reference Figure 19 , which specifically records the sizes and properties of each layer of the (Bi-)LSTM and (Bi-)GRU networks above.

[0124] In addition to comparing with (Bi-)LSTM and (Bi-)GRU, ablation experiments were also conducted on the self-attention mechanism to evaluate the performance and differences of six different network models on unknown activity sequences. We divided S1, S2, and S3 in the dataset into three layers and used a three-layer cross-validation experiment, with no information overlap between the training set and the test set.

[0125] Reference Figure 20 , the evaluation results of six different continuous human activity recognition models are given, including frame-by-frame accuracy and segment score evaluation, and the threshold K for F1 and STT scores is 0.5. Here, split = 1 means using activity sequences S2 and S3 as the training set, and S1 as the test set. In continuous human activity detection, the bidirectional LSTM and GRU networks outperform the unidirectional networks, with F1 scores of 0.6422 and 0.6461 respectively, and STT scores of 0.7814 and 0.7902 respectively. The proposed network model R-ATCN obtained the highest Acc score of 0.9321 in the experiment, with an F1 score of 0.6872 and an STT score of 0.8325. Without the self-attention mechanism module, the performance of this network decreased slightly, only slightly lower than R-ATCN.

[0126] In terms of inference speed, the R-ATCN network without the self-attention mechanism has the fastest inference speed of 75.47 ms, with 0.693M parameters, so it takes 82.89 ms for prediction generation. It is worth noting that the training time of R-ATCN is only 7.01 seconds, significantly shorter than the comparison networks. This efficiency can be attributed to the recurrent neural network structure for sequential prediction, where the activation of each intermediate step depends on the previous time step. In contrast, R-ATCN calculates the activation of each time step simultaneously, thus accelerating the training and testing processes.

[0127] Reference Figure 21 , shows the predictions made by various models for the selected activity samples, and GT represents the ground truth. The LSTM and GRU networks show more incorrect predictions. For example, LSTM misclassifies some actions of X1 as X3. On the other hand, the Bi-LSTM and Bi-GRU networks still have the problem of over-segmentation, identifying multiple segments of actions in the X8 activity. In contrast, the R-ATCN network model proposed in the present invention can accurately identify all activities in the sample regardless of their duration. It is crucial to include the self-attention mechanism module in the network because it can capture the key frame information of the transition of continuous activities, thus enabling more accurate judgment in action transition segmentation.

[0128] The above describes a preferred complete embodiment of the present invention, and does not limit the protection scope of the present invention. Without departing from the core design principle of the present invention, any improvement or equivalent substitution made to the technical solution of the present invention in terms of principle, structure or method shall be regarded as falling within the protection scope covered by the claims of the present invention.

Claims

1. A continuous human action recognition method based on millimeter-wave radar, characterized in that, It includes: Millimeter-wave radar signal processing; The millimeter-wave radar collects a continuous human action data set; a dual-dilated one-dimensional time convolutional network framework is proposed, an attention mechanism is added, and by parallel processing the sequence data of continuous HAR, spatial features and long-term and short-term temporal dependence information are extracted, and sequence prediction results are output; The millimeter-wave radar signal processing includes: millimeter-wave radar signal acquisition parameter configuration; millimeter-wave radar signal preprocessing; The proposed dual-dilated one-dimensional time convolutional network framework includes three core modules: DualDilatedModule, AttentionModule, and PredictionRefinementModule to process the time series data generated by the radar; First, DualDilatedModule expands the receptive field through one-dimensional dilated convolutions with two different dilation factors in parallel, uses 1×1 convolutions for channel adjustment, and uses a combination of exponentially increasing and decaying dilation factors to evenly expand features at different levels. At the same time, ReLU activation, Dropout, and residual connections are combined to stabilize training and prevent gradient vanishing; Next, AttentionModule introduces a multi-head self-attention mechanism, calculates frame-level relationships through Query (Q), Key (K), and Value (V), and multiple attention heads capture global temporal information in parallel, enhancing the attention to key frames and improving the accuracy of frame-level classification; Finally, PredictionRefinementModule further optimizes the preliminary prediction results, gradually refines features using dilated convolutions, 1×1 convolutions, ReLU activation, and Dropout, and outputs the final frame-level classification results through Softmax; The added attention mechanism includes: self-attention mechanism and multi-head self-attention mechanism, which are used to enhance the global information capture ability of frame-level features and improve the attention to key frames. First, in the self-attention mechanism, the input sequence generates Query (Q), Key (K), and Value (V) through three learnable weight transformation matrices, and obtains the global dependence information at each time step through attention weight calculation, so as to improve the modeling ability of long-term dependence relationships. Second, in order to further improve the modeling effect, a multi-head self-attention mechanism is introduced. The features of different subspaces are calculated in parallel through multiple attention heads (head1, head2,..., head i ), and each attention head independently learns different context information. Finally, the final multi-head output is obtained through concatenation and linear transformation; The extraction of spatial features and long-term and short-term temporal dependence information from the parallel processing of continuous HAR sequence data includes: spatial feature information such as target distance and speed collected by the FMCW radar, and long-term and short-term temporal dependence information; among them, spatial features reflect the relative position and dynamic changes between the target and the radar, while long-term and short-term temporal dependence information helps the model understand the temporal structure and trend of activities, including short-term action transitions and long-term overall activity patterns, and these information jointly support the accurate recognition of continuous human activities; The output of sequence prediction results includes: the prediction of human activity types corresponding to each frame of data, which implies the duration information of the activity and particularly focuses on the capture of key frames of activity transitions; these prediction results are given in the form of a sequence, maintaining the continuity and temporal order of the activity, and achieving high-accuracy prediction through the proposed RA-TCN network; Establish an evaluation index STTScore for radar continuous human action recognition; The evaluation index STTScore for radar continuous human action recognition allows a certain tolerance for the error of activity recognition within an extremely short time interval, so as to more accurately reflect the continuity and transitional ambiguity of human actions; The segment-based F1 score is used for action segmentation and detection tasks; the performance is evaluated by calculating the Intersection over Union (IOU) between each predicted action and the ground truth action, and the calculation of IOU can be expressed as: where P represents the predicted action time interval, and T represents the ground truth action time interval; The segmentation method compares the IOU with a threshold K to determine true positives (TP), false positives (FP), and false negatives (FN); P represents the prediction result, GT represents the ground truth label, and the calculation of F1 is expressed as: If the segment prediction is correct, when the IOU is greater than or equal to the threshold K, it is labeled as TP, and if it is less than K, it is FP; if the segment prediction is incorrect, it is labeled as FN; additionally, when there are multiple correct predictions within the range of a single ground truth action, only one will be classified as TP, and the others will be FP; Then, the precision and recall of the total TP, FP, and FN of all classes are calculated, and the segment-based F1 score is calculated: Based on F1, STTScore is proposed for continuous human action recognition by radar; len p denotes the frame length of the prediction time. When counting TP, FP, and FN, if the comparison between the prediction result of the action transition frame and the ground truth is within 5 frames, it is not included in the calculation of FN. STTScore can be expressed as:

2. The continuous human action recognition method based on millimeter-wave radar according to claim 1, characterized in that The millimeter-wave radar collects a continuous human action dataset; it includes: radar data and synchronized video data of the collector performing 8 daily activities in a real living room environment, and the collector completing a preset activity combination within a specified area; the activities are divided into two categories: Category A: major activities, drinking water, falling, picking up an object, and Category B: minor activities, walking, staying in place, sitting on the sofa, standing up from the sofa, standing up from the ground, and 3 sequences of continuous human activity combinations are designed, and each subject completes the activities in 45 - 60 seconds. Category B activities can be freely interspersed before and after performing Category A activities, and there are no restrictions on the start, end, and transition methods of the actions; the same combination is repeated multiple times.