Activity identification system and method based on Doppler characteristics of millimeter wave radar

Through the activity recognition system of Doppler features of millimeter wave radar, dynamic feature screening and cross-channel information fusion technology are used to solve the problems of low recognition accuracy and high computational complexity in scenarios of insufficient data and imbalance, and efficient real-time human behavior recognition is achieved.

CN120370282AActive Publication Date: 2025-07-25SHANDONG UNIV +1

Patent Information

Application Number
CN202510869104.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

The existing millimeter-wave radar human activity recognition technology has low recognition accuracy and high computational complexity in scenarios such as insufficient data and unbalanced data, which is difficult to meet real-time requirements, and the model generalization ability is insufficient.

Method used

Using an activity recognition system based on the Doppler feature of millimeter wave radar, through dynamic feature screening and cross-channel information fusion, combining fast Fourier transform, short-time Fourier transform and four-stage pyramid architecture, masked splicing modules and gated convolution linear units are designed to achieve efficient feature extraction and classification.

Benefits of technology

In the light change, noise interference and action similarity scenarios, the recognition accuracy is significantly improved, the calculation complexity is reduced, and the real-time reasoning at milliseconds is realized. It is suitable for smart home, medical monitoring and industrial security scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120370282A_ABST
    Figure CN120370282A_ABST
Patent Text Reader

Abstract

The invention discloses an activity identification system and method based on millimeter wave radar Doppler characteristics, and relates to the technical field of radar signal processing and artificial intelligence. Comprising a human body behavior radar acquisition module, a human body behavior information transmission module, a human body behavior information preprocessing module, a micro-Doppler feature extraction module, a data classification and identification module and a human body behavior information application module which are connected in sequence, a millimeter-wave radar is adopted to emit high-frequency millimeter-wave signals, and information of human body actions is obtained by receiving signals reflected from the surface of a human body; and the human body behavior information transmission module is used for transmitting the collected behavior data of the user to edge equipment or a cloud server.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of radar signal processing and artificial intelligence technology, and in particular to an activity recognition system and method based on millimeter wave radar Doppler characteristics. Background Art

[0002] Human activity recognition has important application potential in the fields of smart medical care, smart home and security monitoring. Existing human behavior recognition technologies are mainly divided into three categories: vision-based, wearable device-based and radio frequency-based methods. Vision-based methods use cameras to capture images or videos for recognition, but they are sensitive to lighting conditions and pose privacy risks. Although wearable devices can achieve high accuracy, they require users to wear the devices continuously, interfering with daily activities. In contrast, millimeter-wave radar has become an ideal solution in the field of human behavior recognition due to its advantages such as strong environmental adaptability, no influence of light, and non-contact perception. Millimeter-wave radar extracts point clouds, range-Doppler maps and micro-Doppler features by analyzing radio frequency echo signals. Among them, micro-Doppler features are widely used in classification tasks because they can accurately capture the subtle vibration patterns of the human body.

[0003] Traditional radar human activity recognition methods rely on manual feature extraction, such as multi-layer perceptron, principal component analysis, and support vector machine. Although these methods are effective in specific scenarios, they require domain expertise and are easily affected by environmental interference, resulting in limited classification performance. Deep learning technology has significantly improved recognition accuracy, but it relies on a large amount of data, and the high cost of radar data acquisition limits its practical application. Although the Transformer network can analyze spatiotemporal micro-Doppler features through the self-attention mechanism, it has high computational complexity and is difficult to meet real-time requirements; although the hybrid network integrates spatiotemporal features, it still faces the challenges of large data requirements, high resource consumption, and insufficient real-time performance.

[0004] Although recent studies have improved the generalization ability of micro-Doppler features through lightweight modules and transfer learning, these methods are mostly designed based on scenarios with sufficient data. In actual applications, public radar datasets are scarce and data distribution is uneven, resulting in significant performance degradation when the model is insufficient or unbalanced. Existing methods still have obvious deficiencies in balancing feature extraction efficiency, noise robustness, and computational overhead. There is an urgent need for a human behavior recognition solution that takes into account high accuracy, low resource consumption, and strong generalization ability. Summary of the invention

[0005] The technical problem to be solved by the present invention is to provide an activity recognition system and method based on millimeter-wave radar Doppler characteristics. Through dynamic feature screening and cross-channel information fusion technology, the recognition accuracy and robustness in data insufficient and unbalanced scenarios are significantly improved, and efficient real-time reasoning is achieved while reducing computational complexity. It is suitable for scenarios such as smart homes, medical monitoring and industrial safety.

[0006] The present invention proposes an activity recognition system and method based on the Doppler characteristics of millimeter-wave radar. Through the full-process technical design, from data acquisition, signal processing to efficient feature extraction and classification, the recognition ability in complex environments is comprehensively improved. This method uses a millimeter-wave radar system to collect human activity echo signals, covering a variety of daily activities and complex behaviors, ensuring data diversity and adaptability to actual scenarios. In the preprocessing stage, the range-time graph is extracted through fast Fourier transform, the moving target indication technology is combined to eliminate static background interference, and the short-time Fourier transform is used to convert the time-domain signal into a high-resolution micro-Doppler feature map. An input image containing frequency-domain energy distribution information is generated through color mapping, effectively capturing the subtle vibration patterns of limbs and providing a data basis for subsequent analysis.

[0007] In terms of network architecture design, the masked splicing module locally masks and weights the input features through a dynamically generated random masking matrix, suppressing noise and strengthening the expression of key spatio-temporal regions. This module combines depthwise separable convolution and hierarchical pooling operations to gradually compress the feature dimension, reduce redundant calculations, and at the same time retain high-frequency detail information. The gated convolutional linear unit module extracts local spatial features and global channel features through a two-branch structure, and uses a gating mechanism to adaptively fuse multi-scale information, enhancing the sensitivity to the energy distribution and contour changes in the micro-Doppler features. The overall network adopts a four-stage pyramid architecture to fuse multi-scale spatio-temporal features layer by layer, and finally outputs the classification result through global average pooling and a fully connected layer, achieving millisecond-level real-time inference while ensuring high accuracy.

[0008] Tests on multiple public datasets and self-built datasets show that this method exhibits significant advantages in scenarios of light changes, noise interference, and action similarity. Through dynamic masking and cross-channel interaction, the network can accurately distinguish subtle action differences, and the clarity of the classification boundary is significantly improved compared with traditional methods. For data-scarce scenarios, the adaptive masking mechanism effectively alleviates the problem of model overfitting. Combined with lightweight design, stable performance is still maintained under low computational resource consumption. In practical applications, this method can be integrated into smart home systems to monitor user activities in real time and trigger security responses; in the medical field, it supports non-contact rehabilitation action tracking and quantitative evaluation; in industrial scenarios, it accurately identifies workers' dangerous behaviors, including non-standard operations and high-risk actions, and issues timely warnings to reduce accident risks.

[0009] By integrating dynamic feature optimization, lightweight architecture, and multi-scale information fusion, the present invention provides a highly robust and low-cost solution for millimeter-wave radar human activity recognition systems, promoting the wide application of non-contact sensing technologies in the fields of smart terminals, industrial safety, and medical health. In the future, it can be further combined with multi-modal data fusion and edge computing optimization to expand its application potential in fields such as unattended monitoring and intelligent transportation.

[0010] The present invention realizes the invention purpose by adopting the following technical solutions: An activity recognition system and method based on the Doppler characteristics of millimeter-wave radar, characterized by comprising a human behavior radar acquisition module, a human behavior information transmission module, a human behavior information preprocessing module, a micro-Doppler feature extraction module, a data classification and recognition module, and a human behavior information application module connected in sequence: The human behavior information acquisition module is used to collect the reflected signals in the target area, emit high-frequency millimeter-wave signals by using a millimeter-wave radar, and obtain the information of human actions by receiving the signals reflected from the human body surface; The human behavior information transmission module is used to transmit the collected behavior data of the user to an edge device or a cloud server; The human behavior information preprocessing module is used to receive the radar signal data transmitted from the human behavior information transmission module, perform formatting processing, and then use Fourier transform, moving target signature extraction, and short-time Fourier transform to analyze the frequency distribution of the signals, and extract the features related to human behavior from them; The micro-Doppler feature extraction module is used to input the behavior data of the user preprocessed by the human behavior information preprocessing module into the trained micro-Doppler feature extraction module for advanced feature extraction; The data classification and recognition module is used to train and accurately discriminate the behavior data input into the recognition model.

[0011] The human behavior information application module is used to transmit the obtained human behavior classification and recognition results to the user information interface, so as to realize the display of human information.

[0012] As a further limitation of this technical solution, the micro-Doppler feature extraction module adopts a four-layer pyramid stacking structure, which is a data downsampling unit, a four-layer feature extraction unit, and a fully connected layer classifier connected in sequence. Each layer of the feature extraction unit includes a mask splicing unit, a two-dimensional maximum pooling unit, and a convolutional linear gated unit connected in sequence.

[0013] As a further limitation of this technical solution, the data classification and recognition module includes a loss function calculation and update module and a classification and recognition output module connected in sequence; The loss function calculation and update module is used to calculate the mean and variance of the behavior data of known categories during the training process, and generate a loss function to update the parameters of the recognition model; The classification and recognition output module is used to calculate the determination threshold of known categories and category mapping.

[0014] As a further limitation of the present technical solution, the human behavior information acquisition module includes a millimeter-wave radar transmitting unit, a millimeter-wave radar receiving unit, and a signal amplification and processing unit, which are integrated in a radar acquisition board. The acquisition board is installed at a designated position, and the original data of the intermediate-frequency signal after processing is transmitted by transmitting a frequency-modulated continuous wave to detect the human echo within the detection area.

[0015] As a further limitation of the present technical solution, the human behavior information preprocessing module includes a signal storage unit, a window division unit, a Chirp matrix data reshaping unit, a range Fourier transform unit, a moving target feature extraction unit, and a short-time Fourier transform unit connected in sequence, and is used to obtain a micro-Doppler map from the original data of the millimeter-wave radar.

[0016] As a further limitation of the present technical solution, the human behavior information transmission module is used to transmit the collected behavior data of the user to an edge device or a cloud server through any one of the transmission methods of WiFi, serial communication, or Lora.

[0017] An identification method for an activity recognition system based on the Doppler characteristics of a millimeter-wave radar, characterized by comprising the following steps: S1: Human behavior data acquisition; S2: Human behavior data transmission; S3: Human behavior data preprocessing; S4: Micro-Doppler feature extraction and identification; The micro-Doppler maps preprocessed in S3 are input into the identification model in batches, and a fully connected layer classifier is obtained through training; The fully connected layer classifier includes a GlobalAvgPool2D two-dimensional global average pooling layer and a Linear linear layer connected in sequence; S5: Loss function calculation and network update; In S4, the micro-Doppler feature extraction module generates a feature vector through forward propagation and inputs it into the data classification and identification module for processing. In order to achieve efficient training and optimization, the cross-entropy loss function is used as the main optimization target, and at the same time, the Adam optimizer is used to update the model parameters. Specifically, the cross-entropy loss function is used to measure the difference between the model prediction probability distribution and the true label, and its mathematical expression form is shown in Equation (1): (1); Where: represents the cross-entropy loss value, is the total number of categories, is the one-hot encoded vector of the true label, is the probability distribution predicted by the model; In each round of training, data is input into the model's micro-Doppler feature extraction module in batches. The amount of data in each batch is set to B, where B is any positive integer less than the total amount of data. The Adam optimizer dynamically adjusts the learning rate based on the gradient information of the current batch, thereby efficiently updating the model parameters. The core idea of the Adam optimizer is to combine the advantages of the momentum method and the adaptive learning rate. Its parameter update rules are shown in Equations (2), (3), (4), and (5): (2); (3); (4); (5); Where: Represents the iteration step number, Represents the gradient of the current batch, And Are the first-order moment estimate and second-order moment estimate of the gradient respectively, And Are the bias correction terms for the original first-order moment and second-order moment estimates, And Are the exponential decay rates, Is the initial learning rate, Is a small constant to prevent division by zero, Are the model parameters; S6: Determine the threshold and classification discrimination output; After the training of the micro-Doppler feature extraction module is completed, the central mean of the feature vectors of known classes is statistically calculated, and a confidence threshold is set based on this to judge the reliability of the prediction result during output. First, for each class The set of feature vectors is processed to calculate its class central mean vector , as shown in Equation (6): (6); Where: Represents the central mean vector of class , Is the number of samples in class , Represents the th sample in class In the feature space; Further calculate the distance from all samples in each class to its class central mean, and statistically calculate the mean and standard deviation of these distances, denoted as And , as shown in Equations (7) and (8): (7); (8); Wherein: represents the eigenvector to the class center mean of the Euclidean distance; In order to set the confidence threshold, the distance mean plus a certain multiple of the standard deviation is used to define the confidence range for each class, as shown in Equation (9): (9); Wherein: represents the class of the confidence threshold, is a hyperparameter; In the model inference stage, for the input sample , the model first predicts its belonging class , and then calculates the distance from the sample to the predicted class center mean. If this distance is less than or equal to the corresponding confidence threshold , the prediction result is considered reliable; otherwise, the prediction result is considered unreliable and belongs to an unknown class or needs further confirmation; Step S7: Discriminant output result application; The results output by the closed set and the open set are transmitted in real time to the corresponding application platforms.

[0018] As a further limitation of this technical solution, the specific implementation process of S3 is as follows: S31: Human behavior information data reshaping; When processing the intermediate frequency raw millimeter wave radar data, in order to facilitate subsequent feature extraction and model training, it is first necessary to rearrange the continuous radar signals into a two-dimensional matrix form according to Chirp. The raw millimeter wave radar data is expressed as , where represents the intermediate frequency signal value of the th sampling point, is the total number of sampling points. At the same time, each Chirp contains sampling points, and there are a total of Chirps. The total number of sampling points of the original data satisfies ; The original data S is first divided into multiple Chirps, each Chirp corresponding to a sampling point sequence with a length of . These Chirps are arranged in order to form a two-dimensional matrix , whose dimension is , the two-dimensional matrix is represented as: (10); Where: represents the -th sampling point value of the -th Chirp; S32: Radar data window segmentation; Based on the two-dimensional matrix generated in S31, the data is further segmented according to the time window. The window length defines the number of Chirps included in each window, and the step size determines the overlapping degree or interval between adjacent windows. The starting Chirp and ending Chirp of the -th window can be calculated by Equations (11) and (12): (11); (12); Where: represents the segment number, and the value range is , represents the starting position of the -th segment, represents the ending position of the -th segment, and satisfies ; represents the fixed length of each window, represents the actual length of the -th segment; The total number of the finally generated window segments can be determined by Equation (13): S33: Micro-Doppler processing; In the window matrix obtained in step S32, each row represents the sampling point data of a Chirp. By performing a fast Fourier transform on each row, the time-domain signal is converted into a frequency-domain signal, thereby extracting the distance information of the target, as shown in Equation (14): Where: is the imaginary unit, is the value of the -th row and -th column in the input window matrix, represents the -th complex value at the -th distance cell, represents the frequency at which the signal is sampled, is the number of samples per Chirp and also the number of points for the FFT; After performing the fast Fourier transform, the range dimension data contains static clutter. To highlight the dynamic target signal and suppress the static clutter, a high-pass Butterworth filter is used to perform moving target indication filtering on the range dimension data. For discrete-time signals, the difference equation form of the Butterworth filter is shown in Equation (15): (15); Where: is the input signal, is the output signal, and are the forward and feedback coefficients of the filter respectively; and represent the forward order and the feedback order of the filter; After performing moving target indication filtering, the range dimension data is shown in Equation (16): (16); Where: represents the range dimension data after Butterworth filtering, represents the Butterworth filtering operation, The original range dimension data, at the th range cell, the value at the th time sample, represents the forward coefficient vector, represents the feedback coefficient vector; Perform short-time Fourier transform on the data after moving target indication filtering. The core idea of the short-time Fourier transform is to divide the signal into multiple short-time windows and perform FFT within each window, thereby obtaining the time-frequency distribution of the signal as shown in Equation (17): (17); Where: is the imaginary unit, represents the input signal, represents the window function, represents the time-frequency distribution, represents the center of the local window for the current analysis, represents the frequency component for the current analysis; After being processed by the short-time Fourier transform, the micro-Doppler spectrogram has a dimension of , where represents the frequency axis resolution, represents the time axis resolution; Finally, the micro-Doppler spectrogram is represented by Equation (18): (18); where: represents the energy distribution at frequency and time ; S34: Visualization of micro-Doppler features; Convert the micro-Doppler spectrogram described in S33 into a logarithmic form, visualize it through an image drawing tool, and add label information by naming.

[0019] As a further limitation of this technical solution, the specific implementation process of S4 is as follows: S41: Data downsampling; First, process the micro-Doppler feature map input in S34 through a data downsampling module to reduce its spatial resolution and expand the number of channels. This process is represented by Equation (19): (19); where: represents the output of the data downsampling module, is the Gaussian error linear unit, represents the input feature map, represents the block size for downsampling, and the inputs of the function are the feature map, convolution kernel size, and stride respectively; S42: Micro-Doppler feature extraction; The data processed through S41 is passed to four identical stages for layer-by-layer feature extraction. Each stage is composed of a mask splicing module, two-dimensional max pooling, and a convolutional linear gated unit connected in sequence. This process is represented by Equation (20): (20); where: represents the output of this stage, represents the operation of the convolutional linear gated unit, represents the operation of the mask splicing unit, represents the two-dimensional max pooling operation; After being processed layer by layer through four stages, the input feature map is converted into a high-level feature representation for the final classification or regression task, as shown in Equation (21): (21); where: represents the extracted high-level features, is the initial feature map processed by the data downsampling module; represents the operation processed through four stages; S43: Classification and recognition.

[0020] As a further limitation of this technical solution, the specific implementation process of S42 is as follows: S421: Mask splicing unit; The mask splicing unit captures spatio-temporal features through multi-stage feature fusion and enhances the perception ability of key regions. Let the input feature map be , first use A depth convolution layer with a convolution kernel and a stride of 1 extracts the basic features, and after batch normalization, the intermediate feature is obtained. The adaptive mask mechanism masks the implementation area through the dynamically generated binary mask matrix , and its process is shown in Equation (22): (22); Where: is a learnable scaling factor, and the mask matrix is generated under the control of the initial mask rate ; The masked feature is compressed by a multi-layer perceptron and then added to the original input to generate an enhanced feature as shown in Equation (23): (23); Where: represents a multi-layer perceptron, which is a feedforward neural network structure used for non-linear transformation of input data; The final module output is achieved by concatenating the original input and the enhanced feature along the channel dimension, as shown in Equation (24): (24); The principle of the adaptive mask generation mechanism therein is shown in Equations (25) and (26): (25); (26); Where: represents the initial mask vector, whose length is equal to the initial resolution , and the random permutation function is used to randomly shuffle the elements in , and is used to readjust the one-dimensional vector into a matrix with the same shape as the target feature map ; S422: Convolutional linear gating unit; The convolutional linear gated unit enhances the feature expression ability through cross-channel interaction, and the input features are respectively passed through two independent linear layers and After projection, one branch extracts spatial domain features through a 3×3 depth convolution, and modulates them element-wise with the gating weights of the other branch to generate implicit abstract features , as shown in Equation (27): (27); Where: is the Gaussian error linear unit, represents a 3×3 depth convolution with a stride of 1, is element-wise multiplication. After this feature undergoes a second depth convolution and a linear transformation, a triple summation operation is performed with the original input, as shown in Equation (28): (28); Where: is the linear layer, is element-wise addition, represents the output of the convolutional linear gated unit.

[0021] The MC-GCLU network proposed in the present invention addresses the problem of insufficient feature extraction ability in scenarios with limited data scale or unbalanced distribution. By enhancing the model's adaptive learning ability for sparse data, it significantly improves the feature extraction efficiency and generalization performance. The network exhibits stronger robustness under low sample size or class imbalance conditions, effectively alleviating the restriction of data scarcity on model performance. To solve the balance problem between computational overhead and accuracy, the present invention proposes a joint design scheme of a pyramid architecture and a mask splicing module. This module reduces redundant calculations through a dynamic feature selection mechanism, optimizes the feature processing path, and reduces the computational cost while maintaining the recognition accuracy, meeting the computing power constraint requirements for real-time inference. The present invention proposes a method for human behavior recognition based on millimeter-wave radar. By setting an appropriate loss function, adjusting model parameters, and setting an open-set threshold, the problem that existing classification models cannot recognize data of unknown categories can be solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a schematic diagram of the module composition and connection relationship of the present invention.

[0023] Figure 2 is a schematic diagram of the flow of the human activity recognition method of the present invention.

[0024] Figure 3 is a schematic diagram of the principle of the micro-Doppler feature extraction module in the local human activity recognition method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0025] The following will describe in detail a specific embodiment of the present invention in conjunction with the accompanying drawings. It should be understood that the protection scope of the present invention is not limited by the specific embodiment.

[0026] Embodiment 1

[0027] In the scenario of monitoring the falls of the elderly, there are problems such as traditional visual monitoring infringing on privacy and insufficient night monitoring capabilities. The present invention is deployed in a nursing home monitoring system, and non-contact perception of human micro-motion characteristics is realized through a millimeter-wave radar, so as to achieve all-weather non-intrusive monitoring. The millimeter-wave radar device is installed on the top of the room, emits a frequency-modulated continuous-wave signal to the target area, generates an intermediate-frequency signal after receiving the human reflected echo, and generates a micro-Doppler spectrogram after preprocessing. When detecting high-risk behaviors such as falls, the system automatically triggers an alarm and notifies the nursing staff, and at the same time supports long-term behavior data analysis to evaluate the health status.

[0028] As Figure 1 shown, the present invention includes a human behavior radar acquisition module, a human behavior information transmission module, a human behavior information preprocessing module, a micro-Doppler feature extraction module, a data classification and recognition module, and a human behavior information application module that are connected in sequence: The human behavior information acquisition module is used to collect the reflected signals in the target area. It uses a millimeter-wave radar to emit high-frequency millimeter-wave signals, and obtains the information of human actions by receiving the signals reflected from the human body surface. This module first starts the millimeter-wave radar transmitting unit to emit electromagnetic waves to the target area, and receives the reflected echo signals through the millimeter-wave radar receiving unit. The radar signals are amplified by a low-noise amplifier at the front end, and then enter the signal amplification processing unit, where preliminary time-domain and frequency-domain filtering are performed to remove environmental noise. The data acquisition process continues. When the collected original radar data meets the processing requirements, the data is stored and cached for subsequent transmission and processing. The types of collected data are mainly information such as the frequency change, echo intensity, and time difference of arrival of the radar echo signals.

[0029] The human behavior information transmission module is used to transmit the collected behavior data of the user to the edge device or the cloud server. The appropriate transmission method can be selected according to the scenario where the user is located, and the collected behavior data of the user is transmitted to the cloud server through any one of the transmission methods of WiFi and Lora, and then forwarded to the specified device. Different transmission methods are different in terms of power consumption, transmission distance, transmission speed, etc., so they are suitable for different application scenarios. For the application scenario that is directly an edge device, it can also be quickly transmitted through a serial port for data processing.

[0030] The human behavior information preprocessing module is used to receive the radar signal data transmitted from the human behavior information transmission module and perform formatting processing, including time alignment and window division, to ensure the consistency of millimeter-wave radar data in the time dimension. Then, Fourier transform, moving target feature extraction, and short-time Fourier transform are used to analyze the frequency distribution of the processed data, and features related to human behavior are extracted therefrom.

[0031] The micro-Doppler feature extraction module is used to input the behavior data of the user preprocessed by the human behavior information preprocessing module into the trained micro-Doppler feature extraction module for advanced feature extraction.

[0032] The data classification and recognition module is used to train and accurately discriminate the behavior data input into the recognition model.

[0033] The human behavior information application module is used to transmit the obtained human behavior classification and recognition results to the user information interface, thereby realizing the display of human information.

[0034] This system realizes the precise monitoring of fall-high-risk behaviors and daily behaviors through the time-frequency feature analysis of millimeter-wave radar echo signals. In the micro-Doppler feature extraction module, a four-layer pyramid stacking structure is constructed, and a multi-resolution fusion strategy that combines an adaptive mask mechanism and a convolutional linear gated unit is adopted. By dynamically perceiving the micro-Doppler feature distribution of key motion regions, the recognition accuracy of complex human actions is significantly improved; at the same time, a dynamic threshold determination method based on statistical distribution modeling is proposed to effectively avoid the risk of misjudgment.

[0035] Embodiment 2: The micro-Doppler feature extraction module adopts a four-layer pyramid stacking structure, which is a data downsampling unit, four-layer feature extraction units, and a fully connected layer classifier connected in sequence. Each layer of feature extraction unit includes a mask splicing unit, a two-dimensional max pooling unit, and a convolutional linear gated unit connected in sequence.

[0036] The main function of the data downsampling unit is to sample the preprocessed user behavior data to achieve efficient data compression and feature extraction. This module performs downsampling operations on the input data through specific sampling sizes and activation functions to generate feature representations suitable for subsequent processing. In a specific implementation, the data sampling layer processes data with an original input dimension of 3×224×224. The data sampling layer uses a 4×4 convolutional kernel and combines it with the Gaussian error linear unit for downsampling, thus effectively compressing the data volume during the data processing. By performing convolutional operations on the input data, the sampling layer significantly reduces the spatial dimension of the original data while retaining important feature information. In addition, to adapt to the requirements of the subsequent feature extraction layer, the data sampling layer encodes the number of channels as 32, thereby providing a standardized input format for the subsequent modules.

[0037] The main function of the mask splicing unit is to enhance the robustness of feature extraction and the sensitivity of feature contours through an adaptive mask mechanism. This mechanism dynamically adjusts the input data to enhance the network's resistance to various input noises and interferences, while optimizing the extraction of edge and detail information of features. In addition, through the splicing operation, the mask splicing unit can effectively retain the original features of the data, ensuring that key information is not lost during the entire processing process, and thus improving the accuracy and precision of feature extraction. The downsampled data is first fed into the depth convolutional layer, whose main function is to extract the basic features of the data. This process ensures that the most basic patterns and rules can be captured during further processing, laying a foundation for subsequent deep feature extraction and ensuring the reliability of the feature extraction effect. Then, through the adaptive mask mechanism, a mask matrix that meets the mask rate requirements is generated and combined with the original data to complete the masking process of the data. The masked data will be fed into the convolutional layer to extract the features of the incomplete data. This stage can effectively capture the data missing parts caused by the mask mechanism. Then, through the multi-layer perceptron, the hidden features of the data are extracted, and through the residual structure, the original features can be effectively retained. This structure effectively alleviates the problem of information loss that may occur during the training of deep networks and ensures the stability of network performance. Finally, the initially input data and the processed data are spliced in the channel dimension. This splicing operation enables the data to retain rich context information during the processing process, which helps in the comprehensive understanding of features.

[0038] In the specific implementation, the parameter design of the mask concatenation unit fully considers the balance between the efficiency of feature extraction and model performance. The number of input channels of the module defaults to 64, and the initial value of the mask rate is set to 0.7, which is used to control the proportion of the masked area in the mask matrix. The mask rate is dynamically adjusted through a learnable parameter to ensure that the model can adaptively optimize the mask distribution according to the feedback during the training process. The downsampled data is first fed into the deep convolution layer, which uses a 3×3 convolution kernel with a step size of 1 and adopts a grouped convolution method to ensure that each channel is independently convolved. This design not only reduces the amount of calculation, but also enhances the spatial locality of the features. Subsequently, the data is normalized by a 3×3 convolution kernel, a convolution layer with a step size of 1, and a batch normalization layer to further optimize the data distribution characteristics. In order to further mine hidden features, the module introduces a multi-layer perceptron, whose hidden layer dimension is half of the number of input channels, and the original features are added to the hidden features through residual connections to ensure the integrity of the information. Finally, the originally input data and the processed data are concatenated in the channel dimension to form a feature representation containing rich contextual information.

[0039] The main function of the adaptive mask mechanism is to enhance the local feature representation through the generation and feature extraction of random mask matrices. A dynamic mask generator is constructed based on the two-dimensional spatial dimension of the input feature map. For the input feature map, a basic mask template is generated on the original high-width resolution space through parameterized initialization. The template randomly selects effective perception areas within the preset mask ratio range, randomly rearranges the spatial positions of the feature map, and forms a binary mask matrix that matches the spatial dimension of the input feature map. Subsequently, the generated mask matrix is multiplied pixel by pixel with the original input feature map to achieve selective masking of features in the specified spatial area. In order to maintain the integrity of the feature expression, the masked local features are weighted fused with the original features, and the contribution strength of the masked features is dynamically adjusted through a learnable scaling coefficient, ultimately forming a composite feature representation that has both global structure preservation and local feature enhancement. This process enables the model to adaptively strengthen the feature response of key areas while suppressing irrelevant background noise by dynamically adjusting the mask distribution pattern and feature fusion weights.

[0040] The two-dimensional maximum pooling unit is used to compress input data and reduce the consumption of computing resources. The data value of the data space dimension is compressed to one quarter of the original value through a maximum pooling layer.

[0041] The convolutional linear gating unit is used to coordinate local detail perception and global feature dependency, significantly enhancing the discriminability of feature representations. First, a dual-path parallel processing architecture is implemented for the input features. In the first branch, spatial feature extraction is performed through the cascaded operation of linear transformation and depth convolution, and fine-grained feature responses are generated by combining with a non-linear activation function. The second branch adopts a two-level linear layer stacking structure to capture high-order correlation features in the channel dimension through linear projection transformation. Subsequently, the results of the dual-path processing are subjected to cross-modal feature interaction and fusion, generating implicit abstract feature representations by element-wise multiplication. To enhance the feature expression ability, linear transformation and depth convolution operations are respectively applied to the fused features, and the outputs of the two are weighted and superimposed to form a composite feature with both local receptive field characteristics and global context awareness. Finally, by introducing a residual connection mechanism, the composite feature is adaptively fused with the original input feature, while retaining the integrity of the initial feature, dynamically enhancing the cross-channel information interaction ability.

[0042] The core function of the convolutional linear gating unit is to significantly enhance the discriminability of feature representations by coordinating local detail perception and global feature dependency. The module first implements a dual-path parallel processing architecture for the input features. In the first branch, fine-grained features in the spatial dimension are extracted through the cascaded operation of linear transformation and depth convolution. The input and output dimensions of the linear layer are both 32, and the depth convolution layer uses a 3×3 convolution kernel with a stride of 1 and adopts grouped convolution with the number of groups equal to the number of input channels, i.e., 32. The second branch adopts a two-level linear layer stacking structure to capture high-order correlation features in the channel dimension. The input and output dimensions of the linear layer are also 32, and a non-linear transformation is introduced through the Gaussian error linear unit. Subsequently, the results of the dual-path processing are subjected to cross-modal feature interaction and fusion by element-wise multiplication, generating implicit abstract feature representations. To enhance the feature expression ability, the fused features are respectively passed through linear transformation and depth convolution operations. The input and output dimensions of the linear layer are still 32, and the depth convolution layer also uses a 3×3 convolution kernel with a stride of 1 and the number of groups is 32. The outputs of the two are combined by weighted superposition to form a composite feature with both local receptive field characteristics and global context awareness. Finally, the composite feature is adaptively fused with the original input feature through the residual connection mechanism, while retaining the integrity of the initial feature and dynamically enhancing the cross-channel information interaction ability, thereby achieving efficient modeling and expression of complex features.

[0043] The multi-resolution attention parallel convolution module is used for extracting deep features from sensor data of different resolutions; the multi-resolution attention parallel convolution module includes a same-resolution data splicing unit, a parallel self-attention unit, and a parallel convolution unit connected in sequence; the same-resolution data splicing unit splices the 4 kinds of resolution data respectively generated from the X, Y, and Z axis data of the acceleration sensor, the X, Y, and Z axis data of the angular velocity sensor, and the X, Y, and Z axis data of the magnetometer, a total of 9-axis sensor data processed by the first data multi-resolution fusion convolution unit, according to the principle of the same resolution, so that the data of different resolutions all include 9-axis sensor data; the parallel self-attention unit automatically weights the important features in the data of different resolutions and gives more attention to the important features; the parallel convolution unit further extracts deep features from the data processed by the self-attention unit. The fusion attention convolution module is used for deep fusion of feature data of different resolutions; the fusion attention convolution module includes a second data multi-resolution fusion convolution unit, a fusion self-attention unit, and a fusion convolution unit connected in sequence; the calculation method of the second data multi-resolution fusion convolution unit is the same as that in the first data multi-resolution fusion convolution unit, and is used for fusing parallel feature data of different resolutions into feature data with a length of 80, the size of the original sensor data window; the calculation method of the fusion self-attention unit is the same as that of the above parallel self-attention unit, and is used for automatically extracting important features in the fused deep feature data; the fusion convolution unit further extracts deep features from the data features. The closed-set output module is used for outputting the closed-set classification recognition result and outputting the open-set feature vector. The closed-set output module includes a fully connected unit and a closed-set classification discrimination output unit connected in sequence; the fully connected unit is a multi-layer fully connected layer that transforms the data in the feature space; the closed-set classification discrimination output unit calculates through the Softmax layer and outputs the recognition result.

[0044] The fully connected classifier is used for classifying the extracted abstract features. First, global feature aggregation is performed on the high-dimensional feature map processed by the feature extraction network. The two-dimensional feature matrix is compressed into a one-dimensional channel feature vector through a two-dimensional global average pooling operation, while eliminating the spatial position sensitivity and retaining the discriminative information in the channel dimension. Subsequently, channel dimension calibration is performed on the dimensionality-reduced feature vector, and key channel features with category discrimination are screened through feature compression operations. Finally, the refined feature representation is input into a learnable linear projection layer to map the high-dimensional feature space to the target classification dimension, generating a probability distribution output that matches the preset number of categories. This classification architecture effectively suppresses redundant feature interference while reducing the model parameter quantity through a spatial-channel collaborative optimization mechanism, improving the reliability and robustness of classification decisions.

[0045] The data classification and recognition module includes a loss function calculation and update module and a classification and recognition output module connected in sequence; The loss function calculation and update module is used to calculate the mean and variance of the known-class behavior data during the training process, generate a loss function therefrom to update the parameters of the recognition model; during the model training phase, after each batch of forward inferences is completed, this module first receives the predicted probability distribution and the labeled ground truth output by the model, calculates the feature discrimination error of the current batch through the cross-entropy criterion. This error value characterizes the model's ability to distinguish class features, and then triggers the reverse error propagation mechanism to parse the contribution degree of each network parameter to the error layer by layer along the computational graph, generating the corresponding gradient tensor. The built-in adaptive optimizer in the module is synchronously started to dynamically smooth the original gradients generated in the current batch, constructing a continuous gradient correction amount. After each round of training, the constructed loss function is used to adjust the parameters of the recognition model, and this process is continuously repeated until convergence.

[0046] The classification and recognition output module is used to calculate the decision threshold for known classes and class mapping. In the class threshold determination unit, first, all the feature data points correctly classified in each class are extracted from the training data, and the distances between these data points and the class center mean point are calculated. Subsequently, the calculated distance values are normalized to limit the distance values within a unified range to eliminate the influence brought by the distribution difference. Based on the normalized distance distribution, a suitable statistic is selected as the decision threshold for this class. In the recognition output unit, the distances between the data points to be classified and the class center mean points of each class are calculated in turn, and these distance values are compared with the decision thresholds of the corresponding classes to determine the optimal matching class to which the data points belong. Finally, the determined class is mapped to the corresponding behavior label to generate the final output result. Through the above process, the module realizes the complete classification and mapping process from feature data to behavior labels.

[0047] The human body behavior information acquisition module includes a millimeter-wave radar transmitting unit, a millimeter-wave radar receiving unit, and a signal amplification and processing unit, which are integrated in the radar acquisition board. The acquisition board is installed at a specified position, and the original data of the intermediate-frequency signal after processing is transmitted by transmitting frequency-modulated continuous waves to detect the human body echo within the detection area.

[0048] The human body behavior information preprocessing module includes a signal storage unit, a window division unit, a Chirp matrix data reshaping unit, a distance Fourier transform unit, a moving target feature extraction unit, and a short-time Fourier transform unit connected in sequence, and is used to obtain the micro-Doppler map from the millimeter-wave radar raw data.

[0049] The signal storage unit is used to store the millimeter-wave radar behavior data collected in real time. These data contain the behavior characteristics of the user at different time points. The reflected signals captured by the radar sensor record the motion state of the human body in space.

[0050] The window division unit is used to divide the continuous long-time millimeter-wave radar data into several data segments. By setting appropriate window lengths and step sizes, the original data stream can be effectively divided into multiple independent data segments.

[0051] The Chirp matrix data reshaping unit is used to convert one-dimensional input data into a two-dimensional matrix form. Since the millimeter-wave radar uses the Chirp waveform, each Chirp wave corresponds to a specific time period and frequency range. Therefore, when processing, the one-dimensional data needs to be reorganized into a two-dimensional matrix according to the number of Chirp waves, so as to facilitate the subsequent generation and analysis of the micro-Doppler map. This reshaping process not only retains the time and frequency information of the original data, but also makes the data structure more regular, which is beneficial to subsequent calculations and visualizations.

[0052] The range Fourier transform unit is used to perform a Fourier transform on the two-dimensional original data to convert it into range-time data. This process can separate the target information at different ranges, enabling subsequent processing to more clearly identify the position changes of each target.

[0053] The moving target feature extraction unit is used to use a Butterworth filter to suppress the stationary or slowly changing background clutter. By filtering out these interference signals, the characteristics of dynamic targets can be highlighted, thereby improving the detection accuracy of moving targets.

[0054] The short-time Fourier transform unit is used to perform a short-time Fourier transform on the data after moving target feature extraction to obtain frequency-time data, that is, the micro-Doppler map. This transformation method can capture the speed changes of the target within a short period of time, further refining the description of the behavior characteristics.

[0055] The human behavior information application module faces various application platforms. The classification and recognition results are transmitted to the databases of each application platform in real time for storage and management, and the user's behavior is visualized and managed in real time.

[0056] The human behavior information transmission module is used to transmit the collected user behavior data to the edge device or cloud server through any one of the transmission methods of WiFi, serial communication, or Lora.

[0057] Embodiment 3: An identification method for an activity recognition system based on the Doppler characteristics of a millimeter-wave radar, which is implemented by the open-set human behavior recognition system based on multi-resolution fusion convolution described in any one of Embodiments 1 or 2. As shown in Figure 2 the following steps are included, taking the management application platform as an example: S1: Human behavior data acquisition; The millimeter-wave radar transmitting unit transmits a frequency-modulated continuous-wave signal to the target area through the transmitting antenna. The receiving antenna synchronously captures the echo signal reflected from the human body surface, and after being amplified by the low-noise amplifier, it enters the signal processing unit. At this stage, environmental noise is removed through time-domain filtering and frequency-domain filtering, and the processed intermediate-frequency data is received.

[0058] S2: Human behavior data transmission; The compressed original radar data packet is transmitted to the edge computing node or the cloud server through WiFi or LoRa. The transmission protocol uses UDP to reduce latency. The data packet header contains a timestamp, a device ID, and a check code to ensure the integrity of the transmission. If the signal strength is insufficient, serial communication can also be used for wired transmission and directly transmitted to the edge device for processing.

[0059] S3: Preprocessing of human behavior data; The collected radar data of the user is successively segmented, reshaped, Doppler processed, visualized, normalized, and labeled, so as to obtain a behavior micro-Doppler map with dynamic energy feature information and labels.

[0060] S4: Micro-Doppler feature extraction and identification; The micro-Doppler map preprocessed in S3 is input into the recognition model in batches, and a fully connected layer classifier is obtained through training; The fully connected layer classifier includes a GlobalAvgPool2D two-dimensional global average pooling layer and a Linear linear layer connected in sequence; S5: Loss function calculation and network update; In S4, the micro-Doppler feature extraction module generates a feature vector through forward propagation and inputs it into the data classification and recognition module for processing. In order to achieve efficient training and optimization, the cross-entropy loss function is used as the main optimization objective, and at the same time, the Adam optimizer is combined to update the model parameters. Specifically, the cross-entropy loss function is used to measure the difference between the model prediction probability distribution and the true label, and its mathematical expression form is shown in Equation (1): (1); Where: represents the cross-entropy loss value, is the total number of categories, Is the one-hot encoded vector of the true label, Is the probability distribution predicted by the model; In each round of training, the data is input into the model micro-Doppler feature extraction module in batches. The data volume of each batch is set to B, and B is any positive integer less than the total data volume. The Adam optimizer dynamically adjusts the learning rate according to the gradient information of the current batch, so as to efficiently update the model parameters. The core idea of the Adam optimizer is to combine the advantages of the momentum method and the adaptive learning rate. Its parameter update rules are shown in equations (2), (3), (4) and (5): (2); (3); (4); (5); Where: Represents the iteration step, Represents the gradient of the current batch, And Are the first-order moment estimate and the second-order moment estimate of the gradient respectively, And Are the bias correction terms for the original first-order moment (mean) and second-order moment (variance) estimates, And Are the exponential decay rates, Is the initial learning rate, Is a small constant to prevent division by zero, Are the model parameters; During the training process, the model is first supervised and trained using the cross-entropy loss function and the Softmax function to ensure that the model can accurately predict the labels of known classes. Finally, through continuous iterative optimization, the model gradually converges to the optimal state, thus showing higher classification accuracy and robustness on the test set.

[0061] S6: Determine the threshold and the classification discrimination output; After the micro-Doppler feature extraction module is trained, in order to improve the discrimination ability in practical applications, the central mean of the feature vectors of known classes is statistically calculated, and a confidence threshold is set accordingly to judge the reliability of the prediction results during output. Specifically, first for each class Of the feature vector set is processed to calculate its class central mean vector , As shown in equation (6): (6); Where: Represents the class Of the central mean vector, For the category the number of samples in indicating the category the th sample's feature vector in the feature space; On this basis, further calculate the distances from all samples in each category to the mean of their category centers, and statistically calculate the mean and standard deviation of these distances, denoted as and respectively, as shown in Equations (7) and (8): (7); (8); Where: represents the Euclidean distance from the feature vector to the mean of the category center ; To set the confidence threshold, use the mean distance plus a certain multiple of the standard deviation to define the confidence range for each category, as shown in Equation (9): (9); Where: represents the confidence threshold of the category , is a hyperparameter used to control the strictness of the threshold, usually adjusted according to actual needs; In the model inference stage, for the input sample , the model first predicts its belonging category , and then calculates the distance from this sample to the mean of the predicted category center. If this distance is less than or equal to the corresponding confidence threshold , the prediction result is considered reliable; otherwise, the prediction result is considered unreliable, belonging to an unknown category or requiring further confirmation; this method can effectively improve the robustness of the model in the open-set scenario and reduce the risk of misclassification at the same time.

[0062] Step S7: Apply the discrimination output result; The results output by the closed set and the results output by the open set are transmitted to the corresponding application platforms in real time.

[0063] The specific implementation process of the above S3 is as follows: S31: Reshape the human behavior information data; When processing the intermediate-frequency raw millimeter-wave radar data, in order to facilitate subsequent feature extraction and model training, it is first necessary to rearrange the continuous radar signals into a two-dimensional matrix form according to Chirp. The raw millimeter-wave radar data is expressed as where represents the intermediate frequency signal value of the th sampling point, is the total number of sampling points. At the same time, each Chirp contains sampling points, and there are a total of Chirps. The total number of sampling points of the original data satisfies ; The original data S is first divided into multiple Chirps, and each Chirp corresponds to a sequence of sampling points with a length of . These Chirps are arranged in order to form a two-dimensional matrix , whose dimension is . This two-dimensional matrix is expressed as: (10); where: represents the th sampling point value of the th Chirp; S32: Radar data window segmentation; Based on the two-dimensional matrix generated in S31, the data is further segmented according to the time window. The window length defines the number of Chirps contained in each window, and the step size determines the overlap degree or interval between adjacent windows. The starting Chirp and ending Chirp of the th window can be calculated by equations (11) and (12): (11); (12); where: represents the segment number, and its value range is , represents the starting position of the th segment, represents the ending position of the th segment, and satisfies ; represents the fixed length of each window, represents the actual length of the The total number of finally generated window segments can be determined by equation (13): (13); S33: Micro-Doppler processing; In the window matrix , each row represents the sampled point data of a Chirp. By performing a fast Fourier transform on each row, the time-domain signal is converted into a frequency-domain signal, thereby extracting the distance information of the target, as shown in Equation (14): (14); Where: is the imaginary unit, is the value at the th row and the th column in the input window matrix, represents the th complex value at the th range cell, represents the frequency at which the signal is sampled, is the number of samples per Chirp and also the number of points of the FFT; After the fast Fourier transform, the range dimension data contains static clutter. To highlight the dynamic target signal and suppress the static clutter, a high-pass Butterworth filter is used to perform moving target indication filtering on the range dimension data. For discrete-time signals, the difference equation form of the Butterworth filter is shown in Equation (15): (15); Where: is the input signal, is the output signal, and are the forward and feedback coefficients of the filter, respectively; and represent the forward order and the feedback order of the filter; After the moving target indication filtering, the range dimension data is shown in Equation (16): (16); Where: represents the range dimension data after Butterworth filtering, represents the Butterworth filtering operation, The original range dimension data, the value at the th range cell and the th time sample, represents the forward coefficient vector, represents the feedback coefficient vector; To further extract the micro-Doppler features of the target, the data after the moving target indication filtering is subjected to a short-time Fourier transform. The core idea of the short-time Fourier transform is to divide the signal into multiple short-time windows and perform FFT within each window, thereby obtaining the time-frequency distribution of the signal as shown in Equation (17): (17); Wherein: is the imaginary unit, represents the input signal, represents the window function, represents the time-frequency distribution, represents the center of the local window for the current analysis, represents the frequency component for the current analysis; After being processed by the short-time Fourier transform, the micro-Doppler spectrogram has a dimension of , where represents the frequency-axis resolution, represents the time-axis resolution; Finally, the micro-Doppler spectrogram is represented by Equation (18): (18); Wherein: represents the energy distribution at frequency and time ; S34: Visualization of micro-Doppler features; To visually display the micro-Doppler features, the micro-Doppler spectrogram described in S33 is converted into a logarithmic form and visualized through an image drawing tool, and label information is added by naming.

[0064] The specific implementation process of the described S4 is as follows: S41: Data downsampling; The micro-Doppler feature map input to S34 is first processed by a data downsampling module to reduce its spatial resolution and expand the number of channels. This process can be represented by Equation (19): (19); Wherein: represents the output of the data downsampling module, is the Gaussian error linear unit, represents the input feature map, represents the block size for downsampling, which is set to 4; the inputs of the function are the feature map, the convolution kernel size, and the stride, and these parameters jointly determine the downsampling ratio; After the convolution operation, the spatial resolution of the feature map is effectively compressed, and layer normalization is introduced to ensure the stability of the numerical distribution. The Gaussian error linear unit further performs a non-linear transformation on the features to enhance the expression ability of the model. Finally, the number of channels of the feature map is expanded to 32, thereby reducing the computational complexity while retaining the key feature information and laying a foundation for subsequent feature extraction.

[0065] S42: Micro-Doppler feature extraction; After the downsampled feature map, the input feature map changes from to , which compresses the resolution of the feature map and can effectively reduce the computational load and improve the computational speed in subsequent calculations. The data processed by S41 is sent to four identical stages for layer-by-layer feature extraction. Each stage is composed of a mask splicing module, a two-dimensional max pooling, and a convolutional linear gating unit connected in sequence. This process is expressed as Equation (20): (20); Where: represents the output of this stage, represents the operation of the convolutional linear gating unit, represents the operation of the mask splicing unit, represents the two-dimensional max pooling operation; during the micro-Doppler feature extraction process, after the input feature map is processed by each stage, its spatial resolution is halved, while the number of channels doubles. In each stage, the two-dimensional max pooling operation halves the spatial resolution by downsampling the feature map. The mask splicing module doubles the number of channels by splicing the front and back data.

[0066] After four stages of layer-by-layer processing, the input feature map is converted into a high-level feature representation for the final classification or regression task. This process can be expressed as shown in Equation (21): (21); Where: represents the extracted high-level features, usually used to predict the target activity or behavior category; is the initial feature map processed by the data downsampling module; represents the operation after four stages of processing. The size of the feature map changes from to ; S43: Classification and recognition.

[0067] The fully connected layer classifier includes a GlobalAvgPool2D two-dimensional global average pooling layer and a Linear linear layer connected in sequence.

[0068] In S42, through the upsampling and downsampling methods again, the fusion convolution calculation of different resolutions is carried out, and the feature extraction of the sensor uniaxial data at different resolutions is further carried out. The specific implementation process is as follows: S421: Mask splicing unit; The mask splicing unit captures spatio-temporal features through multi-stage feature fusion and enhances the perception ability of key regions. Let the input feature map be , first, a depth convolutional layer with a convolutional kernel and a stride of 1 is used to extract basic features. After batch normalization, intermediate features are obtained . The adaptive masking mechanism uses a dynamically generated binary masking matrix to perform regional masking on . The process is shown in Equation (22): (22); where: is a learnable scaling factor, and the masking matrix is generated under the control of the initial masking rate ; Specifically, it is implemented through random permutation and matrix reshaping to ensure random masking of non-critical regions. A convolutional layer with a convolutional kernel and a stride of 1 is used to extract masking features ; The masked features are compressed by a multi-layer perceptron and then added to the original input to generate enhanced features as shown in Equation (23): (23); where: represents a multi-layer perceptron, which is a feed-forward neural network structure used for non-linear transformation of input data; The final module output is achieved by concatenating the original input and the enhanced features along the channel dimension, as shown in Equation (24): (24); The principle of the adaptive masking generation mechanism is shown in Equations (25) and (26): (25); (26); where: represents the initial masking vector, whose length is equal to the initial resolution , and the random permutation function is used to randomly shuffle the elements in , and is used to reshape the one-dimensional vector into a matrix with the same shape as the target feature map ; In , the first positions are set to 1, and the remaining positions are set to 0, representing the selective activation region of the initial mask. After random permutation and shape reshaping, the generated masking matrix has the same shape as the input feature map The same spatial dimension can be used for subsequent masking operations or feature selection tasks.

[0069] S422: Convolutional Linear Gated Unit; The convolutional linear gated unit enhances the feature expression ability through cross-channel interaction. The input features are respectively passed through two independent linear layers and After projection, one branch extracts spatial domain features through a 3×3 depth convolution, and modulates them element-wise with the gating weights of the other branch to generate implicit abstract features , as shown in Equation (27): (27); Where: is the Gaussian Error Linear Unit, represents a 3×3 depth convolution with a stride of 1, is element-wise multiplication. After this feature undergoes a second depth convolution and a linear transformation, a triple summation operation is performed with the original input, as shown in Equation (28): (28); Where: is the linear layer, is element-wise addition, represents the output of the convolutional linear gated unit.

[0070] By using a gating mechanism to fuse multi-scale features, a 3×3 convolutional kernel realizes local context modeling, and the cross-channel summation operation retains the original features while strengthening semantic associations. Compared with traditional attention mechanisms, this design significantly reduces complexity through parameter sharing and parallel computing.

[0071] The specific embodiments of the present invention disclosed above are only examples. However, the present invention is not limited thereto, and any changes that can be conceived by those skilled in the art should fall within the protection scope of the present invention.

Claims

1. An activity recognition system based on the Doppler characteristics of millimeter-wave radar, characterized in that, It includes a human behavior radar acquisition module, a human behavior information transmission module, a human behavior information preprocessing module, a micro-Doppler feature extraction module, a data classification and recognition module, and a human behavior information application module connected in sequence: The human behavior information acquisition module is used to collect the reflected signals in the target area. It uses a millimeter-wave radar to transmit high-frequency millimeter-wave signals, and obtains the information of human actions by receiving the signals reflected from the human body surface; The human behavior information transmission module is used to transmit the collected behavior data of the user to the edge device or the cloud server; The human behavior information preprocessing module is used to receive the radar signal data transmitted from the human behavior information transmission module, perform formatting processing, and then use Fourier transform, moving target feature extraction, and short-time Fourier transform to analyze the frequency distribution of the signal, and extract the features related to human behavior from it; The micro-Doppler feature extraction module is used to input the behavior data of the user preprocessed by the human behavior information preprocessing module into the trained micro-Doppler feature extraction module for advanced feature extraction; The data classification and recognition module is used to train and accurately discriminate the behavior data input into the recognition model; The human behavior information application module is used to transmit the obtained human behavior classification and recognition results to the user information interface, so as to realize the display of human information.

2. The activity recognition system based on the Doppler characteristics of millimeter-wave radar according to claim 1, characterized in that: The micro-Doppler feature extraction module adopts a four-layer pyramid stacking structure, which is a data downsampling unit, a four-layer feature extraction unit, and a fully connected layer classifier connected in sequence. Each layer of the feature extraction unit includes a mask splicing unit, a two-dimensional maximum pooling unit, and a multi-convolution linear gating unit connected in sequence.

3. The activity recognition system based on the Doppler characteristics of millimeter-wave radar according to claim 2, wherein: The data classification and recognition module includes a loss function calculation and update module and a classification and recognition output module connected in sequence; The loss function calculation and update module is used to calculate the mean and variance of the behavior data of known categories during the training process, and generate a loss function to update the parameters of the recognition model; The classification and recognition output module is used to calculate the decision threshold for known categories and category mapping.

4. The activity recognition system based on the Doppler characteristics of millimeter-wave radar according to claim 3, characterized in that: The human behavior information acquisition module includes a millimeter-wave radar transmitting unit, a millimeter-wave radar receiving unit, and a signal amplification and processing unit, which are integrated in the radar acquisition board. The acquisition board is installed at a specified position, and the original data of the intermediate-frequency signal after processing is transmitted by transmitting a frequency-modulated continuous wave to detect the human echo within the detection area range.

5. The activity recognition system based on the Doppler characteristics of millimeter-wave radar according to claim 4, characterized in that: The human behavior information preprocessing module includes a signal storage unit, a window division unit, a Chirp matrix data reshaping unit, a range Fourier transform unit, a moving target feature extraction unit, and a short-time Fourier transform unit connected in sequence, and is used to obtain the micro-Doppler map from the millimeter-wave radar raw data.

6. The activity recognition system based on the Doppler characteristics of millimeter-wave radar according to claim 1, characterized in that: The human behavior information transmission module is used to transmit the collected behavior data of the user to the edge device or the cloud server through any one of the transmission methods of WiFi, serial communication, or Lora.

7. A recognition method using the activity recognition system based on the Doppler characteristics of millimeter-wave radar described in claim 5, characterized in that, It includes the following steps: S1: Human behavior data acquisition; S2: Human behavior data transmission; S3: Human behavior data preprocessing; S4: Micro-Doppler feature extraction and recognition; The micro-Doppler maps after S3 preprocessing are input into the recognition model in batches, and a fully connected layer classifier is obtained through training; The fully connected layer classifier includes a GlobalAvgPool2D two-dimensional global average pooling layer and a Linear linear layer connected in sequence; S5: Loss function calculation and network update; In S4, the micro-Doppler feature extraction module generates feature vectors through forward propagation and inputs them into the data classification and recognition module for processing. To achieve efficient training and optimization, the cross-entropy loss function is used as the main optimization objective, and at the same time, the Adam optimizer is combined to update the model parameters. Specifically, the cross-entropy loss function is used to measure the difference between the model prediction probability distribution and the true label, and its mathematical expression form is shown in Equation (1): (1); Wherein: represents the cross-entropy loss value, is the total number of categories, is the one-hot encoded vector of the true label, is the probability distribution predicted by the model; In each round of training process, the data is input into the micro-Doppler feature extraction module of the model in batches. The data volume of each batch is set to B, and B is any positive integer less than the total data volume. The Adam optimizer dynamically adjusts the learning rate according to the gradient information of the current batch, thereby efficiently updating the model parameters. The core idea of the Adam optimizer is to combine the advantages of the momentum method and the adaptive learning rate, and its parameter update rules are shown in Equations (2), (3), (4), and (5): (2); (3); (4); (5); Wherein: represents the iteration step number, represents the gradient of the current batch, and are the first - order moment estimate and the second - order moment estimate of the gradient respectively, and are the bias correction terms for the original first - order moment and second - order moment estimates, and is the exponential decay rate, is the initial learning rate, is a small constant to prevent division by zero, are the model parameters; S6: Determine the threshold and classification discriminant output; After the training of the micro-Doppler feature extraction module is completed, the central mean of the feature vectors of known classes is statistically calculated, and a confidence threshold is set based on this to judge the reliability of the prediction result during output. First, for each class process the set of feature vectors to calculate its class central mean vector , as shown in Equation (6): (6); Wherein: represents the central mean vector of the class , is the number of samples in the class , represents the feature vector of the -th sample in the feature space of the class ; Further calculate the distances from all samples in each category to the mean of their category centers, and statistically calculate the mean and standard deviation of these distances, denoted as and , as shown in Equations (7) and (8): (7); (8); Wherein: represents the Euclidean distance from the feature vector to the class center mean; To set the confidence threshold, the distance from the mean plus a certain multiple of the standard deviation is used to define the confidence range for each category, as shown in Equation (9): (9); Wherein: represents the confidence threshold of the category , and is a hyperparameter; ​ During the model inference stage, for the input sample , the model first predicts the class it belongs to , and then calculates the distance from this sample to the central mean of the predicted class . If this distance is less than or equal to the corresponding confidence threshold , the prediction result is considered reliable; otherwise, the prediction result is considered untrustworthy and belongs to an unknown class or requires further confirmation; Step S7: Apply the discriminant output result; The results output by the closed set and the results output by the open set are transmitted to the corresponding application platforms in real time.

8. The recognition method according to claim 7, wherein: The specific implementation process of S3 is as follows: S31: Reshape the human behavior information data; When processing the raw millimeter-wave radar data at intermediate frequency, in order to facilitate subsequent feature extraction and model training, it is first necessary to rearrange the continuous radar signals into a two-dimensional matrix form according to Chirp. The raw millimeter-wave radar data is represented as , where represents the intermediate frequency signal value of the th sampling point, is the total number of sampling points. At the same time, each Chirp contains sampling points, and there are a total of Chirps. The total number of sampling points of the original data satisfies ; The original data S is first divided into multiple Chirps, each Chirp corresponding to a sequence of sampling points with a length of , and these Chirps are arranged in order to form a two-dimensional matrix , whose dimension is , and this two-dimensional matrix is represented as: (10); Wherein: represents the th sampling point value of the th Chirp; S32: Segment the radar data window; The two-dimensional matrix generated in S31 , based on this, the data is further segmented according to a time window, and the window length defines the number of Chirps included in each window, and the step size determines the overlapping degree or interval between adjacent windows. The starting Chirp and ending Chirp of the th window can be calculated by equations (11) and (12): (11); (12); Wherein: represents the segment number, and the value range is , represents the starting position of the th segment, represents the ending position of the th segment, and satisfies ; represents the fixed length of each window, Indicates the actual length of the section; The total number of finally generated window segments can be determined by Equation (13): (13); S33: Micro-Doppler processing; The window matrix obtained in step S32 , where each row represents the sampled point data of a Chirp. By performing a fast Fourier transform on each row, the time-domain signal is converted into a frequency-domain signal, thereby extracting the distance information of the target, as shown in Equation (14): (14); Wherein: is the imaginary unit, is the value at the th row and the th column in the input window matrix, represents the th complex value on the th distance cell, represents the frequency for sampling the signal, is the number of sampling points for each Chirp and also the number of points for FFT; After fast Fourier transform, the range dimension data contains static clutter. To highlight the dynamic target signal and suppress the static clutter, a high-pass Butterworth filter is used to perform moving target indication filtering on the range dimension data. For discrete-time signals, the difference equation form of the Butterworth filter is shown in Equation (15): (15); Wherein: is the input signal, is the output signal, and are the forward and feedback coefficients of the filter respectively; and represent the forward order of the filter and the feedback order of the filter; After moving target indication filtering, the range dimension data is shown in Equation (16): (16); Wherein: represents the range dimension data after Butterworth filtering, represents the Butterworth filtering operation, the original range dimension data, at the th range bin, the value at the th time sample, represents the forward coefficient vector, represents the feedback coefficient vector; Perform short-time Fourier transform on the data after moving target indication filtering. The core idea of short-time Fourier transform is to divide the signal into multiple short-time windows and perform FFT within each window, so as to obtain the time-frequency division of the signal as shown in Equation (17): (17); Wherein: is the imaginary unit, represents the input signal, represents the window function, represents the time-frequency distribution, represents the center of the local window for the current analysis, represents the frequency component for the current analysis; After short-time Fourier transform processing, the micro-Doppler spectrogram has a dimension of , where represents the frequency-axis resolution, and represents the time-axis resolution; Finally, the micro-Doppler map is represented by Equation (18): (18); Wherein: represents the energy distribution at frequency and time; S34: Visualization of micro-Doppler features; Convert the micro-Doppler map described in S33 into a logarithmic form and perform visualization through an image drawing tool, and add label information by naming.

9. The recognition method according to claim 7, wherein: The specific implementation process of S4 is as follows: S41: Data downsampling; The micro-Doppler feature map input in S34 is first processed by the data downsampling module to reduce its spatial resolution and expand the number of channels. This process is represented by Equation (19): (19); Wherein: represents the output of the data downsampling module, is the Gaussian error linear unit, represents the input feature map, represents the block size for downsampling, and the inputs of the function are the feature map, the convolutional kernel size, and the stride respectively; S42: Micro-Doppler feature extraction; The data processed in S41 is transmitted to four identical stages for layer-by-layer feature extraction. Each stage consists of a mask splicing module, two-dimensional maximum pooling, and a convolutional linear gated unit connected in sequence. This process is represented by Equation (20): (20); Wherein: represents the output of this stage, represents the operation of the convolutional linear gating unit, represents the operation of the mask splicing unit, represents the two-dimensional max pooling operation; After four stages of layer-by-layer processing, the input feature map is converted into a high-level feature representation for the final classification or regression task, as shown in Equation (21): (21); Wherein: represents the extracted high-level features, which is the initial feature map processed by the data downsampling module; represents the operations processed through four stages; S43: Classification and recognition.

10. The recognition method according to claim 9, wherein: The specific implementation process of S42 is as follows: S421: Mask splicing unit The mask splicing unit captures spatio-temporal features through multi-stage feature fusion and enhances the perception ability of key regions. Let the input feature map be , first, a depth convolution layer with a convolution kernel and a stride of 1 is used to extract basic features. After batch normalization, the intermediate feature is obtained. The adaptive mask mechanism uses the dynamically generated binary mask matrix to implement regional occlusion, and its process is shown in Equation (22): (22); Wherein: is a learnable scaling factor, and the mask matrix is generated under the control of an initial masking rate ; Features after masking After being compressed by the multi-layer perceptron and the original input Perform feature addition to generate enhanced features As shown in Equation (23): (23); Wherein: represents a multi-layer perceptron, which is a feedforward neural network structure used for non-linearly transforming input data; Final module output It is achieved by concatenating the original input and the enhanced features along the channel dimension, as shown in Equation (24): (24); The principle of the adaptive mask generation mechanism therein is as shown in Equations (25) and (26): (25); (26); Wherein: represents an initial mask vector, the length of which is equal to the initial resolution , a random permutation function for randomly shuffling the elements in for reshaping a one-dimensional vector into a matrix with the same shape as the target feature map ; S422: Convolutional linear gating unit The convolutional linear gated unit enhances the feature representation ability through cross-channel interaction, and the input features are respectively passed through two independent linear layers and After projection, one branch extracts spatial features through 3×3 depth convolution and modulates them element-wise with the gating weights of the other branch to generate implicit abstract features , as shown in Equation (27): (27); Wherein: is a Gaussian error linear unit, represents a 3×3 depth convolution with a stride of 1, is element-wise multiplication. After the feature undergoes two-depth convolution and linear transformation, a triple summation operation is performed with the original input, as shown in Equation (28): (28); Wherein: is a linear layer, is element-wise addition, represents the output of the convolutional linear gated unit.

Citation Information

Patent Citations

  • Distributed optical fiber vibration signal feature extraction and identification method

    CN111242021A

  • Multi-complexity behavior recognition system and method based on adaptive feature extraction

    CN116956222A

  • Lightweight human body posture recognition method and system based on indoor millimeter wave radar

    CN117310646A

  • Human respiration and heartbeat frequency detection method and system based on millimeter wave radar in multi-target scene

    CN118276082A

  • Personnel perception quantity statistical method and system based on millimeter wave radar

    CN119397198A

Cited By

  • Radar target template signal establishment method based on depth model adaptive segmentation

    CN120849932A

  • A human body fall risk early warning method, system, device and storage medium

    CN121354282B