Millimeter wave radar human body behavior recognition system and method based on global pixel characteristics

By decomposing the three-dimensional convolution in the millimeter-wave radar human behavior recognition system into two-dimensional and one-dimensional convolutions, combined with the spatiotemporal attention and feature separation module of global pixel features, the problem of mixing time and space features is solved, the recognition accuracy is improved and the calculation amount is reduced, and the effective recognition of complex human behavior is achieved.

CN120491016AActive Publication Date: 2025-08-15SHANDONG UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510976315.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-08-15
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

The existing millimeter-wave radar human behavior recognition technology is difficult to fully explore the time-related information in the multi-frame distance-Doppler feature map, and the time and spatial characteristics are mixed, resulting in low recognition accuracy and large model calculations, making it difficult to achieve effective recognition of complex human behaviors.

Method used

The millimeter-wave radar human behavior recognition system based on global pixel features is adopted. By decomposing traditional three-dimensional convolution into spatial domain two-dimensional convolution and time domain one-dimensional convolution, we explicitly distinguish between time and space features, and use the spatiotemporal attention module and spatiotemporal feature separation module based on global pixel features to reduce the calculation amount and the number of parameters and improve the recognition accuracy.

Benefits of technology

It improves the recognition accuracy of confusing behaviors, reduces the computational complexity, realizes effective recognition of complex human behaviors, and improves the recognition performance of existing models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120491016A_ABST
    Figure CN120491016A_ABST
Patent Text Reader

Abstract

The invention discloses a millimeter wave radar human body behavior recognition system and method based on global pixel characteristics, and relates to the technical field of human body behavior recognition. Comprising a millimeter-wave radar original data acquisition module, a millimeter-wave radar original data transmission module, a millimeter-wave radar original data preprocessing module, a global pixel feature-based time extraction network distance-Doppler feature extraction and classification module and a human body behavior information application module which are connected in sequence. The millimeter wave radar original data acquisition module comprises a millimeter wave radar transmitting unit, a receiving unit, a signal mixing unit and a power amplification unit. The technical problem to be solved by the invention is to provide the millimeter wave radar human body behavior recognition system and method based on the global pixel characteristics, so that the calculation amount and the number of parameters are reduced, the recognition accuracy of easily-confused behaviors is improved, the effective recognition of complex human body behaviors is realized, and the recognition performance of an existing model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human behavior recognition, and in particular to a millimeter wave radar human behavior recognition system and method based on global pixel features. Background Art

[0002] Human Activity Recognition (HAR) is a technology that uses sensors to collect physiological information about human motion and posture, and then processes and analyzes this data using artificial intelligence algorithms to automatically identify human behavior. The core goal of HAR technology is to accurately classify and predict different behavioral patterns by extracting and modeling human behavioral data, providing decision support for intelligent applications.

[0003] HAR technology has demonstrated widespread application value in multiple fields in recent years, becoming a key technology driving the development of an intelligent society. In the smart home sector, HAR technology uses sensors to capture real-time information about family members' activities, such as monitoring children for dangerous areas, and analyzes daily behavior patterns to provide users with personalized health management solutions and safety assurance. In intelligent assisted driving scenarios, HAR technology can monitor driver status, such as fatigue or inattention, and identify passenger behavior to enhance driving safety. In the medical and rehabilitation fields, HAR technology can be used to detect in real time whether patients are engaging in dangerous behaviors or abnormal conditions, such as sudden falls or abnormal body posture, thereby issuing timely alerts to prevent accidents. Furthermore, HAR technology can remotely and continuously monitor a patient's health status, including activity levels, sleep quality, and changes in daily behavior patterns, providing comprehensive data support for medical staff to more effectively assess and manage patients' health. With the development of artificial intelligence, the application of HAR technology in human-computer interaction and virtual reality is rapidly expanding. HAR technology enables users to use gestures to naturally control smart homes, game interfaces, or virtual environments, providing a more intuitive way for human-computer interaction. HAR technology can also infer a user's emotional state based on motion and facial expressions, giving intelligent assistants or robots more human-like interaction capabilities. Furthermore, in education, HAR systems can monitor students' focus, engagement, and emotional changes in real time, helping teachers optimize teaching strategies and improve classroom effectiveness. These applications not only enhance the intelligence of interactions but also bring innovative user experiences to a wide range of fields.

[0004] Based on the sensor type and data collection method, HAR technology can be divided into three main categories: wearable device-based HAR, video-based HAR, and radio frequency-based HAR. Radio frequency devices include radar and Bluetooth.

[0005] With advancements and iterations in related technologies, HAR algorithms and systems have entered a period of rapid development. Currently, research on HAR primarily focuses on classification models based on deep neural networks. Deep neural networks can automatically extract behavioral features, enabling end-to-end behavior recognition and effectively improving recognition accuracy. Convolutional neural networks (CNNs) are one of the most widely used deep neural networks. CNNs are multi-layered, deep feedforward networks that process input data layer by layer, integrating information and transforming raw data into high-level feature representations more closely linked to the output targets. Finally, a classifier completes label mapping. Compared to CNNs, recurrent neural networks (RNNs) focus more on the temporal characteristics of data and can capture correlations between these features. Therefore, RNNs and their variants, such as long short-term memory (LSTM) and gated recurrent units (GRU), are widely used in HAR model construction.

[0006] Radar-based HAR technology has garnered widespread attention in recent years, but many challenges remain. When extracting spatiotemporal features, many existing networks struggle to fully exploit the temporal correlations in multi-frame range-Doppler maps, and the mixing of temporal and spatial features remains unresolved. In network design, many algorithms prioritize high recognition accuracy over lightweight model design. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a millimeter-wave radar human behavior recognition system and method based on global pixel features. The method can solve the problem of insufficient extraction of time correlation information in multi-frame range-Doppler feature maps, decompose the traditional three-dimensional convolution into two-dimensional convolution in the spatial domain and one-dimensional convolution in the time domain, explicitly distinguish the two features, avoid feature mixing, and at the same time reduce the amount of calculation and the number of parameters, improve the recognition accuracy of easily confused behaviors, and realize the effective recognition of complex human behaviors, thereby improving the recognition performance of existing models.

[0008] The present invention adopts the following technical solutions to achieve the invention objectives: A millimeter-wave radar human behavior recognition system and method based on global pixel features, characterized by comprising: a millimeter-wave radar raw data acquisition module, a millimeter-wave radar raw data transmission module, a millimeter-wave radar raw data preprocessing module, a time extraction network distance-Doppler feature extraction and classification module based on global pixel features, and a human behavior information application module connected in sequence; The millimeter wave radar raw data acquisition module includes a millimeter wave radar transmitting unit, a receiving unit, a signal mixing unit and a power amplification unit; The millimeter wave radar raw data preprocessing module includes a raw data storage unit, a raw data partitioning unit, a chirp matrix reshaping unit, a range-dimensional fast Fourier transform unit, a moving target display unit, a velocity-dimensional fast Fourier transform unit, and a range-Doppler characteristic map visualization unit; The temporal extraction network based on global pixel features, the distance-Doppler feature extraction and classification module, includes a spatial downsampling unit, a spatio-temporal attention module based on global pixel features (STAM-GPF), a spatio-temporal feature separation module (STFSM), and a fully connected layer classifier.

[0009] As a further limitation of the present technical solution, the millimeter-wave radar original data transmission module supports multiple transmission protocols such as Wi-Fi, serial communication, Ethernet, LoRa and ZigBee.

[0010] As a further limitation of the present technical solution, the spatiotemporal attention module based on global pixel features includes a spatial downsampling unit, a self-attention unit, a global mean filtering unit and a residual link unit connected in sequence; The spatiotemporal feature separation module includes a channel downsampling unit, a spatial feature extraction unit, a temporal feature extraction unit and a residual connection unit.

[0011] As a further limitation of the present technical solution, the fully connected layer classifier of the time extraction network distance-Doppler feature extraction and classification module based on global pixel features maps the features into specific behavior category probabilities through the normalized exponential function Softmax to identify different behaviors.

[0012] As a further limitation of this technical solution, the human behavior information application module includes smart medical care, smart transportation and smart home scenarios.

[0013] A recognition method for a millimeter-wave radar human behavior recognition system based on global pixel features comprises the following steps: S1: Millimeter-wave radar raw data acquisition: The millimeter-wave radar in the millimeter-wave radar raw data acquisition module collects raw data of user behavior. The raw data is in the form of IQ data and is stored in the form of a binary file; S2, millimeter wave radar raw data transmission, transmitting the raw data collected by the millimeter wave radar raw data acquisition module to the raw data storage unit of the millimeter wave radar raw data preprocessing module via an Ethernet cable; S3. Millimeter-wave radar raw data preprocessing: The raw data storage unit of the millimeter-wave radar raw data preprocessing module cuts the radar raw data into frames with a 25% overlap rate. Each frame contains 128 chirps. S4: Construct a time extraction network distance-Doppler feature extraction and classification module based on global pixel features to perform human behavior recognition and classification; the 6 frames of distance-Doppler feature maps that have completed feature preprocessing are input into the time extraction network distance-Doppler feature extraction and classification module based on global pixel features in batches, and then the spatiotemporal attention module based on global pixel features is used to perform temporal modeling, and the spatiotemporal feature separation module is used to extract and separate temporal and spatial features. Finally, the fully connected layer classifier outputs the discrimination result; S5: Construct a spatiotemporal attention module based on global pixel features to perform temporal modeling on the input feature map and fully exploit the temporal correlation features between multi-frame range-Doppler feature maps; S6: Construct a spatiotemporal feature separation module to extract temporal and spatial features from the input feature map, which can reduce the impact of feature confounding. S7: Output action recognition results. After multi-dimensional spatiotemporal feature extraction and fusion, the obtained feature data is input into the fully connected layer. The number of hidden units in the fully connected layer is the number of behavior categories contained in the data. The data after the fully connected layer is passed through the Softmax classifier to calculate the probability of the corresponding behavior. The behavior type with the highest probability is the final judgment result. S8: Display and application of human behavior information. Through the human behavior information application module, the behavior recognition results are visualized and used in the fields of smart medical care, smart transportation and smart home.

[0014] As a further limitation of this technical solution, S3 includes: S31: Data framing and chirp matrix reconstruction; Assume that the original signal sampling rate is , each frame duration , the inter-frame overlap rate is 25%, so the frame division formula is shown in formula (1): (1); in: , is the number of sampling points in a single frame, is the number of frames; Will M The frame data is reconstructed into a two-dimensional matrix with the dimension of , each column is a sampling sequence of a single chirp; S32: distance-dimensional fast Fourier transform; perform a 128-point fast Fourier transform on each column of chirp data, and the calculation method is shown in formula (2): (2); in: is the result matrix of fast Fourier transform, is the input signal matrix, is the rotation factor, , Represents the sampling point number of the input signal matrix in the distance dimension, Represents the column number of the input signal matrix, Indicates the distance unit number; S33: Matched filtering; increasing the pulse width will reduce the resolution, while pulse compression achieves both through signal modulation. The matched filter compresses the received wide pulse into a narrow pulse, and its output is shown in formula (3): (3); in: is the output of the matched filter, is the matched filter input signal, is the complex conjugate of the transmitted signal, is the integral variable, which represents the time dimension of the signal. The integral range covers the entire signal duration. is the delay variable, which represents the time offset between the received signal and the transmitted signal; S34: Moving target display; the input matrix is filtered using a fourth-order Butterworth filter. The transfer function of the fourth-order Butterworth filter is shown in Equation (4): (4); in: is the fourth-order Butterworth filter transfer function, and is the coefficient, is the delay operator; S35: Velocity-dimensional fast Fourier transform; perform a 256-point fast Fourier transform on each row of distance cells, and the calculation method is shown in formula (5): (5); in: is the result matrix of fast Fourier transform, is the input signal matrix, , Indicates the distance unit number, Indicates the pulse number, is the summation variable.

[0015] As a further limitation of this technical solution, S5 includes: S51: Spatiotemporal attention module based on global pixel features; performs temporal modeling on the input feature map to fully exploit the temporal correlation features between multi-frame range-Doppler feature maps; The dimension of the input feature map is ,in is the number of channels, is the time dimension, and are the height and width of the space respectively; S52: Spatial downsampling: The input feature map is processed by spatial downsampling to compress the spatial dimension of the feature map and reduce the computational complexity. The convolution kernel size of the three-dimensional convolution layer is , the stride size is 2, the filling method is the three-dimensional convolution module, and the calculation method of the three-dimensional convolution module is shown in formula (6): (6); in: represents the output feature map of the three-dimensional convolution, represents the input feature map of the 3D convolution, represents the three-dimensional convolution function; S53: Self-attention mechanism; the self-attention mechanism allows the model to dynamically focus on the importance of different parts when processing the input sequence; The dimension is The input sequence passes through three linear transformation layers to generate the query matrix , key matrix Sum Matrix , the calculation method is shown in formula (7): (7); in: Represents the output feature map of 3D convolution; The attention score matrix can be obtained by dot product operation , the calculation method is shown in formula (8): (8); in: represents the characteristic dimension of the bond matrix, is the scaling factor, softmax () is the Softmax function; Finally, the context vector is obtained through weighted summation operation , the calculation method is shown in formula (9): (9); S54: Global mean filtering; first calculate the interaction between any two positions to directly capture long-range dependencies, rather than being limited to adjacent points, the correlation function The calculation method is shown in formula (10): (10); in: represents the inner product, and are the position matrices respectively; S55: Feature weighting and residual connection; through attention score matrix A Pair Matrix V Weighted, aggregated global context information, residual connection to the output feature map of the three-dimensional convolution Directly superimpose on the global features to avoid information loss; The calculation method of the residual connection result is shown in formula (11): (11).

[0016] As a further limitation of this technical solution, S6 includes: S61: Spatiotemporal feature separation module; decomposes the traditional 3D convolution kernel into a combination of 2D convolution in spatial domain and 1D convolution in temporal domain; S62: Channel compression: First, channel compression is performed on the input feature map to reduce computational complexity by pruning redundant channels. S63: Spatial feature extraction: Two parallel convolution branches are used to extract spatial features of different dimensions. Branch 1 uses a 1×3×3 convolution kernel to focus on extracting the spatial relationship in the horizontal-vertical direction. Branch 2 uses a 1×1×1 convolution kernel to achieve cross-channel interaction. The calculation method of the two branches is shown in formula (12). The calculation method of the feature fusion result is shown in formula (13): (12); (13); in: Extract local spatial features, Pay attention to channel interaction, is the spatial convolution, It is a 1×1×1 convolution; S64: Temporal feature extraction; a 3×1×1 convolution kernel covers 3 consecutive frames, and a weighted sum is calculated to capture short-term motion patterns. It slides only on the time axis, and the spatial dimension remains independently processed.

[0017] Compared with related technologies, the millimeter-wave radar human behavior recognition system and method based on global pixel features provided by the present invention has the following beneficial effects: The technical solution provided by this invention includes a millimeter-wave radar raw data acquisition module, a millimeter-wave radar raw data transmission module, a millimeter-wave radar raw data preprocessing module, a TEN-GPF range-Doppler feature extraction and classification module, and a human behavior information application module. The TEN-GPF range-Doppler feature extraction and classification module includes a spatial downsampling unit, a spatiotemporal attention module based on global pixel features, and an STFSM spatiotemporal feature separation module. This system better utilizes the rich temporal and spatial information in multi-frame range-Doppler feature maps, improving the recognition accuracy of easily confusing behaviors and enhancing the recognition performance of existing models. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a structural schematic diagram of the present invention.

[0019] Figure 2 Schematic diagram of the structure of TEN-GPF of the present invention.

[0020] Figure 3 Schematic diagram of the structure of the spatiotemporal attention module based on global pixel features of the present invention.

[0021] Figure 4 It is a structural diagram of the STFSM spatiotemporal feature separation module of the present invention. DETAILED DESCRIPTION

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0023] A millimeter-wave radar human behavior recognition system based on global pixel features, comprising: The millimeter-wave radar raw data acquisition module, millimeter-wave radar raw data transmission module, millimeter-wave radar raw data preprocessing module, Time Extraction Network based on Global Pixel Feature (TEN-GPF) range-Doppler feature extraction and classification module, and human behavior information application module are connected in sequence; The millimeter wave radar raw data acquisition module includes a millimeter wave radar transmitting unit, a receiving unit, a signal mixing unit and a power amplification unit; In this embodiment of the present invention, the millimeter-wave radar raw data acquisition module uses frequency-modulated continuous wave (FMCW) technology to achieve high-precision detection. The transmitting unit is based on the Texas Instruments IWR1843 Boost radar system, operating in the 77-80 GHz frequency band with a bandwidth of 2998.2 MHz. Signals are transmitted and received via an antenna array consisting of one transmitting antenna and one receiving antenna (with a gain of 40 dB). The system sampling rate is 10,000 kps, with each frame containing 128 chirps, and the frame period is set to 500 ms.

[0024] In the embodiment of the present invention, after the receiving unit collects the target reflected signal, the signal mixing unit performs key processing: mixing the received signal with the original transmitted signal and its orthogonal signal respectively to generate an orthogonal intermediate frequency signal containing target distance and speed information.

[0025] In this embodiment, the power amplifier unit utilizes a three-stage amplification structure: a pre-stage low-noise amplifier provides 20dB gain to compensate for transmission losses, an intermediate frequency amplifier with 30dB gain optimizes the signal-to-noise ratio, and a final programmable gain amplifier achieves 40dB of dynamic range adjustment. The system's detection range is set to a 5m x 4m rectangular area in front of the radar, with an installation height of 1.2m, ensuring complete coverage of human activity.

[0026] In the present invention, experimental data collection shows that the system can clearly identify six daily activities, such as bending, falling, and jogging, within a range of 1-5 meters. Through distance-dimensional FFT processing of 128 sampling points and a 2998.2MHz bandwidth, the theoretical distance resolution reaches 5cm.

[0027] The millimeter wave radar raw data preprocessing module includes a raw data storage unit, a raw data partitioning unit, a chirp matrix reshaping unit, a range-dimensional fast Fourier transform unit, a moving target display unit, a velocity-dimensional fast Fourier transform unit, and a range-Doppler characteristic map visualization unit; The raw data division unit divides the raw data according to a specific period; the chirp matrix reshaping unit reorganizes the radar raw signal into a two-dimensional matrix according to the chirp sequence; the distance-dimensional fast Fourier transform unit extracts the distance information of the signal through a fast Fourier transform operation; the speed-dimensional fast Fourier transform unit extracts the speed information of the signal through a fast Fourier transform operation; the range-Doppler feature map visualization unit generates a range-speed feature map containing target motion information; the matched filter unit uses pulse compression technology to improve the resolution of the radar signal, thereby achieving more accurate ranging and speed measurement; In this embodiment of the present invention, the data preprocessing module segments the raw radar signal, converting the continuous signal into discrete frames suitable for feature extraction. The raw ADC sampled data first enters the segmentation unit, where it is segmented into 128 chirp cycles, each containing 128 sampling points. The segmented data is then reassembled by the chirp matrix reshaping unit to form a two-dimensional matrix. This reassembly effectively preserves the time-frequency characteristics of the radar signal, establishing the data structure foundation for subsequent processing. Experiments have shown that using a Hamming window for windowing can effectively mitigate spectrum leakage. The fast Fourier transform unit performs spectrum analysis in the range and velocity dimensions to extract target range and velocity information. The range-dimensional Fourier transform converts the time-domain sampling sequence into the frequency domain. The peak frequency corresponds to the target's range. The different echo delays for targets at different distances result in different intermediate frequencies. The fast Fourier transform decomposes the mixed time-domain signal into discrete frequency bins, each corresponding to a specific range gate, enabling multi-target range resolution. The phase difference for static targets is zero (zero-frequency component), while moving targets produce a non-zero frequency shift. The range-Doppler characteristic diagram, generated by the velocity-dimensional Fast Fourier Transform, can intuitively distinguish stationary clutter from moving targets. The moving target display unit effectively suppresses static background interference through digital filtering technology.

[0028] The temporal extraction network based on global pixel features, the distance-Doppler feature extraction and classification module, includes a spatial downsampling unit, a spatio-temporal attention module based on global pixel features (STAM-GPF), a spatio-temporal feature separation module (STFSM), and a fully connected layer classifier.

[0029] In this embodiment of the present invention, a spatiotemporal attention module based on global pixel features performs channel-wise dimensionality reduction on input features through 1×1 convolution, projecting high-dimensional features into a low-dimensional latent space. This compression preserves critical spatial information while reducing subsequent computational burden. The feature matrix is reshaped into a two-dimensional form, laying the foundation for pixel-level correlation analysis. Simultaneously, a self-attention mechanism is employed to process features in parallel: the query matrix captures the importance of target features, the key matrix establishes cross-temporal and spatial correlations, and the value matrix preserves the original feature content. Attention weights are calculated using scaled dot products, and temporal position encoding is introduced to enable the model to perceive temporal order. Temporal sequence is modeled by calculating the similarity between each pixel, and finally, the original features are fused with the attention features through residual connections. Layer normalization stabilizes the training process, and the DropPath technique prevents overfitting. This design preserves underlying details while enhancing high-level semantic features. This mechanism significantly improves feature discrimination in complex scenarios through joint spatiotemporal modeling and dynamic feature selection.

[0030] In this embodiment of the present invention, the STFSM spatiotemporal feature separation module decouples the traditional three-dimensional convolution kernel (3×3×3) into two independent processing paths in the spatial domain and the temporal domain: Spatial processing path: A 1×3×3 two-dimensional spatial convolution kernel is used to focus on extracting spatial features within the image plane, capturing spatial relationships in the horizontal and vertical directions, and effectively extracting geometric features such as edges and textures. At the same time, a 1×1×1 convolution kernel is added to achieve cross-channel feature recombination and enhance important feature channels through linear combination.

[0031] Temporal processing path: A 3×1×1 one-dimensional temporal convolution kernel is used. The 3×1×1 convolution kernel slides along the time axis to model the dynamic changes between three consecutive frames. The spatial dimension is processed independently to avoid interference between spatiotemporal features and focus on analyzing the motion features of the temporal dimension.

[0032] Through convolution kernel decomposition, the number of parameters is greatly reduced, and the migration and use of two-dimensional convolution pre-trained weights are supported; in multi-scale feature fusion, the spatial dual branches capture local details and global context respectively, the temporal path establishes short-term motion dependencies, and the residual connection retains the original feature information.

[0033] The spatial downsampling unit compresses the feature map size through the convolutional layer.

[0034] The millimeter-wave radar raw data transmission module supports multiple transmission protocols such as Wi-Fi, serial communication, Ethernet, LoRa and ZigBee.

[0035] In this embodiment of the present invention, the data transmission module adopts a multi-protocol compatible design, which can select the optimal transmission method according to different application scenarios. In this embodiment of the present invention, the data transmission module adopts a multi-protocol compatible design, which can select the optimal transmission method according to different application scenarios. In local data transmission scenarios, the system uses the 1000Base-T Gigabit Ethernet protocol for high-speed transmission.

[0036] The spatiotemporal attention module based on global pixel features includes a spatial downsampling unit, a self-attention unit, a global mean filter unit and a residual link unit connected in sequence; the self-attention unit calculates the global correlation weight of the spatiotemporal features; the global mean filter unit smoothes the attention weight; the residual link unit retains the original features to avoid gradient disappearance; The spatiotemporal feature separation module includes a channel downsampling unit, a spatial feature extraction unit, a temporal feature extraction unit and a residual connection unit; the channel downsampling unit can reduce the channel dimension; the spatial feature extraction unit separates spatial dimension features through two-dimensional spatial convolution; the temporal feature extraction unit uses one-dimensional temporal convolution to capture dynamic changes in behavior; and the residual connection unit realizes cross-layer feature reuse.

[0037] The fully connected layer classifier of the time extraction network distance-Doppler feature extraction and classification module based on global pixel features maps features into specific behavior category probabilities through the normalized exponential function Softmax to identify different behaviors.

[0038] The human behavior information application module includes smart medical care, smart transportation and smart home scenarios.

[0039] At the application level, this system has broad application in multiple intelligent scenarios. In healthcare monitoring, it can detect falls by analyzing micro-motion signatures; in intelligent transportation, it can predict pedestrian behavior using motion trajectories from consecutive frames; and in smart homes, it can enable contactless control by recognizing specific gestures. The system utilizes an end-to-end deep learning architecture to directly map raw radar signals to behavioral categories, significantly simplifying traditional processing.

[0040] The innovations of this invention are primarily reflected in three aspects: First, the proposed spatiotemporal attention mechanism overcomes the limitations of traditional local receptive fields and achieves truly global feature modeling; second, the spatiotemporal separation feature extraction architecture effectively solves the problem of unclear mixed feature representation; and finally, the entire system uses millimeter-wave radar as a sensing method, achieving all-weather, high-precision behavior recognition while protecting privacy. These technological innovations significantly improve the system's recognition performance in complex scenarios.

[0041] A recognition method for a millimeter-wave radar human behavior recognition system based on global pixel features comprises the following steps: S1: Millimeter-wave radar raw data acquisition: The millimeter-wave radar in the millimeter-wave radar raw data acquisition module collects the raw data of the user's behavior. The format of the raw data is IQ data and is stored in the form of a binary file.

[0042] S2. Millimeter-wave radar raw data transmission: The raw data collected by the millimeter-wave radar raw data acquisition module is transmitted to the raw data storage unit of the millimeter-wave radar raw data preprocessing module via an Ethernet cable.

[0043] S3. Millimeter-wave radar raw data preprocessing: The raw data storage unit of the millimeter-wave radar raw data preprocessing module cuts the radar raw data into frames with a 25% overlap rate. Each frame contains 128 chirps. The chirp matrix reshaping unit reconstructs the chirp matrix into a two-dimensional matrix, where the two dimensions are chirp and frame respectively; The distance-dimensional fast Fourier transform unit performs fast Fourier transform on the chirp matrix along the fast time dimension to extract the distance information of the target. The result of the fast Fourier transform will show the intensity of different frequency components. Since the frequency is proportional to the target distance, this can be directly converted into distance information; The matched filter unit uses pulse compression technology to enhance the detection capability of useful signals and improve the resolution of radar signals, thereby achieving more accurate ranging and speed measurement; The moving target display unit uses a fourth-order Butterworth filter to separate the characteristics of high-frequency dynamic targets from the raw data collected by the millimeter-wave radar, while suppressing static or slowly changing background clutter; The speed-dimensional fast Fourier transform unit performs fast Fourier transform along the slow time dimension to extract the speed information of the target. The speed of the target can be calculated by analyzing the frequency offset in the fast Fourier transform result.

[0044] The S3 includes: S31: Data framing and chirp matrix reconstruction; Assume that the original signal sampling rate is , each frame duration , the inter-frame overlap rate is 25%, so the frame division formula is shown in formula (1): (1); in: , is the number of sampling points in a single frame, is the number of frames; Will M The frame data is reconstructed into a two-dimensional matrix with the dimension of , each column is a sampling sequence of a single chirp; S32: distance-dimensional fast Fourier transform; perform a 128-point fast Fourier transform on each column of chirp data, and the calculation method is shown in formula (2): (2); in: is the result matrix of fast Fourier transform, is the input signal matrix, is the rotation factor, , Represents the sampling point number of the input signal matrix in the distance dimension, Represents the column number of the input signal matrix, Indicates the distance unit number; S33: Matched filtering; increasing the pulse width will reduce the resolution, while pulse compression achieves both through signal modulation. The matched filter compresses the received wide pulse into a narrow pulse, and its output is shown in formula (3): (3); in: is the output of the matched filter, is the matched filter input signal, is the complex conjugate of the transmitted signal, is the integral variable, which represents the time dimension of the signal. The integral range covers the entire signal duration. is the delay variable, which represents the time offset between the received signal and the transmitted signal; S34: Moving target display; the input matrix is filtered using a fourth-order Butterworth filter. The transfer function of the fourth-order Butterworth filter is shown in Equation (4): (4); in: is the fourth-order Butterworth filter transfer function, and is the coefficient, is the delay operator; S35: Velocity-dimensional fast Fourier transform; perform a 256-point fast Fourier transform on each row of distance cells, and the calculation method is shown in formula (5): (5); in: is the result matrix of fast Fourier transform, is the input signal matrix, , Indicates the distance unit number, Indicates the pulse number, is the summation variable.

[0045] S4: Construct a time extraction network distance-Doppler feature extraction and classification module based on global pixel features to perform human behavior recognition and classification; input the 6 frames of distance-Doppler feature maps that have completed feature preprocessing into the time extraction network distance-Doppler feature extraction and classification module based on global pixel features in batches, and then use the spatiotemporal attention module based on global pixel features to perform temporal modeling, and the spatiotemporal feature separation module to display the extraction of separated temporal features and spatial features, and finally the fully connected layer classifier outputs the discrimination result.

[0046] S5: Construct a spatiotemporal attention module based on global pixel features to perform temporal modeling on the input feature map and fully exploit the temporal correlation features between multi-frame range-Doppler feature maps; The input feature map is used to reduce the dimensionality of spatial features in the spatial downsampling unit, thereby extracting features and reducing the amount of subsequent calculations. The self-attention unit calculates the spatiotemporal dependencies between global pixels and establishes cross-frame feature associations through the Query-Key-Value mechanism; The global mean filter unit calculates the correlation of global pixels through matrix multiplication, which can model both time series and spatial information. The residual connection unit adds the original input and the global features output by the global mean filter unit to achieve collaborative modeling of local details and global context.

[0047] The S5 includes: S51: Spatiotemporal attention module based on global pixel features; performs temporal modeling on the input feature map to fully exploit the temporal correlation features between multi-frame range-Doppler feature maps; for a pixel in the image, its return value is obtained by the weighted sum of all other pixels, which means that each pixel is associated with all other pixels in the image, thus achieving global attention; The dimension of the input feature map is ,in is the number of channels, is the time dimension, and are the height and width of the space respectively; S52: Spatial downsampling: The input feature map is processed by spatial downsampling to compress the spatial dimension of the feature map and reduce the computational complexity. The convolution kernel size of the three-dimensional convolution layer is , the stride size is 2, the filling method is the three-dimensional convolution module (SAME), and the calculation method of the three-dimensional convolution module is shown in formula (6): (6); in: represents the output feature map of the three-dimensional convolution, represents the input feature map of the 3D convolution, represents the three-dimensional convolution function; Original feature map H × W The computational complexity of attention is , after downsampling to , which can greatly reduce the amount of subsequent calculations; S53: Self-attention mechanism; the self-attention mechanism allows the model to dynamically focus on the importance of different parts when processing the input sequence; compared to recurrent neural networks, the self-attention mechanism can parallelize the output of each time step, significantly improving the training speed and efficiency of the model; The dimension is The input sequence passes through three linear transformation layers to generate the query matrix , key matrix Sum Matrix , the calculation method is shown in formula (7): (7); in: Represents the output feature map of 3D convolution; The attention score matrix can be obtained by dot product operation , the calculation method is shown in formula (8): (8); in: represents the characteristic dimension of the bond matrix, is a scaling factor that can prevent the dot product result from being too large and causing gradient instability. softmax () is the Softmax function; Finally, the context vector is obtained through weighted summation operation , the calculation method is shown in formula (9): (9); In radar-based human action recognition, the self-attention mechanism calculates the correlation between different positions in the feature map, captures long-range dependencies, enhances attention to key areas, and improves classification accuracy. S54: Global mean filtering; first calculate the interaction between any two positions to directly capture long-range dependencies, rather than being limited to adjacent points. This is equivalent to constructing a convolution kernel as large as the feature map size, which can maintain more information, the correlation function The calculation method is shown in formula (10): (10); in: represents the inner product, and are position matrices respectively; measure the similarity between two positions; the exponential function converts The inner product result is mapped to the positive real number space to ensure the correlation is non-negative. Similar to the Gaussian kernel function, it makes the similarity calculation have a smooth attenuation characteristic and can better suppress noise. Traditional convolution can only capture local receptive fields (such as 3×3 neighborhoods), while formula (10) calculates the correlation between any two points globally; S55: Feature weighting and residual connection; through attention score matrix A Pair Matrix V Weighted, aggregated global context information, residual connection to the output feature map of the three-dimensional convolution Directly superimpose on the global features to avoid information loss; The calculation method of the residual connection result is shown in formula (11): (11).

[0048] S6: Construct a spatiotemporal feature separation module to extract temporal and spatial features from the input feature map, which can reduce the impact of feature confounding. The input feature map completes the dimensionality reduction of the number of channels in the channel downsampling unit, significantly reducing the number of parameters and computational complexity of subsequent layers. By reducing redundant feature channels, key information is retained, improving the generalization ability of the model; The spatial feature extraction unit uses two-dimensional spatial convolution with convolution kernel sizes of 1×3×3 and 1×1×1 to extract spatial features; The temporal feature extraction unit uses a one-dimensional temporal convolution with a convolution kernel size of 3×1×1 to extract temporal features; The residual connection unit extracts high-level semantic features and is more conducive to capturing global context; The S6 includes: S61: Spatiotemporal feature separation module; the traditional three-dimensional convolution kernel The convolution kernel is decomposed into a combination of two-dimensional convolution in the spatial domain and one-dimensional convolution in the temporal domain to reduce the amount of computation and training parameters of the model. This not only reduces the amount of computation and parameters, but also fully utilizes the advantages of pre-training on the two-dimensional convolutional network. S62: Channel compression: First, channel compression is performed on the input feature map to reduce computational complexity by pruning redundant channels. S63: Spatial feature extraction: Two parallel convolution branches are used to extract spatial features of different dimensions. Branch 1 uses a 1×3×3 convolution kernel to focus on extracting the spatial relationship in the horizontal-vertical direction, and can focus on capturing geometric features such as edges and textures in local areas of the image. Branch 2 uses a 1×1×1 convolution kernel to achieve cross-channel interaction, linearly combine channel dimensions, and enhance or suppress the feature response of specific channels. The calculation method of the two branches is shown in formula (12), and the calculation method of the feature fusion result is shown in formula (13): (12); (13); in: Extract local spatial features, Pay attention to channel interaction, is the spatial convolution, It is a 1×1×1 convolution; S64: Temporal feature extraction; a 3×1×1 convolution kernel covers three consecutive frames and calculates a weighted sum to capture short-term motion patterns. It slides only on the time axis, while the spatial dimensions remain independently processed, thus reducing the impact of feature confounding to a certain extent.

[0049] S7: Output action recognition results. After multi-dimensional spatiotemporal feature extraction and fusion, the obtained feature data is input into the fully connected layer. The number of hidden units in the fully connected layer is the number of behavior categories contained in the data. The data after the fully connected layer is passed through the Softmax classifier to calculate the probability of the corresponding behavior. The behavior type with the highest probability is the final judgment result. S8: Display and application of human behavior information. Through the human behavior information application module, the behavior recognition results are visualized and used in smart medical care, smart transportation, smart home and other fields.

[0050] The present invention realizes the classification and recognition of human behavior based on millimeter-wave radar by collecting and analyzing radar data, which makes up for the shortcomings of video-based behavior recognition such as high cost, susceptibility to interference, and poor privacy. The attention mechanism and residual structure effectively improve the effectiveness and comprehensiveness of feature extraction and the ability to capture complex dynamics. Compared with mainstream and latest models, it has obvious advantages in adaptability, reliability and practicality.

[0051] The technical solution provided by this invention includes a millimeter-wave radar raw data acquisition module, a millimeter-wave radar raw data transmission module, a millimeter-wave radar raw data preprocessing module, a TEN-GPF range-Doppler feature extraction and classification module, and a human behavior information application module. The spatiotemporal attention module based on global pixel features includes a sequentially connected spatial downsampling unit, a self-attention unit, a global mean filter unit, and a residual link unit. The spatiotemporal feature separation module includes a channel downsampling unit, a spatial feature extraction unit, a temporal feature extraction unit, and a residual link unit. This system better utilizes the temporal features contained in human behavior data, improves the recognition accuracy of easily confused behaviors, and enhances the recognition performance of existing models.

[0052] Each step of the embodiment of the present invention may be performed by an electronic device, including but not limited to a mobile phone, a tablet computer, a portable PC, a desktop computer, etc.

[0053] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0054] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A millimeter wave radar human behavior recognition system based on global pixel features, characterized by: include: The millimeter-wave radar raw data acquisition module, the millimeter-wave radar raw data transmission module, the millimeter-wave radar raw data preprocessing module, the time extraction network distance-Doppler feature extraction and classification module based on global pixel features, and the human behavior information application module are connected in sequence; The millimeter wave radar raw data acquisition module includes a millimeter wave radar transmitting unit, a receiving unit, a signal mixing unit and a power amplification unit; The millimeter wave radar raw data preprocessing module includes a raw data storage unit, a raw data partitioning unit, a chirp matrix reshaping unit, a range-dimensional fast Fourier transform unit, a moving target display unit, a velocity-dimensional fast Fourier transform unit, and a range-Doppler characteristic map visualization unit; The time extraction network distance-Doppler feature extraction and classification module based on global pixel features includes a spatial downsampling unit, a spatiotemporal attention module based on global pixel features, a spatiotemporal feature separation module and a fully connected layer classifier.

2. The millimeter-wave radar human behavior recognition system based on global pixel features according to claim 1 is characterized by: The millimeter-wave radar raw data transmission module supports multiple transmission protocols such as Wi-Fi, serial communication, Ethernet, LoRa and ZigBee.

3. The millimeter-wave radar human behavior recognition system based on global pixel features according to claim 1 is characterized by: The spatiotemporal attention module based on global pixel features includes a spatial downsampling unit, a self-attention unit, a global mean filter unit and a residual link unit connected in sequence; The spatiotemporal feature separation module includes a channel downsampling unit, a spatial feature extraction unit, a temporal feature extraction unit and a residual connection unit.

4. The millimeter-wave radar human behavior recognition system based on global pixel features according to claim 3 is characterized by: The fully connected layer classifier of the time extraction network distance-Doppler feature extraction and classification module based on global pixel features maps features into specific behavior category probabilities through the normalized exponential function Softmax to identify different behaviors.

5. The millimeter-wave radar human behavior recognition system based on global pixel features according to claim 1 is characterized by: The human behavior information application module includes smart medical care, smart transportation and smart home scenarios.

6. A recognition method using the millimeter wave radar human behavior recognition system based on global pixel features according to claim 4, characterized in that: The following steps are involved: S1: Millimeter-wave radar raw data acquisition: The millimeter-wave radar in the millimeter-wave radar raw data acquisition module collects raw data of user behavior. The raw data is in the form of IQ data and is stored in the form of a binary file; S2, millimeter wave radar raw data transmission, transmitting the raw data collected by the millimeter wave radar raw data acquisition module to the raw data storage unit of the millimeter wave radar raw data preprocessing module via an Ethernet cable; S3. Millimeter-wave radar raw data preprocessing: The raw data storage unit of the millimeter-wave radar raw data preprocessing module cuts the radar raw data into frames with a 25% overlap rate. Each frame contains 128 chirps. S4: Construct a time extraction network distance-Doppler feature extraction and classification module based on global pixel features to perform human behavior recognition and classification; the 6 frames of distance-Doppler feature maps that have completed feature preprocessing are input into the time extraction network distance-Doppler feature extraction and classification module based on global pixel features in batches, and then the spatiotemporal attention module based on global pixel features is used to perform temporal modeling, and the spatiotemporal feature separation module is used to extract and separate temporal and spatial features. Finally, the fully connected layer classifier outputs the discrimination result; S5: Construct a spatiotemporal attention module based on global pixel features to perform temporal modeling on the input feature map and fully exploit the temporal correlation features between multi-frame range-Doppler feature maps; S6: Construct a spatiotemporal feature separation module to extract temporal and spatial features from the input feature map, which can reduce the impact of feature confounding. S7: Output action recognition results. After multi-dimensional spatiotemporal feature extraction and fusion, the obtained feature data is input into the fully connected layer. The number of hidden units in the fully connected layer is the number of behavior categories contained in the data. The data after the fully connected layer is passed through the Softmax classifier to calculate the probability of the corresponding behavior. The behavior type with the highest probability is the final judgment result. S8: Display and application of human behavior information. Through the human behavior information application module, the behavior recognition results are visualized and used in the fields of smart medical care, smart transportation and smart home.

7. The identification method according to claim 6, wherein: The S3 includes: S31: Data framing and chirp matrix reconstruction; Assume that the original signal sampling rate is , each frame duration , the inter-frame overlap rate is 25%, so the frame division formula is shown in formula (1): (1); in: , is the number of sampling points in a single frame, is the number of frames; Will M The frame data is reconstructed into a two-dimensional matrix with the dimension of , each column is a sampling sequence of a single chirp; S32: distance-dimensional fast Fourier transform; perform a 128-point fast Fourier transform on each column of chirp data, and the calculation method is shown in formula (2): (2); in: is the result matrix of fast Fourier transform, is the input signal matrix, is the rotation factor, , Represents the sampling point number of the input signal matrix in the distance dimension, Represents the column number of the input signal matrix, Indicates the distance unit number; S33: Matched filtering; increasing the pulse width will reduce the resolution, while pulse compression achieves both through signal modulation. The matched filter compresses the received wide pulse into a narrow pulse, and its output is shown in formula (3): (3); in: is the output of the matched filter, is the matched filter input signal, is the complex conjugate of the transmitted signal, is the integral variable, which represents the time dimension of the signal. The integral range covers the entire signal duration. is the delay variable, which represents the time offset between the received signal and the transmitted signal; S34: Moving target display; the input matrix is filtered using a fourth-order Butterworth filter. The transfer function of the fourth-order Butterworth filter is shown in Equation (4): (4); in: is the fourth-order Butterworth filter transfer function, and is the coefficient, is the delay operator; S35: Velocity-dimensional fast Fourier transform; perform a 256-point fast Fourier transform on each row of distance cells, and the calculation method is shown in formula (5): (5); in: is the result matrix of fast Fourier transform, is the input signal matrix, , Indicates the distance unit number, Indicates the pulse number, is the summation variable.

8. The identification method according to claim 7, wherein: The S5 includes: S51: Spatiotemporal attention module based on global pixel features; performs temporal modeling on the input feature map to fully exploit the temporal correlation features between multi-frame range-Doppler feature maps; The dimension of the input feature map is ,in is the number of channels, is the time dimension, and are the height and width of the space respectively; S52: Spatial downsampling: The input feature map is processed by spatial downsampling to compress the spatial dimension of the feature map and reduce the computational complexity. The convolution kernel size of the three-dimensional convolution layer is , the stride size is 2, the filling method is the three-dimensional convolution module, and the calculation method of the three-dimensional convolution module is shown in formula (6): (6); in: represents the output feature map of the three-dimensional convolution, represents the input feature map of the 3D convolution, represents the three-dimensional convolution function; S53: Self-attention mechanism; the self-attention mechanism allows the model to dynamically focus on the importance of different parts when processing the input sequence; The dimension is The input sequence passes through three linear transformation layers to generate the query matrix , key matrix Sum Matrix , the calculation method is shown in formula (7): (7); in: Represents the output feature map of 3D convolution; The attention score matrix can be obtained by dot product operation , the calculation method is shown in formula (8): (8); in: represents the characteristic dimension of the bond matrix, is the scaling factor, softmax () is the Softmax function; Finally, the context vector is obtained through weighted summation operation , the calculation method is shown in formula (9): (9); S54: Global mean filtering; first calculate the interaction between any two positions to directly capture long-range dependencies, rather than being limited to adjacent points, the correlation function The calculation method is shown in formula (10): (10); in: represents the inner product, and are the position matrices respectively; S55: Feature weighting and residual connection; through attention score matrix A Pair Matrix V Weighted, aggregated global context information, residual connection to the output feature map of the three-dimensional convolution Directly superimpose on the global features to avoid information loss; The calculation method of the residual connection result is shown in formula (11): (11)。 9. The identification method according to claim 8, wherein: The S6 includes: S61: Spatiotemporal feature separation module; decomposes the traditional 3D convolution kernel into a combination of 2D convolution in spatial domain and 1D convolution in temporal domain; S62: Channel compression: First, channel compression is performed on the input feature map to reduce computational complexity by pruning redundant channels. S63: Spatial feature extraction: Two parallel convolution branches are used to extract spatial features of different dimensions. Branch 1 uses a 1×3×3 convolution kernel to focus on extracting the spatial relationship in the horizontal-vertical direction. Branch 2 uses a 1×1×1 convolution kernel to achieve cross-channel interaction. The calculation method of the two branches is shown in formula (12). The calculation method of the feature fusion result is shown in formula (13): (12); (13); in: Extract local spatial features, Pay attention to channel interaction, is the spatial convolution, It is a 1×1×1 convolution; S64: Temporal feature extraction; a 3×1×1 convolution kernel covers 3 consecutive frames, and a weighted sum is calculated to capture short-term motion patterns. It slides only on the time axis, and the spatial dimension remains independently processed.

Citation Information

Patent Citations

  • Millimeter wave radar head action recognition method based on multi-domain fusion deep learning

    CN115063884A

  • Vital sign detection method based on millimeter wave radar

    CN115644840A

  • Denoising and classifying method for Doppler feature map of human body action

    CN115761353A

  • Human behavior identification method based on FMCW millimeter wave radar

    CN118191827A