Human body action recognition method based on millimeter wave radar enhanced point cloud

By employing static clutter elimination, dynamic threshold filtering, and an improved DBSCAN clustering combined with temporal frame sequence voxelization and dual-view projection feature reconstruction method, the problems of poor point cloud quality and insufficient robustness in human motion recognition by millimeter-wave radar are solved, achieving high-accuracy motion recognition.

CN121482858APending Publication Date: 2026-02-06ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511488647.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing millimeter-wave radar human motion recognition methods suffer from poor point cloud quality and insufficient recognition robustness, especially limiting the accuracy of motion recognition in complex environments.

Method used

A three-level point cloud enhancement strategy, consisting of static clutter cancellation, dynamic threshold filtering, and improved DBSCAN clustering, is adopted. Combined with temporal frame sequence voxelization and dual-view projection feature reconstruction methods, a high-accuracy action classification is achieved through an improved Transformer network.

Benefits of technology

It significantly improves the purity and integrity of point clouds, enhances the accuracy and robustness of action recognition, and is suitable for applications in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482858A_ABST
    Figure CN121482858A_ABST
Patent Text Reader

Abstract

The invention discloses a millimeter wave radar enhanced point cloud-based human body action recognition method, which comprises the following steps of: firstly, processing a radar echo signal through frame difference processing, 2D-FFT and a secondary density clustering algorithm, and enhancing the quality of a three-dimensional point cloud; then, organizing continuous seven frames of point clouds into a time sequence sample at a sliding step length of 2, independently performing voxelization processing on each frame of point clouds, performing maximum intensity projection frame by frame along the Y axis and the X axis, and generating a double-view projection image sequence of XOZ and YOZ planes; and finally, inputting the double-view-angle image sequence into an improved DVP-TransNet network which adopts double-branch ResNet50 to perform feature extraction, adaptively fusing double-view-angle features through a trainable weight module, and capturing time sequence dynamic features by using a Transformer encoder to finally realize high-precision classification of human body actions. According to the method, the action recognition accuracy rate up to 99.67% is achieved in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of millimeter wave radar signal processing, point cloud enhancement and action recognition, and particularly relates to a human action recognition method based on millimeter wave radar enhanced point cloud. The method is aimed at the core problem of sparse millimeter wave radar point cloud and low signal-to-noise ratio. Through the fusion design of point cloud enhancement strategy and deep learning, the human action recognition accuracy can be significantly improved in low signal-to-noise ratio and complex environment. It is suitable for security monitoring (such as fall detection for the elderly living alone), smart home (such as gesture control of household appliances), health monitoring (such as daily action behavior analysis) and other scenes. BACKGROUND

[0002] With the rapid development of artificial intelligence, Internet of Things and wearable devices, human action recognition as a core technology of human-computer interaction and security monitoring has wide application prospects in the fields of security, medical treatment and human-computer interaction. Traditional action recognition technology mainly relies on optical cameras or wearable sensors, but has the following problems:

[0003] Light and shielding influence: optical cameras have poor recognition effect in weak light or shielding environment.

[0004] Privacy leakage risk: image acquisition may infringe on user privacy.

[0005] High dependence on equipment: wearable sensors need to be worn by users, which is inconvenient to use.

[0006] Millimeter wave radar has the advantages of all-weather, anti-light change, penetration of light and thin shielding objects, no privacy leakage, etc. In recent years, it has been gradually applied to human action recognition. However, the existing millimeter wave radar action recognition method has the following shortcomings: sparse point cloud, uneven spatial distribution; low signal-to-noise ratio, many noise points; insufficient robustness in complex environment, limited action recognition accuracy, etc.

[0007] Therefore, how to enhance the quality of point cloud and combine an efficient neural network structure to realize robust action recognition has become a technical problem that needs to be solved in the field of millimeter wave radar action recognition. SUMMARY

[0008] In order to overcome the defects of poor point cloud quality and insufficient recognition robustness in the existing millimeter wave radar human action recognition method, the present application provides a human action recognition method based on millimeter wave radar enhanced point cloud. Through the three-level point cloud enhancement strategy of "static clutter elimination-dynamic threshold screening-improved DBSCAN clustering", the purity and integrity of point cloud are improved; combined with the feature reconstruction method of "time sequence frame sequence voxelization-double view projection", the problem of sparse point cloud is solved; finally, the improved Transformer network (DVP-TransNet) is used to realize high-accuracy action classification, meeting the application requirements in complex scenes.

[0009] To solve the above technical problems the present application proposes the following technical solutions:

[0010] The human action recognition method based on millimeter wave radar enhanced point cloud provided by the present application comprises the following steps:

[0011] Step 1: build a millimeter wave FMCW radar experimental platform, set the radar parameters, and collect the human action echo signal;

[0012] Step 2: the human completes the preset action, the millimeter wave radar transmits the FMCW signal and receives the echo, obtains the intermediate frequency signal through mixing, and stores it as a.bin file in the form of complex numbers;

[0013] Step 3: first, eliminate static clutter through frame difference processing to enhance the signal features of the moving target, and then perform two-dimensional fast Fourier transform (2D-FFT) on each intermediate frequency signal to obtain a range-doppler map (RDM, Range-Doppler Map);

[0014] Step 4: perform two-dimensional ordered statistical constant false alarm rate detection (2D_OS_CFAR) on the generated RDM to obtain the initial candidate point cloud of the human action, perform dynamic threshold screening on the candidate points, use the improved DBSCAN clustering algorithm (hereinafter referred to as the two-stage density clustering algorithm) to eliminate isolated noise points, enhance the point cloud cluster structure, improve the point cloud quality, and obtain the effective target points;

[0015] Step 5: for the extracted effective target points, calculate the radial distance R and radial velocity Z according to the radar configuration, then obtain the two-dimensional plane coordinates X and Y according to the target azimuth angle, and combine the radial velocity Z to obtain the target three-dimensional point cloud (X, Y, Z);

[0016] Step 6: organize the continuous 7 frames of three-dimensional point clouds into time sequence samples in a sliding step of 2, independently perform voxelization processing on each frame of point cloud in the sample, divide it into a three-dimensional voxel grid of 10(X)×32(Y)×32(Z), and count the number of points in each non-empty voxel to generate a discrete voxel sequence; then, for each frame of tensor in the voxel sequence, perform maximum intensity projection along the Y axis and the X axis respectively to obtain the dual-view projection image sequence of the XOZ plane and the YOZ plane;

[0017] Step 7: input the dual-view projection image into the improved deep learning network DVP-TransNet (Dual-View Projection Transformer Network), which is composed of a double-branch feature extraction module of ResNet50, a self-learning weight feature fusion module, and a Transformer encoder module, to extract the time sequence features and perform action classification recognition.

[0018] Further, in step 1, the millimeter wave radar experimental platform refers to IWR1843 millimeter wave FMCW radar developed by Texas Instruments, the working frequency range of which is from 76GHz to 81GHz, which is configured as 1 transmitting antenna and 4 receiving antennas, the starting frequency and Chirp slope are set as 77GHz and 40MHz / μs respectively, and the ADC sampling rate is 2MHz. The number of Chirps per frame is 128, the number of ADC samples per Chirp is 128, and 50 frames are collected per action.

[0019] Further, in step 2, the human body performs nine actions in front of the radar, including walking, sitting, standing, squatting, falling, bending, bending to straightening, waving and stretching, each action lasting about 1-2s, and the intermediate frequency signal containing the echo information of the human body action received by the ADC1000EVM acquisition board is stored as a.bin file in the form of a complex number.

[0020] Further, the process of step 3 is:

[0021] (a) According to the set radar parameters, the original intermediate frequency signal data collected by the acquisition card is divided into multiple channel data and saved as complex I / Q data, that is, the intermediate frequency echo.bin file is obtained, then the data of each receiving antenna is read from the file and arranged into a three-dimensional array [n_ADC, n_Chirp, n_Frame], wherein n_Frame represents the number of frames sampled, n_Chirp represents the number of Chirps contained in each frame, and n_ADC represents the number of sampling points contained in each Chirp, then frame difference processing is performed, that is, the data of the latter frame is subtracted from the data of the former frame, so as to eliminate background clutter and improve signal-to-noise ratio, and output the frame difference two-dimensional I / Q data;

[0022] (b) Based on the two-dimensional I / Q data after frame difference processing, two-dimensional fast Fourier transform (2D FFT) is performed on the frame difference matrix corresponding to each receiving antenna, wherein distance FFT is performed along the n_ADC dimension and Doppler FFT is performed along the n_Chirp dimension, so as to obtain the RDM of each frame.

[0023] In the present application, the process of step 4 is:

[0024] (a) 2D_OS_CFAR detection algorithm is applied to RDM for target detection processing, a sliding window and a protection unit are set in the distance dimension and the Doppler dimension at the same time, the background noise estimate value is extracted in the two-dimensional region, and the signal strength of each unit is judged according to the preset false alarm probability threshold, so as to identify the potential human target point as the initial candidate point for subsequent point cloud construction and clustering recognition;

[0025] (b) After the initial detection is completed, first, the background energy is extracted in the region not containing the initial candidate points, the energy median is calculated as the background noise reference value, then the average of the energy median of the initial candidate points and the background noise reference value is calculated as the dynamic threshold value for subsequent discrimination;

[0026] (c) Based on the energy distribution characteristics, the initial candidate points are subjected to secondary screening, if the target point has a density higher than a set value in the current frame RDM, the background noise reference value is used as the screening threshold to remove the candidate points with energy lower than the threshold, otherwise the dynamic threshold value in (b) is used as the threshold to perform energy re-screening on the initial candidate points, to obtain a target point set (RD_target_new) with higher confidence; The purpose of setting the dynamic threshold is to balance the false alarm caused by too small conventional CFAR detection threshold and the missed alarm caused by too large threshold, to improve the purity and recognition robustness of point cloud extraction;

[0027] (d) Improved DBSCAN clustering (i.e. the second level processing of the two-level density clustering algorithm) is performed on RD_target_new: the distance dimension and Doppler dimension coordinates of each target point are extracted from RD_target_new, based on the spatial distribution characteristics of the millimeter wave radar point cloud, dynamic clustering parameters are set: neighborhood radius epsilon, minimum point number MinPts, each target point is traversed, if the number of target points contained in the epsilon neighborhood of the point is greater than or equal to MinPts, the point is marked as a core point; if the epsilon neighborhood of the point does not contain a core point, but the distance from the core point is less than or equal to epsilon, the point is marked as a boundary point; if neither a core point nor a boundary point, it is determined as an isolated noise point and removed; finally, the core points and boundary points are merged into point cloud clusters according to connectivity to construct a clustering label graph, to obtain effective target points; this level of processing can further eliminate the isolated noise points remaining after the first level of screening, optimize the spatial continuity of the point cloud cluster, and ensure that the effective target point cloud matches the actual shape of the human body action.

[0028] In the present application, in step 5, the distance dimension and Doppler dimension indexes of the target points can be obtained according to the clustering label graph, the radial distance R and radial velocity Z of the target points relative to the radar can be calculated according to the related formula and radar parameters; the azimuth angle of the target point can be obtained by using the MUSIC algorithm; the coordinates X and Y of the target point in the two-dimensional plane rectangular coordinate system can be calculated, and the radial velocity Z together constitute the three-dimensional point cloud features (X, Y, Z) of the target point; the point cloud data of the target point is stored in xlsx format according to the action category, the receiving antenna number, the frame number and the point number, and the data of each action category is divided into training set and test set in the order of acquisition time in the ratio of 8:2 for subsequent use.

[0029] In the present application, in step 6, the continuous 7 frames of three-dimensional point cloud data are constructed into a time sequence sample (sliding step is 2), and the voxelization processing is independently carried out on each frame of point cloud in the sequence, and the point cloud is quantized into a three-dimensional voxel grid of 10*32*32 (X*Y*Z), the number of points in each non-empty voxel is counted, and a discrete three-dimensional voxel sequence is generated; the maximum intensity projection is carried out on the voxel sequence frame by frame: the XOZ view image is obtained by projecting along the Y axis, and the YOZ view image is obtained by projecting along the X axis, so as to generate a double-view projection image sequence.

[0030] In the present application, the process of step 7 is:

[0031] (a) Training data preparation: input the xlsx format point cloud data divided according to the proportion of 8:2 into the "three-dimensional point cloud accumulation fusion-voxelization-double view projection" process, generate the XOZ plane projection image sequence and the YOZ plane projection image sequence corresponding to the training set point cloud one by one, and jointly constitute the training data set of DVP-TransNet;

[0032] (b) Double-branch feature extraction: two independent ResNet50 feature extraction branches with the same structure are used to extract features from the XOZ view projection image sequence and the YOZ view projection image sequence respectively; each ResNet50 branch removes the last two layers of full connection layer and classification layer of the original network, and only retains to layer4 layer, and outputs two view feature tensors with the same dimension;

[0033] (c) Trainable multi-scale feature fusion: the two view feature tensors are adaptively fused through a fusion weight module, the fusion weight module sets one trainable parameter w (w [0,1]), which is used to dynamically control the fusion ratio of XOZ branch feature and YOZ branch feature, and the fusion formula is: f=w×F_XOZ+(1-w)×F_YOZ; Wherein f is the fused feature tensor, F_XOZ is the XOZ branch output feature, and F_YOZ is the YOZ branch output feature; the initial value of the trainable parameter w is 0.5, which represents that the two view feature weights are equal in the initial state, and the optimal feature combination in different action scenes is realized by automatically optimizing the model convergence through the back propagation algorithm in the training process;

[0034] (d) Transformer temporal modeling: the fused feature tensor f is input into a Transformer encoder module, which includes 2 serially connected Transformer Encoder Layers, each Encoder Layer internally sequentially setting a multi-head self-attention sublayer and a feedforward fully connected sublayer; wherein the multi-head self-attention sublayer adopts a 4-head attention mechanism for capturing the temporal correlation information in the feature tensor; the feedforward fully connected sublayer adopts a three-layer structure of Linear-ReLU-Linear, and the activation function is ReLU;

[0035] (e) Classification prediction and network optimization: after the Transformer encoder module models the input temporal features, a feature vector is output for each temporal sample, which is mapped to a 9-dimensional output vector through a 1-layer fully connected layer, corresponding to 9 human action categories of walking, sitting, standing, squatting, falling, bending, bending to straightening, waving, and stretching, and the probabilities of each category are calculated through a Softmax function to realize action classification prediction;

[0036] (f) Network layer parameter setting: all convolutional layers in the double-branch feature extraction and the feedforward fully connected sublayer of the Transformer encoder are set with ReLU activation function and BatchNorm (batch normalization) layer to reduce the risk of gradient disappearance, improve the model convergence speed and training stability; all pooling layers adopt maximum pooling operation, and the pooling kernel size is fixed at 2x2 and the pooling step is fixed at 2 to realize feature dimension reduction and key information preservation.

[0037] The beneficial effects of the present application are:

[0038] 1. The present application improves the traditional density clustering algorithm DBSCAN, based on the adaptive threshold mechanism and the spatial distribution of effective target points, which can effectively eliminate noise points and correct boundary points in complex environments, thereby significantly improving the spatial consistency and integrity of point cloud data, and enhancing the robustness and anti-interference ability of human action point cloud;

[0039] 2. The present application uses the time accumulation and three-dimensional voxelization resampling method of multiple frames of point cloud to convert sparse point cloud into structured regular point cloud, while maintaining the temporal continuity and spatial features of human action, avoiding the problems of traditional single-frame point cloud, and enhancing the point cloud density and stability, which is helpful for network learning of more discriminative action features;

[0040] 3. The present application extracts features from the enhanced point cloud through double-view (XOZ, YOZ plane projection), which enriches the diversity of input features, effectively avoids the information loss caused by single view, and improves the generalization ability and robustness of action recognition;

[0041] 4. The application is improved on the basis of the traditional ResNet50 network, combines the Transformer module with lightweight design, can capture local spatial geometric features and time correlation at the same time, realizes the efficient feature extraction and fusion of enhanced point cloud sequence; The accuracy of the method in the human action recognition task can reach 99.71%. BRIEF DESCRIPTION OF DRAWINGS

[0042] Figure 1 is the implementation process flow chart of the action recognition method of the application;

[0043] Figure 2 is the flow chart of the improved two-level density clustering algorithm of the application;

[0044] Figure 3 is a point cloud voxel block double-view projection diagram example of nine kinds of actions defined by the embodiment of the application. Each subgraph respectively corresponds to the XOZ and YOZ double-view projection of a representative sample of one kind of action category at a single frame time, wherein (a) walking, (b) sitting, (c) standing, (d) squatting, (e) falling down, (f) bending, (g) bending to straight up, (h) waving hands, (i) stretching lazily;

[0045] Figure 4 is the network flow chart of the DVP-TransNet network of the application, wherein (a) is the residual block structure, and (b) is the Transformer encoder structure;

[0046] Figure 5 is the action recognition confusion matrix diagram of the application under the embodiment. DETAILED DESCRIPTION

[0047] The preferred embodiments of the application will be described in detail below with reference to the accompanying drawings, so that the advantages and characteristics of the application can be more easily understood by those skilled in the art, and the protection scope of the application can be more clearly and definitely defined.

[0048] Reference Figures 1-5 A human action recognition method based on enhanced point cloud of millimeter wave radar, comprising the following steps:

[0049] (1) Build a millimeter wave FMCW radar experimental platform, set the radar parameters, and collect the human action echo signal;

[0050] (2) The human completes the preset action, the millimeter wave radar transmits the FMCW signal and receives the echo, gets the intermediate frequency signal through frequency mixing, and stores it as a.bin file in complex form;

[0051] (3)Firstly, static clutter is eliminated by frame difference processing to enhance the signal characteristics of the moving target, and then two-dimensional fast Fourier transform (2D-FFT) is performed on the intermediate frequency signal in each frame to obtain a range-doppler map (RDM);

[0052] (4)The generated RDM is subjected to two-dimensional ordered statistical constant false alarm rate detection (2D_OS_CFAR) to obtain an initial candidate point cloud of human motion, the candidate points are subjected to dynamic threshold screening, isolated noise points are removed by using an improved DBSCAN clustering algorithm, the point cloud cluster structure is enhanced, the point cloud quality is improved, and effective target points are obtained;

[0053] (5)For the extracted effective target points, the radial distance R and the radial velocity Z are calculated according to the radar configuration, and then the two-dimensional plane coordinates X and Y are obtained according to the target azimuth angle, and the target three-dimensional point cloud (X, Y, Z) is obtained in combination with the radial velocity Z;

[0054] (6)Subsequently, a time sequence feature accumulation and three-dimensional voxelization method is adopted. Seven consecutive three-dimensional point clouds are organized into time sequence samples in a sliding step of 2 frames, each frame of point cloud in the sample is independently subjected to voxelization processing, and is divided into a three-dimensional voxel grid of 10 (X) × 32 (Y) × 32 (Z), and the number of points in each non-empty voxel is counted to generate a discrete voxel tensor. Subsequently, each frame of tensor in the voxel sequence is respectively projected along the Y axis and the X axis to obtain a double-view projection image sequence of the XOZ plane and the YOZ plane;

[0055] (7)The double-view projection image is input into an improved deep learning network DVP-TransNet, which is composed of a double-branch feature extraction module of ResNet50, a self-learning weight feature fusion module (the initial value of the trainable parameter w is 0.5) and a Transformer encoder module, to extract time sequence features and perform action classification and recognition.

[0056] In this embodiment, the TI IWR1843 radar and the ADC1000EVM data acquisition card are used for echo signal acquisition, and in this example, one transmitting antenna and four receiving antennas are enabled. In this embodiment, the radar operating frequency range is 76GHz-81GHz, the starting frequency is set to 77GHz, the frequency modulation slope is 40MHz / μs, the ADC sampling rate is 2MHz, each frame contains 128 Chirps, each Chirp contains 128 ADC sampling points, and 50 frames of data are collected for each action. The human subject completes the preset actions (walking, sitting, standing, squatting, falling, bending, bending to straightening, waving, stretching, etc.) about 2m in front of the radar, and the intermediate frequency echo signal is stored as a.bin file in complex form after being collected by the ADC.

[0057] In this embodiment, the original intermediate frequency signal is divided into a three-dimensional array [n_ADC, n_Chirp, n_Frame] according to the receiving channel. Frame difference operation is performed on adjacent frame signals to eliminate static clutter and enhance the characteristics of the moving target signal. Then, two-dimensional FFT is performed on the frame difference matrix: distance direction FFT in n_ADC dimension and Doppler direction FFT in n_Chirp dimension, so as to obtain the range-Doppler map (RDM) of each frame.

[0058] In this embodiment, two-dimensional ordered statistics constant false alarm rate detection (2D_OS_CFAR) is applied to the RDM to obtain initial candidate target points, and then secondary density clustering is performed. The algorithm flow is as shown in Figure 2 (b) dynamic threshold screening module (first level processing): according to the target point density of the current frame RDM, the screening threshold is adaptively selected, and RD_target_new is output; (c) improved DBSCAN clustering module (second level processing): load dynamic clustering parameters (ε, MinPts), complete the judgment and clustering of core points / boundary points / isolated noise points, and output effective target points with clustering labels. In this embodiment, according to the radar parameters, the range and Doppler information of the effective target points are converted into radial distance R and radial velocity Z; the target azimuth angle is estimated by combining the MUSIC algorithm, the plane coordinates (X, Y) are calculated, and the complete three-dimensional point cloud (X, Y, Z) is obtained. The enhanced point cloud is stored as an xlsx file according to the action category, frame number and point number.

[0059] Then, the three-dimensional point cloud is constructed into a time sequence sample of 7 consecutive frames with a sliding step of 2. Each frame of point cloud in the sample is independently divided into a 10×32×32 (X×Y×Z) voxel grid, the number of points in each non-empty voxel is counted, and a discrete three-dimensional voxel sequence is generated; the maximum intensity projection is performed on the voxel sequence frame by frame: the XOZ view image is obtained by projecting along the Y axis, and the YOZ view image is obtained by projecting along the X axis, so as to obtain a dual-view projection image sequence. Figure 3 The XOZ and YOZ dual-view projection images of a representative sample of each action category at a single frame moment are shown.

[0060] The DVP-TransNet network structure proposed in this embodiment is as shown in Figure 4As shown, it comprises: 1) a double-branch ResNet50 module: the left and right two branches respectively take the XOZ view image sequence and the YOZ view image sequence as input, and extract deep features of the projection images of different views through convolution layers, multiple groups of residual blocks (such as 3 residual blocks of residual layer 1, 4 residual blocks of residual layer 2, etc.) and global average pooling operations. Different from the traditional ResNet50, the end fully connected layer is removed, and only the features are extracted, not directly classified. 2) a trainable fusion module: receiving the features output by the double-branch ResNet50 module, setting the learnable fusion weight, the initial weight is 0.5, and then the combination ratio of XOZ and YOZ features can be automatically optimized to obtain the fusion features. 3) a Transformer encoder module: containing 2 Transformer encoders, each encoder has multi-head self-attention and feed-forward sublayer, and the fusion features are time series modeled to capture the dynamic changes of the action, and then time series average pooling processing is performed. 4) classification layer: the features after time series modeling are further processed through the fully connected layer, and finally the action category prediction result is output.

[0061] In this embodiment, the collected nine types of action data are respectively divided into training set and test set according to 8:2, and are trained by using cross-entropy loss function and Adam optimizer. Finally, the nine human actions in this embodiment are tested, and the confusion matrix of the test result is as shown in Figure 5 As shown, the average accuracy in the action recognition task reaches 99.67%, which verifies the effectiveness of the human action recognition method based on the enhanced point cloud of millimeter wave radar proposed by the present application.

[0062] The above is only an embodiment of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent process transformation using the content of the specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method for human motion recognition based on millimeter-wave radar-enhanced point clouds, characterized in that, The method includes the following steps: Step 1: Set up a millimeter-wave FMCW radar experimental platform, set radar parameters, and collect human motion echo signals; Step 2: The human body completes the preset action, the millimeter-wave radar transmits FMCW signal and receives the echo, mixes it to obtain intermediate frequency signal, and stores it as a complex .bin file; Step 3: First, static clutter is eliminated through frame difference processing to enhance the signal characteristics of moving targets. Then, a two-dimensional fast Fourier transform is performed on the intermediate frequency signal of each frame to obtain the range-Doppler diagram (RDM). Step 4: Perform two-dimensional ordered statistical constant false alarm rate detection on the generated RDM to obtain the initial candidate point cloud of human movement. Perform dynamic threshold screening on the candidate points, use the improved DBSCAN clustering algorithm to remove isolated noise points, enhance the point cloud cluster structure, improve the point cloud quality, and obtain effective target points. Step 5: For the extracted effective target points, calculate their radial distance R and radial velocity Z according to the radar configuration, then obtain the two-dimensional plane coordinates X and Y according to the target azimuth angle, and combine the radial velocity Z to obtain the target three-dimensional point cloud (X,Y,Z); Step 6: Organize 7 consecutive frames of 3D point clouds into a time series sample with a sliding step size of 2. Perform voxelization on each frame of the sample independently, dividing it into a 10(X)×32(Y)×32(Z) 3D voxel grid. Count the number of points in each non-empty voxel to generate a discrete voxel sequence. Then, perform maximum intensity projection on the tensor of each frame in the voxel sequence along the Y-axis and X-axis respectively to obtain a dual-view projection image sequence of the XOZ plane and the YOZ plane. Step 7: Input the dual-view projection image into the improved deep learning network DVP-TransNe, which consists of a ResNet50 dual-branch feature extraction module, a self-learning weight feature fusion module, and a Transformer encoder module to extract temporal features and perform action classification and recognition.

2. The human motion recognition method based on millimeter-wave radar-enhanced point clouds according to claim 1, characterized in that, In step 1, the millimeter-wave radar FMCW experimental platform refers to the IWR1843 millimeter-wave FMCW radar developed by Texas Instruments. Its operating frequency range is from 76 GHz to 81 GHz, and it is configured with 1 transmitting antenna and 4 receiving antennas. The starting frequency and chirp slope are set to 77 GHz and 40 MHz / μs, respectively. The ADC sampling rate is 2 MHz, the number of chirps per frame is 128, the number of ADC samples per chirp is 128, and 50 frames of data are collected for each action.

3. A human motion recognition method based on millimeter-wave radar-enhanced point clouds according to claim 1 or 2, characterized in that, In step 2, the human body performs nine actions in front of the radar: walking, sitting, standing, squatting, falling, bending over, bending over and straightening up, waving and stretching. Each action lasts for about 1-2 seconds. The ADC1000EVM acquisition board stores the received intermediate frequency signal containing the human body action echo information as a complex .bin file.

4. The human motion recognition method based on millimeter-wave radar-enhanced point clouds according to claim 3, characterized in that, The process of step 3 is as follows: (a) According to the radar parameters, the raw intermediate frequency signal data acquired by the acquisition card is divided into multiple channel data and saved as I / Q data in complex form, which is the intermediate frequency echo .bin file. Then, the data of each receiving antenna is read from it and organized into a three-dimensional array [n_ADC, n_Chirp, n_Frame], where n_Frame represents the number of sampling frames, n_Chirp represents the number of Chirps in each frame, and n_ADC represents the number of sampling points in each Chirp. Then, frame difference processing is performed on it, that is, the data of the previous frame is subtracted from the data of the next frame in order to eliminate static clutter, improve the signal-to-noise ratio, and output the two-dimensional I / Q data after frame difference. (b) Based on the two-dimensional I / Q data after frame difference processing, 2D-FFT is performed on the frame difference matrix corresponding to each receiving antenna, where range FFT is performed along the n_ADC dimension and Doppler FFT is performed along the n_Chirp dimension, thereby obtaining the RDM of each frame.

5. The human motion recognition method based on millimeter-wave radar-enhanced point clouds according to claim 4, characterized in that, The process of step 4 is as follows: (a) Apply the 2D_OS_CFAR detection algorithm to RDM for target detection processing. By setting a sliding window and guard unit in both the distance dimension and the Doppler dimension, the background noise estimate is extracted in the two-dimensional region. The signal strength of each unit is judged according to the preset false alarm probability threshold, thereby identifying potential human target points as initial candidate points for subsequent point cloud construction and cluster recognition. (b) After the initial detection is completed, firstly, the background energy is extracted in the region that does not contain the initial candidate points, and its median energy is calculated as the background noise reference value; then, the average value of the median energy of the initial candidate points and the background noise reference value is calculated as the dynamic threshold value for subsequent discrimination. (c) Based on the energy distribution characteristics, the initial candidate points are screened a second time. If the density of the target point in the current frame RDM is higher than the set value, the background noise reference value is used as the screening threshold to remove candidate points with energy lower than the threshold. Otherwise, the dynamic threshold value in (b) above is used as the threshold to screen the initial candidate points again to obtain a target point set RD_target_new with higher confidence. The purpose of setting the dynamic threshold is to balance the false alarms caused by the conventional CFAR detection threshold being too small and the false alarms caused by the threshold being too large, thereby improving the purity of point cloud extraction and recognition robustness. (d) Improved DBSCAN clustering of RD_target_new: Extract the distance dimension and Doppler dimension coordinates of each target point from RD_target_new. Based on the spatial distribution characteristics of millimeter-wave radar point cloud, set dynamic clustering parameters: neighborhood radius ε, minimum number of points MinPts. Traverse each target point. If the number of target points contained in the ε neighborhood of the point is ≥ MinPts, then mark it as a core point; if the ε neighborhood of the point does not contain a core point, but the distance to the core point is ≤ ε, then mark it as a boundary point. If a point is neither a core point nor a boundary point, it is identified as an isolated noise point and removed. Finally, the core points and boundary points are merged into point cloud clusters according to their connectivity, and a clustering label map is constructed to obtain the effective target points. This level of processing can further eliminate the isolated noise points remaining after the first level of screening, optimize the spatial continuity of the point cloud clusters, and ensure that the effective target point cloud matches the actual shape of the human body movement.

6. The human motion recognition method based on millimeter-wave radar-enhanced point clouds according to claim 5, characterized in that, In step 5, the range dimension and Doppler dimension index of the target point can be obtained from the clustering label map. The radial distance R and radial velocity Z of the effective target point relative to the radar are calculated according to relevant formulas and radar parameters. Then, the azimuth angle of the target point is obtained using the MUSIC algorithm. The coordinates X and Y of the target point in the two-dimensional Cartesian coordinate system can be calculated, and together with the radial velocity Z, they constitute the three-dimensional point cloud features (X,Y,Z) of the target point. The point cloud data of the target point is stored in xlsx format according to action category, receiving antenna number, frame number, and point number. The data of each action category is divided into training set and test set according to the acquisition time sequence and a set ratio for subsequent use.

7. The human motion recognition method based on millimeter-wave radar-enhanced point clouds according to claim 6, characterized in that, In step 6, multiple consecutive frames of 3D point cloud data are accumulated and fused to enhance the target point cloud in the time dimension; then the fused 3D point cloud frame is uniformly divided into a 3D voxel grid with a size of 10×32×32 (x×y×z) to make the sparse point cloud denser. The number of point clouds in each non-empty voxel block is counted to generate a three-dimensional voxel tensor; then the voxel feature tensor is compressed along the Y-axis and X-axis directions respectively to generate XOZ and YOZ dual-view projection image sequences.

8. The human motion recognition method based on millimeter-wave radar-enhanced point clouds according to claim 7, characterized in that, The process of step 7 is as follows: (a) Training data preparation: Input the training set xlsx format point cloud data divided in the 8:2 ratio in claim 6 into the "temporal frame sequence voxelization-dual view projection" process described in claim 7 to generate XOZ plane projection image sequence and YOZ plane projection image sequence that correspond one-to-one with the training set point cloud, which together constitute the training dataset of DVP-TransNet. (b) Dual-branch feature extraction: Two independent ResNet50 feature extraction branches with completely identical structures are used to extract features from the XOZ view projection image sequence and the YOZ view projection image sequence respectively; each ResNet50 branch removes the last two fully connected layers and the classification layer of the original network, retaining only layer 4, and outputs two view feature tensors with the same dimension. (c) Trainable multi-scale feature fusion: The feature tensors of the two perspectives are adaptively fused through a fusion weight module. The fusion weight module is set with one trainable parameter w, w∈[0,1], which is used to dynamically control the fusion ratio of XOZ branch features and YOZ branch features. The fusion formula is: f=w×F_XOZ+(1-w)×F_YOZ; where f is the fused feature tensor, F_XOZ is the XOZ branch output feature, and F_YOZ is the YOZ branch output feature. The initial value of the trainable parameter w is set to 0.5, which means that the feature weights of the two perspectives are equal in the initial state. During the training process, the backpropagation algorithm is used to automatically optimize as the model converges to achieve the optimal feature combination under different action scenarios. (d) Transformer Temporal Modeling: The fused feature tensor f is input into the Transformer encoder module, which contains two cascaded Transformer Encoder Layers. Each Encoder Layer contains a multi-head self-attention sub-layer and a feedforward fully connected sub-layer. The multi-head self-attention sub-layer uses a 4-head attention mechanism to capture the temporal correlation information in the feature tensor. The feedforward fully connected sub-layer uses a three-layer structure of "Linear-ReLU-Linear" with ReLU activation function. (e) Classification prediction and network optimization: After modeling the input temporal features, the Transformer encoder module outputs a feature vector for each temporal sample. This feature vector is mapped to a 9-dimensional output vector through a fully connected layer, corresponding to 9 human action categories: walking, sitting, standing, squatting, falling, bending over, bending over to straighten up, waving, and stretching. The probability of each category is calculated through the Softmax function to achieve action classification prediction. (f) Network layer parameter settings: All convolutional layers and feedforward fully connected sub-layers of the Transformer encoder in the dual-branch feature extraction are set with ReLU activation function and BatchNorm to reduce the risk of gradient vanishing and improve the convergence speed and training stability of the model; all pooling layers adopt max pooling operation, with the pooling kernel size fixed at 2×2 and the pooling stride fixed at 2 to achieve feature dimensionality reduction and key information preservation.