A Human Action Recognition Method Based on Deep Learning and Ultra-Wideband Radar

Through deep learning-based methods and ultra-wideband radar technology, a three-dimensional feature data set is constructed and the PointNet network model is used to solve the problems of insufficient resolution and waste of human movement recognition in the existing technology, and high-accuracy action recognition is achieved in multi-objective scenarios.

CN113850204BActive Publication Date: 2025-06-03TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111145029.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-28
Publication Date
2025-06-03
Estimated Expiration
2041-09-28

AI Technical Summary

Technical Problem

When the prior art extracts Doppler features of human body movements from signals collected from ultra-wideband radars, the resolution of the two-dimensional time-frequency graph is not high, information wastes are serious, and the time-frequency features are easy to alias and difficult to segment in multi-target scenarios.

Method used

Using a deep learning-based method, a data set of human body movement information is obtained through ultra-wideband radar, distance-Doppler imaging, adaptive threshold detection and multi-object segmentation processing is performed, a three-dimensional feature data set is constructed, and a PointNet network model is used for training and testing to identify human body movements.

Benefits of technology

It improves the accuracy of human body movement recognition, can easily separate the action features of different targets in multiple scenes, reduces information waste, and improves the resolution of two-dimensional feature images and the integrity of action information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113850204B_ABST
    Figure CN113850204B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of radar detection, and discloses a human action recognition method based on deep learning and ultra-wideband radar, including: obtaining a human action information data set through an ultra-wideband radar module, and performing feature extraction to extract the action features of the target; successively performing range-Doppler imaging, adaptive threshold detection and multi-target segmentation processing on the data of the human action information data set, constructing a three-dimensional feature data set including time, range and velocity, and performing label marking according to the actual occurrence time of each action; dividing the three-dimensional feature data set into a training set and a test set, building a PointNet network model to train and test the network; extracting human action features from the real-time collected radar signals and constructing three-dimensional feature data, and inputting the trained PointNet network model to recognize the action types. The present invention can realize human action classification and recognition with a higher recognition rate, and can be applied in aspects such as physical security and intelligent detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of radar detection, and particularly relates to a human action recognition method based on deep learning and ultra-wideband radar. Background Art

[0002] During hostage rescue and personnel search and rescue, for the blind area of cameras, using radar to monitor the real-time situation of hostages, terrorists or trapped people can create favorable conditions for the rescue tasks of relevant personnel. In addition, in private places such as bathrooms, dressing rooms, and wards, using radar to monitor human postures can alarm for abnormal situations of the elderly, children or patients, and taking timely and effective treatment measures is very important for protecting their health and life safety. At present, certain progress has been made in human action recognition technology. Common human detection means mainly include video monitoring, infrared sensor monitoring, wearable devices, and radar system monitoring, etc. Among them, radar system monitoring has a wider detection range (from a few meters to several kilometers), is not sensitive to light changes, has better robustness to visual obstacles, and is also less affected by weather factors such as rain and haze. Compared with cameras, radar monitoring technology can better protect personnel privacy. Compared with technologies such as infrared imaging, radar has better anti-interference ability in complex environments.

[0003] After the radar echo signal is collected, the analysis of the radar echo signal usually uses the micro-Doppler effect for time-frequency analysis in the time-frequency domain and analysis in the time domain using high range resolution profiles. However, both have their limitations: the former will ignore the range information of human actions, while the latter will ignore the Doppler information, that is, the velocity information. At present, researchers use the micro-Doppler effect of radar signals to extract features of human actions more, and use time-frequency analysis methods such as short-time Fourier transform, wavelet transform, generalized S transform, Hilbert-Huang transform, and Wigner-Ville distribution to analyze the echo signals of slow-time targets. These time-frequency analysis methods combined with deep learning classification and recognition algorithms for human action recognition can basically achieve action classification and recognition. However, the obtained time-frequency diagram only contains the time and Doppler information of the action, and there will be a phenomenon of time-frequency diagram aliasing when detecting a multi-person scene, and it is difficult to accurately identify the actions of each target in a multi-person scene, and the ignored range information will cause waste of information.

[0004] Therefore, making full use of the time, range, and Doppler information in the echo signal will help improve the accuracy of human action recognition. Summary of the Invention

[0005] The present invention overcomes the deficiencies of the prior art. When extracting the Doppler characteristics of human movements from the signals collected by an ultra-wideband radar, there are problems such as low image resolution and information waste in the two-dimensional time-frequency diagram, and it is difficult to segment the time-frequency characteristics when there are multiple targets. The present invention provides a human movement recognition method based on deep learning and an ultra-wideband radar.

[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A human movement recognition method based on deep learning and an ultra-wideband radar, comprising the following steps:

[0007] S1. Obtain a human movement information data set through an ultra-wideband radar module, and perform feature extraction to extract the movement features of the target.

[0008] S2. Perform distance-Doppler imaging, adaptive threshold detection, and multi-target segmentation processing on the data of the human movement information data set in sequence, construct a three-dimensional feature data set including time, distance, and speed, and perform label marking according to the actual occurrence time of each movement.

[0009] S3. Divide the three-dimensional feature data set obtained in step S2 into training set data and test set data, build a PointNet network model, and train and test the network.

[0010] S4. After the network training is completed, extract the human movement features from the real-time collected radar signals through data processing, then construct three-dimensional feature data, and input it into the trained PointNet network model to identify the types of movements.

[0011] In the above step S1, obtaining the human movement information data set specifically includes the following steps:

[0012] Use an ultra-wideband radar to collect human movement information data and preprocess the data.

[0013] Use a publicly available motion data set to establish a human model in motion, and set the same signal parameters as the ultra-wideband radar in the simulation model to collect human simulated motion data.

[0014] Take the collection of human simulated motion data and the human movement information data collected by the ultra-wideband radar as the data of the human movement information data set.

[0015] The motion data in the human movement information data set includes single-person motions and multi-person motions.

[0016] The motion data includes seven motions: jogging, walking, jumping, climbing stairs, bending over, sitting down, and standing up, and each motion has at least 1000 samples.

[0017] In the step S2, the specific method for sequentially performing range-Doppler imaging, adaptive threshold detection, and multi-target segmentation on the data of the human motion information dataset is as follows:

[0018] Perform range-Doppler imaging operation, and use the range migration compensation method based on Keystone transform to eliminate the Doppler migration effect;

[0019] Adopt the CA-CFAR algorithm for adaptive threshold detection;

[0020] Finally, separate the point cloud features of multiple human targets through the DBSCAN clustering algorithm.

[0021] In the range-Doppler imaging process of the described human motion recognition method based on deep learning and ultra-wideband radar, first, the received signal is divided into multiple sub-bands after Fourier transform. The sub-bands contain the phase shift introduced by the initial target range, the Doppler frequency shift of the center frequency, and the fast frequency / slow time coupling term generated due to migration. Then, the coupling term is compensated through time scaling to eliminate the migration effect in the fast frequency / slow time domain. Finally, Sinc interpolation is used to resample the slow time of the Keystone shape matrix for inverse Fourier transform.

[0022] The specific steps of adopting the CA-CFAR algorithm for adaptive threshold detection are as follows:

[0023] Use a sliding 2D CFAR window to scan pixel by pixel in the whole image to extract valid targets in the image. The 2D CFAR window is divided into an internal cell under test covering the target features, an external reference cell covering the background area around the target pixels, and a guard cell between the cell under test and the reference cell;

[0024] Then, compare the energy ratio between the cell under test and the reference cell with the set threshold to determine whether the cell under test is a target.

[0025] The specific steps of separating the point cloud features of multiple human targets through the DBSCAN clustering algorithm in the described human motion recognition method based on deep learning and ultra-wideband radar are as follows:

[0026] Arbitrarily select a data object point from the target point set, select appropriate neighborhood radius and density threshold. If the number of points within the neighborhood radius of this point is greater than the density threshold, classify it as a core point, and then find another point within its neighborhood radius to determine whether it is a core point. All core points and the data object points within their neighborhood radii form a clustering cluster, and repeat the operation until all points are processed.

[0027] In the step S3, the ratio of the training set data to the test set data is 8:2.

[0028] The PointNet network model includes a T-Net network, a multi-layer perceptron MLP, and a Max pooling layer.

[0029] The PointNet network model uses the negative log-likelihood loss function NLLLoss as the loss function, performs log_softmax processing on the classification score map, and obtains the loss value by removing the negative sign from the value corresponding to the actual label in the result and then summing and taking the average; uses the mean intersection over union mIoU to evaluate the performance of the network model, that is, after summing the ratios of the intersection to the union of the prediction results and the ground truth values for each type of action, the average is taken to obtain the result; selects Adam as the optimizer, dynamically adjusts the learning rate of each parameter using the first-order moment estimate and second-order moment estimate of the gradient, updates the weights of the network model, and saves the optimal result.

[0030] The present invention has the following beneficial effects compared with the prior art:

[0031] 1. The present invention provides a human action recognition method based on deep learning and ultra-wideband radar, which extracts three-dimensional action features of actions from radar echo signals, can comprehensively utilize time, distance, and speed information, and can improve the recognition rate of actions; moreover, for multi-person scenarios, the aliasing part of the three-dimensional point cloud features is greatly reduced, and the action features of different targets can be easily separated, which is beneficial to the subsequent implementation of the segmentation and recognition of specific actions.

[0032] 2. The present invention combines a point cloud segmentation network in deep learning algorithms to segment and accurately recognize continuous human actions, which can greatly improve problems such as limited resolution of two-dimensional feature images, incomplete action information, low recognition rate under small sample data sets, and difficulty in extracting action features in multi-target scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a schematic flow chart of the human action recognition method based on deep learning and ultra-wideband radar provided by the present invention;

[0034] Figure 2 is a network structure diagram of the PointNet network used in the embodiment of the present invention to achieve action segmentation and classification recognition. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0036] As Figure 1 shown, to solve the problems such as limited resolution of two-dimensional feature images, incomplete motion information, low recognition rate under small sample data sets, and difficulty in segmenting motion features in multi-target scenarios in human motion detection, an embodiment of the present invention provides a method for human motion recognition based on a deep learning and ultra-wideband radar monitoring system, including the following steps.

[0037] S1. Obtain a human motion information data set through an ultra-wideband radar module, and perform feature extraction to extract the motion features of the target.

[0038] Specifically, in this embodiment, obtaining the human motion information data set specifically includes the following steps:

[0039] Use an ultra-wideband radar to collect human motion information data and preprocess the data;

[0040] Use a publicly available motion data set to establish a human model in motion, and set signal parameters identical to those of the ultra-wideband radar in the simulation model to collect human simulation motion data;

[0041] Regard the collection of human simulation motion data and the human motion information data collected by the ultra-wideband radar as the data of the human motion information data set.

[0042] Furthermore, in this embodiment, the PulsON 440 (P440) ultra-wideband radar module designed by Domain Company in the United States is used to collect actual human motion information. The data collected in the experiment needs to be processed for environmental noise reduction: before collecting the motion information of the human body, first use the P440 module to detect the echo signal in the no-person scenario in the experimental area. In subsequent human motion detection experiments, subtract the experimental data in the no-person scenario from the collected echo signal to reduce the influence of environmental factors on the detection result, and then use the simple average elimination method to filter out clutter, directly subtract the signal mean from the radar echo signal to remove the direct reflection echo of the obstacle, and finally use a four-tap difference filter to perform motion filtering on the signal.

[0043] Use the publicly available motion data set MOCAP to establish a human model in motion, set signal parameters identical to those of the P440 module in the simulation software to simulate the collection of human motion data, and regard it together with the experimental collection data as the data source of the human motion information collected by the radar. The data set contains seven actions: jogging, walking, jumping, climbing stairs, bending, sitting down, and standing up. Each action has 1000 samples, and they are divided into a training set and a test set according to the ratio of 8:2.

[0044] S2. Perform distance-Doppler imaging, adaptive threshold detection, and multi-target segmentation on the data of the human motion information dataset in sequence to construct a three-dimensional feature dataset including time, distance, and speed, and perform label marking according to the actual occurrence time of each action.

[0045] Specifically, in step S2, the specific methods for performing distance-Doppler imaging, adaptive threshold detection, and multi-target segmentation on the data of the human motion information dataset in sequence are as follows:

[0046] (1) Perform distance-Doppler imaging operation, and use the distance migration compensation method based on Keystone transform to eliminate the Doppler migration effect;

[0047] (2) Adopt the CA-CFAR algorithm for adaptive threshold detection;

[0048] (3) Finally, separate the point cloud features of multiple-person targets through the DBSCAN clustering algorithm.

[0049] Since the Doppler migration effect of the ultra-wideband radar will seriously reduce the radar's ability to distinguish similar velocity targets, and the target velocity can be represented by the Doppler frequency. When the Doppler spectra of two targets overlap, it may lead to the inability to distinguish targets with similar velocities. In this embodiment, Fourier transform is used to perform distance-Doppler imaging on the data frames collected by the radar, and Keystone transform is used to achieve distance migration compensation to weaken the influence brought by the migration effect.

[0050] The received signal contains three parts of information: the phase shift introduced by the initial target distance, the Doppler frequency shift of the center frequency, and the fast frequency / slow time coupling term generated due to migration. Compensating the coupling term through time scaling in the fast frequency / slow time domain can alleviate the migration effect, re-adjust the time axis for each frequency to obtain the output of the high-resolution matched filter, and improve the velocity resolution of radar detection. Since after the Keystone transform, the time scales in each sub-band become different, it is necessary to resample the signal along the slow time in order to perform the inverse Fourier transform.

[0051] The following is the specific calculation process for realizing distance migration compensation by using Keystone transform in this embodiment.

[0052] The P440 module is an ultra-wideband radar module with a center frequency of 4.3 GHz and a frequency range of 3.1 - 5.3 GHz. Its transmitted pulse is a Gaussian monocycle pulse signal p(t):

[0053] ; (1)

[0054] Among them, Ais the amplitude parameter, set to 1; a is the slope constant of the pulse signal, , f c represents the center frequency. The human backscattering model can be represented by the coherent accumulation of the echoes of multiple scatterers. The echo signal of a single scatterer collected by the simulated radar system is , and its expression is:

[0055] ; (1)

[0056] where t’ represents the fast time and t s represents the slow time; τ 0 represents the initial time delay, , and R 0 represents the initial position of the single scatterer; , v represents the velocity of the single scatterer, and c represents the speed of light. To eliminate the migration effect in the fast frequency / slow time domain, first perform a Fourier transform on the fast time variable t’ of to obtain the expression:

[0057] ; (2)

[0058] where P ( f ) is the Fourier transform of p(t’). Generalize a single scatterer to a set of N moving point scatterers to obtain the complete formula:

[0059] ; (3)

[0060] In this process, range migration will occur. Use to re-adjust the time axis for each frequency to obtain the output of the high-resolution matched filter. Rewrite S ( f , t s ) to S ( f , t s ’) domain as:

[0061] ; (4)

[0062] Since after the Keystone transform, the time scales in each sub-band become different, use Sinc interpolation to resample the slow time of the Keystone matrix, and then perform the inverse Fourier transform:

[0063] ; (5)

[0064] For the slow time variablet s Perform a Fourier transform to obtain a frame of range-Doppler image:

[0065] ; (6)

[0066] Adopt the CA-CFAR algorithm for adaptive threshold detection to extract the motion information of each part of the human body in the range-Doppler image in a noise and clutter environment. When extracting target information, use a sliding 2D CFAR window to scan the entire image pixel by pixel to extract valid targets in the image. The 2D CFAR window is divided into an internal cell under test covering the target pixel, an external reference cell covering the background area around the target pixel, and a guard cell between the cell under test and the reference cell; calculate the energy values of the cell under test and the reference cell, compare the ratio of the two with a set threshold, and determine whether the cell under test is a target. Arrange the extracted target data in time frames to obtain three-dimensional motion features in continuous time, and then perform downsampling to reduce the data volume and speed up the calculation.

[0067] Specifically, in this embodiment, use a sliding 2D CFAR window to scan the entire image pixel by pixel to extract valid targets in the image. Set the 2D CFAR window as a 9×9 cell grid. Among them, the cell under test is 3×3 cells at the center, the reference cell is the outermost 3 layers of cells, and the middle 3 layers are reference cells. Calculate the energy of the cell under test CUT and the reference cell RC:

[0068] ; (7)

[0069] Among them, Y E represents the total energy value of all pixel cells in the cell under test, X E represents the total energy value of all pixel cells in the reference cell, and CUT(i, j) and RC(i, j) respectively represent the individual pixel values corresponding to the CUT area and the RC area in the detection image.

[0070] Compare the energy ratio of the cell under test and the reference cell with the set threshold to determine whether the cell under test is a target:

[0071] ; (8)

[0072] Among them, T represents the threshold, and an appropriate threshold can be selected according to the image characteristics. In this embodiment, set T = 0.5, The value represents whether the target cell is a valid feature; in addition, in the present invention, the energy of the reference cell should not be close to zero to avoid excessive false alarm detections and prevent the problem of division by zero.

[0073] Mark the corresponding digital tags for the extracted target points according to the specific actions that occur at the corresponding time. Jogging, walking, jumping, climbing stairs, bending down, sitting down, and standing up correspond to natural numbers 1-7 respectively. Shuffle the dataset in a ratio of 8:2 and divide it into a training set and a test set. The training set contains 5,600 samples, and the test set contains 1,400 samples.

[0074] Specifically, in this embodiment, for the multi-target scenario, the DBSCAN clustering algorithm is used to separate the point cloud features of different target persons, that is, the multi-person action information is segmented into single-person action information for recognition. Set the neighborhood radius ε and the point density M within the neighborhood radius centered on the core point, and divide all the points in the point cloud into core points, boundary points, and noise points. First, select a core point. When the density of points within the neighborhood radius of the selected core point is not less than M, randomly select a point from this range to determine whether it also meets the conditions. If it meets, add it to the clustering of the core point. If it does not meet, determine whether there is a core point within its neighborhood range. If there is, add it to the clustering of the boundary point. If not, it belongs to the clustering of the noise point. Divide the core points and boundary points in the same point cluster into one category to achieve the separation of the action features of the target in the multi-person scenario.

[0075] Specifically, in this embodiment, the two important parameters involved in the DBSCAN algorithm, the neighborhood radius ε and the density threshold M at the core point, are set to 0.5 and 15 respectively. Assume x ∈ X, is the neighborhood of x, and ρ(x) = |N ε (x)| is the density of x. If ρ(x) ≥ M, then x is a core point of X. Denote the set of core points as Xc, and the set of non-core points as Xnc; if x ∈ Xnc and falls within the neighborhood range of a certain core point, then x is a boundary point. The set of boundary points is X bd , and other points that are neither core points nor boundary points are called noise points. Divide the core points and boundary points in the same point cluster into one category, which can achieve the separation of the action features of the target in the multi-person scenario.

[0076] Establish a three-dimensional feature dataset of human body movements through the above operations, mark the obtained point cloud data with labels, and divide it into a training set and a test set in an appropriate ratio.

[0077] S3. Divide the three-dimensional feature dataset obtained in step S2 into training set data and test set data, build a PointNet network model, and train and test the network.

[0078] Build a PointNet network, and its network structure is as Figure 2As shown. The PointNet network mainly includes three processes: extracting the local features and global features of the point cloud, and merging the local features and global features. To build the PointNet network, its network structure is as Figure 2 As shown. The PointNet network mainly includes three processes: extracting the local features and global features of the point cloud, and merging the local features and global features. The specific method is as follows.

[0079] Each sample in the training set is data of n×7 dimensions (n represents the number of points, and the features of each point include its three-dimensional coordinate position, the RGB value representing color, and the digital label corresponding to the point). The coordinate data (n×3) among them is input into the T-Net network, and the T-Net network is used to train a 3×3 matrix to perform coordinate transformation on the input points, that is, multiplying it by the trained matrix to obtain the transformed point cloud coordinates; then, through a multi-layer perceptron of (64,64), the features of each point are extended to 64 dimensions, and then sent into the T-Net network again to multiply with a 64×64 matrix to perform feature transformation on the local features of each point, and a regularization term is added to obtain the local features of the point cloud, and the regularization term is added to reduce the information loss of the point cloud; the local features are input into a multi-layer perceptron of (64,128,1024), and the global features of each point are extracted by using the multi-layer perceptron of (64,128,1024) and the Maxpooling function, that is, the global features of the point cloud are obtained after performing the max pooling operation; after merging the local features and global features of the point cloud, a score map of n×m is obtained through two multi-layer perceptrons of (512,256,12,8) and (128, m). The multi-layer perceptron includes an input layer, a hidden layer, and an output layer, and all layers are fully connected. Among them, the output layer of the last multi-layer perceptron is a Dropout layer to prevent the model from overfitting to improve the generalization ability. The point cloud is segmented into different subclasses by classifying and predicting each point, and the action category represented by each subclass is judged. Among them, m represents the number of action types, which is 7 in this example.

[0080] Then, the above network is trained according to the training set data. Since the number of points in each sample file varies greatly, all samples are divided into sub-samples with 4096 points in each sample according to the principle of local correlation. In this embodiment, 32 sub-samples are set to be read each iteration, and the number of iterations is determined by the total number of sub-samples.

[0081] During the iteration process, the maximum likelihood loss function NLLLoss is used to calculate the loss value, and the classification score map is processed by softmax to obtain y pre After taking the logarithm of it, the value corresponding to the actual label y true in the result is removed the negative sign and summed and averaged to obtain the loss value loss, and the expression is:

[0082] ; (9)

[0083] The Adam optimizer is selected to optimize the performance of the network model. Adam combines the advantages of Adagrad in dealing with sparse gradients and RMSprop in dealing with non-stationary targets, has a small memory requirement, and dynamically adjusts the learning rate of each parameter by using the first-order moment estimate and second-order moment estimate of the gradient:

[0084] ; (10)

[0085] where g t represents the gradient of the t-th round of training, m t and n t are the first-order moment estimate and second-order moment estimate of the gradient respectively, β 1 and β 2 are the decay factors of the first-order moment estimate and second-order moment estimate, taking β 1 = 0.9, β 2 = 0.999; m t ’, n t ’ are the bias corrections for m t and n t , approximately unbiased estimates of the expectations, β 1 t and β 2 t represent the cumulative values of the decay factors of the first-order moment estimate and second-order moment estimate after the t-th round of training; θ t-1 represents the parameter obtained by optimizing the parameter θ t at the t-th round of training, and ε = 1e-8 is used to prevent the denominator from being zero. The decay factor weighs the past and current gradient information during the learning rate update process, reduces the impact of the learning rate being significantly reduced due to continuous gradient accumulation, and prevents premature termination of learning.

[0086] In this embodiment, the mean intersection over union (mIoU) is used to evaluate the performance of the model. The mIoU is the result of summing and averaging the ratios of the intersection to the union of the prediction results and the ground truth for each class of the model. The model parameters with the best performance are saved by comparing the calculated mIoU with the mIoU of the previously saved model.

[0087] S4. After the test is completed, the human motion features are extracted from the real-time collected radar signals through data processing, and then three-dimensional feature data is constructed and input into the trained PointNet network model to identify the types of actions.

[0088] In this embodiment, the performance of the trained network model is tested using a test set. First, the training set is loaded. Similar to the operation of processing the training set, the test sample file is divided into subsamples containing 4096 points, and then the network is initialized using the trained model parameters. The local and global features of the points in all subsamples in each batch are learned to obtain a score map for the category to which each sample point belongs, and then the prediction result is obtained. By comparing it with the true result, the effectiveness of the human action recognition method proposed in the present invention based on deep learning and an ultra-wideband radar monitoring system is verified. When identifying the type of action of the radar signal collected in real time, the data can be processed by the data processing method in step S2 for range-Doppler imaging, adaptive threshold detection, and multi-target segmentation processing. After constructing three-dimensional feature data including time, range, and velocity, it is then input into the trained PointNet network model to identify the type of action.

[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A human action recognition method based on deep learning and ultra-wideband radar, characterized in that, it includes the following steps: S1. Obtain a human action information dataset through an ultra-wideband radar module, and perform feature extraction to extract the action features of the target; S2. Sequentially perform range-Doppler imaging, adaptive threshold detection, and multi-target segmentation processing on the data of the human action information dataset, construct a three-dimensional feature dataset including time, range, and velocity, and perform label marking according to the actual occurrence time of each action; S3. Divide the three-dimensional feature dataset obtained in step S2 into training set data and test set data, build a PointNet network model, and train and test the network; S4. After the network training is completed, extract the human action features from the real-time collected radar signals through data processing, then construct three-dimensional feature data, and input it into the trained PointNet network model for action type recognition; In step S2, the specific methods for sequentially performing range-Doppler imaging, adaptive threshold detection, and multi-target segmentation processing on the data of the human action information dataset are: Perform range-Doppler imaging operation, and use the range migration compensation method based on Keystone transform to eliminate the Doppler migration effect; Adopt the CA-CFAR algorithm for adaptive threshold detection; Finally, separate the point cloud features of multiple-person targets through the DBSCAN clustering algorithm.

2. A human action recognition method based on deep learning and ultra-wideband radar according to claim 1, characterized in that, in step S1, obtaining the human action information dataset specifically includes the following steps: Collect human action information data using an ultra-wideband radar, and preprocess the data; Establish a human body model in motion using a publicly available motion dataset, and set the same signal parameters as the ultra-wideband radar in the simulation model to collect human simulation motion data; Take the collection of human simulation motion data and the human action information data collected by the ultra-wideband radar together as the data of the human action information dataset.

3. A human action recognition method based on deep learning and ultra-wideband radar according to claim 1, characterized in that, the action data in the human action information dataset includes single-person actions and multi-person actions; the action data includes seven actions: jogging, walking, jumping, climbing stairs, bending, sitting down, and standing up, and each action has at least 1000 samples.

4. A human action recognition method based on deep learning and ultra-wideband radar according to claim 1, characterized in that, during the range-Doppler imaging processing, first divide the received signal into multiple sub-bands after Fourier transform, and the sub-bands contain the phase shift introduced by the initial target range, the Doppler frequency shift of the center frequency, and the fast frequency / slow time coupling term generated due to migration; Then compensate the coupling term through time scaling to eliminate the migration effect in the fast frequency / slow time domain; finally, resample the slow time of the Keystone shape matrix using Sinc interpolation for inverse Fourier transform.

5. A human action recognition method based on deep learning and ultra-wideband radar according to claim 1, characterized in that, the specific steps of using the CA-CFAR algorithm for adaptive threshold detection are as follows: Use a sliding 2D CFAR window to scan pixel by pixel in the whole image to extract effective targets in the image; the 2D CFAR window is divided into an internal unit under test covering target features, an external reference unit covering the background area around the target pixel, and a protection unit between the unit under test and the reference unit; Then compare the energy ratio of the unit under test and the reference unit with the set threshold to determine whether the unit under test is a target.

6. A human action recognition method based on deep learning and ultra-wideband radar according to claim 1, characterized in that, the specific steps of separating the point cloud features of multiple-person targets by the DBSCAN clustering algorithm are as follows: Arbitrarily select a data object point from the target point set, select appropriate neighborhood radius and density threshold. If the number of points within the neighborhood radius of this point is greater than the density threshold, classify it as a core point, and then find another point within its neighborhood radius to determine whether it is a core point. All core points and data object points within their neighborhood radius form a clustering cluster, and repeat the operation until all points are processed.

7. A human action recognition method based on deep learning and ultra-wideband radar according to claim 1, characterized in that, in step S3, the ratio of the training set and the test set data is 8:

2.

8. A human action recognition method based on deep learning and ultra-wideband radar according to claim 1, characterized in that, The PointNet network model includes a T-Net network, a multi-layer perceptron MLP, and a Max pooling layer.

9. A human action recognition method based on deep learning and ultra-wideband radar according to claim 1, characterized in that, The PointNet network model uses the negative log-likelihood loss function NLLLoss as the loss function, performs log_softmax processing on the classification score map, and takes the average value of the sum of the values corresponding to the actual labels in the result after removing the negative sign to obtain the loss value; uses the mean intersection over union mIoU to evaluate the performance of the network model, that is, after summing the ratios of the intersection and union of the prediction results and the true values for each type of action, take the average to get the result; Select Adam as the optimizer, use the first-order moment estimate and second-order moment estimate of the gradient to dynamically adjust the learning rate of each parameter, update the weights of the network model and save the optimal result.

Citation Information

Patent Citations

  • Radar target recognition method based on micro-Doppler feature extraction and deep learning

    CN108256488A

  • An ultra-wideband radar human body action recognition method based on a deep convolutional neural network

    CN109948532A