A fast multi-scale through-wall radar human behavior recognition method and device

By synchronously extracting and identifying the translation and microDoppler features in the wall-through radar, the problem of low accuracy and high calculation cost of human motion recognition method of the wall-through radar in complex environments is solved, and efficient and fast recognition effect is achieved.

CN116338680BActive Publication Date: 2025-08-05BEIJING INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310293640.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-23
Publication Date
2025-08-05
Estimated Expiration
2043-03-23

AI Technical Summary

Technical Problem

The existing wall-through radar human motion recognition method has low recognition accuracy, high calculation cost, slow training and reasoning speed, and is difficult to deploy in complex environments.

Method used

The lightweight deep neural network model is adopted, combining super-resolution Doppler estimation based on kernel distance and lightweight non-anchor point single-stage object detection network to synchronize translation information and micro-Doppler features in complex environments, and judgment recognition is performed through lightweight convolutional neural network.

Benefits of technology

It realizes high recognition accuracy, reduces calculation costs and training inference time, improves the robustness and real-time nature of the system, and is suitable for complex scene recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116338680B_ABST
    Figure CN116338680B_ABST
Patent Text Reader

Abstract

The present invention discloses a fast multi-scale through-wall radar human behavior recognition method and device. Based on frequency-stepped radar data, the present invention combines a super-resolution Doppler estimation model based on kernel distance, a lightweight non-anchor point single-stage target detection network model, and a lightweight classification network model to construct a lightweight deep neural network capable of mining multi-scale feature information. By simultaneously extracting the translational motion information of the human body in a complex through-wall environment and the generated μ-D features, and combining the information at both scales for judgment and recognition, the through-wall radar human motion recognition result is ultimately provided. The present invention method can achieve a high recognition accuracy, and because its submodule construction methods are all based on lightweight models, the overall computational cost of the algorithm is significantly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of radar signal processing, and in particular relates to a fast multi-scale through-wall radar human behavior recognition method and device. Background Art

[0002] In recent years, there has been growing interest in deploying wall-penetrating radars to detect, track, and monitor human activity. Research areas include security surveillance, falls, breathing, and heartbeat detection. Human motion recognition behind walls is one of the most widely used applications. Humans are slow-moving, non-rigid targets, and the temporally varying motion of their torsos, arms, legs, hands, and feet modulates the carrier frequency of radar signals. In addition to the Doppler shift caused by the overall translation of the target, these internal body movements also produce micro-Doppler (μ-D) signatures in radar echoes. These μ-D components are typically detected using ultra-wideband (UWB) radars, which offer sub-decimeter range resolution. However, due to the effects of walls, such as attenuation, refraction, and multipath, UWB-penetrating radars introduce significant distortion to the echo signal, significantly reducing the accuracy of human motion recognition and significantly increasing the computational cost of available models, making system deployment extremely challenging. Therefore, there is an urgent need to study a multi-scale through-wall radar human behavior recognition method and device that has high recognition accuracy, low model parameter count, less forward reasoning calculation amount, and faster prediction speed.

[0003] Recent research in through-the-wall radar human motion recognition focuses on using μ-D features to classify human activities, distinguishing between armed and unarmed individuals, and modeling the motion patterns of different body parts behind the wall to achieve behavior and gesture recognition. These methods range from heuristic models to more complex statistical learning techniques, including principal component analysis, independent component analysis, empirical mode decomposition, and the Hilbert-Huang transform. Discriminative features are extracted from spectrograms and classified using algorithms such as support vector machines and Bayesian classifiers. Recent research has focused on using deep convolutional neural networks for single-stage feature extraction and action recognition. These neural networks can simultaneously learn modal and boundary information from measured data for classification without requiring any explicit feature extraction algorithms. However, to achieve high-precision feature extraction from through-the-wall radar imaging, corresponding recognition methods require improvements through widening and deepening strategies, which significantly increases the time complexity of training and inference, hindering system deployment and real-world application.

[0004] The main challenges in this field are the limited amount of data available to train machine learning algorithms to improve accuracy, and the computational complexity of the models used for recognition, which often results in slow training and inference. Current research on methods for simultaneously detecting the overall translational state of a target and μ-D features for human motion recognition is very limited. The present invention aims to address these challenges by:

[0005] 1. Propose a lightweight and accurate algorithm model that can simultaneously extract the translational information and μ-D features of the human body in complex wall-penetrating environments, and combine the information of the two scales for judgment and recognition;

[0006] 2. Verify, evaluate and optimize the model through measured data. Summary of the Invention

[0007] In view of this, the present invention provides a fast multi-scale through-wall radar human behavior recognition method and device, which can solve the technical problems caused by the influence of walls, such as attenuation, refraction and multipath effects, such as obvious distortion of ultra-wideband through-wall radar echo signals, a significant decrease in the accuracy of human behavior recognition, a significant increase in the computational cost of available models, and very challenging system deployment.

[0008] The technical solutions for implementing the present invention are as follows:

[0009] A fast multi-scale through-wall radar human behavior recognition method includes the following steps:

[0010] Step 1: Acquire the original image of the human behavior echo of the through-wall radar, and preprocess the original image to obtain frequency-stepped radar data, including the translational information and micro-Doppler characteristics of the human movement behind the wall; the preprocessing includes data enhancement, clutter and noise suppression;

[0011] Step 2: Input the frequency-stepped radar data into the backbone CNN of the human translation information feature extraction module. After passing through four layers of convolution, pooling, inter-layer normalization, and efficient attention weighting, a frequency-stepped radar data heat map containing some wall clutter and human motion Doppler features is obtained. The heat map is input into a four-layer FPN, where the scale of each layer is 1 / 2 of the previous layer. The heat map is further weighted using the efficient attention module, and the image is mapped into two vectors containing classification and regression box information using the detection head. Finally, the center point of each regression box is solved, and a curve fitting method based on Lagrange interpolation is used to perform curve annotation on the frequency-stepped radar data image. The resulting annotated image is the human translation information.

[0012] The human body translation information feature extraction module consists of four parts: a backbone convolutional neural network (CNN) for extracting features, a feature pyramid network (FPN) at the neck of the network, an efficient attention module embedded at the connection between each part, and a detection head using a classification branch and a regression branch;

[0013] Step 3: Perform a DTOF scan on the frequency-stepped radar data generated in Step 1 and the frequency-stepped radar data heat map generated in Step 2. Parameter estimation is performed using the echo amplitude to determine whether a suspicious target exists. The micro-Doppler time-frequency analysis method is then used to analyze the target characteristics of the suspicious object. A Gaussian kernel probability function is introduced into each variable during the parameter estimation process. Pixels whose estimated results fall within the human motion speed range are annotated and output. The output dot trace image is the micro-Doppler signature of human motion.

[0014] Step 4: For the human body translation information and micro-Doppler feature maps generated in steps 2 and 3, a feature dataset is generated using the image stitching method. Based on the generated feature dataset, a lightweight convolutional neural network is used for judgment and recognition, and finally the through-wall radar human behavior recognition result is given.

[0015] Furthermore, in step 1, the human body is modeled using a node model. Assuming that the human target is divided into P strong scattering points, the received echo is expressed as:

[0016]

[0017] Among them, S W (t) and S N (t) represents the wall echo and noise respectively, and t is the propagation time of the electromagnetic wave emitted by the radar in space. By using the time delay characteristics of the echo and summing the point target, we can obtain the delayed focusing imaging result of the echo. Represents the total echo signal of the human target. The round-trip delay corresponding to different nodes of the human body can be split into p scattered components, corresponding to p echoes, and the amplitude gain of each echo component is a p , the time domain response is s, and the delay is τ p .

[0018] Furthermore, the background-eliminated target echo signal Z(m,n) with n and m fast and slow time points respectively can be expressed as:

[0019] Z(m,n)=κ(m,n)-B(m,n)

[0020] =κ(m,n)-w T (n)φ r(m,n)

[0021] =w *T (n)φ r (m,n)

[0022] Among them, w *T (n)=[1-w0(n),-w1(n),…,-w L (n)] T To load the time domain echo φ after background removal for estimation r The weight on (m,n), φ r (m,n) is the time domain echo component used for background estimation when the fast-slow time points are n and m respectively, κ(m,n) is the complete time domain echo when the fast-slow time points are n and m respectively, B(m,n) is the background estimation result when the fast-slow time points are n and m respectively; w T (n)=[w0(n),w1(n),…,w L (n)] T is the time domain echo component φ loaded into the background estimation r The weights on (m,n), where T is the matrix transpose operator.

[0023] Furthermore, the specific calculation process of the efficient attention module is as follows: first, the input feature map is subjected to a global average pooling operation; secondly, a one-dimensional convolution operation with a convolution kernel size of 3 is performed, and the weights of each channel are obtained through a Sigmoid activation function; finally, the weights are multiplied by the corresponding elements of the original input feature map to obtain the final output feature map.

[0024] Furthermore, the behavioral states of the human body behind the wall are divided into seven categories, including the scene of the human body moving back and forth parallel to the wall, the scene of the human body moving back and forth diagonally behind the wall, the scene of the human body rotating behind the wall, the scene of the human body squatting behind the wall, the scene of the human body stepping on the spot behind the wall, the scene of the human body standing still behind the wall and the empty scene.

[0025] A fast multi-scale through-wall radar human behavior recognition device, comprising an image processing module, a feature extraction module, and a classification module;

[0026] Image processing module: pre-processes the original image of the human behavior echo obtained by the through-wall radar; the pre-processing includes data enhancement, clutter and noise suppression, and obtains the processed image;

[0027] Feature Extraction Module: This module combines a kernel-distance-based super-resolution Doppler estimation model, a lightweight non-anchor single-stage target detection neural network (DET), and a lightweight classification neural network (LNN) to construct a lightweight deep neural network capable of mining multi-scale feature information. It simultaneously extracts the translational motion information and the resulting μ-D features of the human body in complex wall-penetrating environments, and combines the information at both scales to generate a feature dataset.

[0028] Classification module: The obtained feature data set is judged and classified through a lightweight convolutional neural network to obtain the final recognition result.

[0029] Beneficial effects:

[0030] The method of the present invention can achieve a high recognition accuracy, and because the construction method of its sub-modules is based on a lightweight model, the overall computational cost of the algorithm is greatly reduced, making it possible for system deployment and rapid real-time recognition.

[0031] Compared with the prior art, the present invention has the following technical effects:

[0032] (1) The present invention discloses a fast multi-scale through-wall radar human behavior recognition method and device, which differs from traditional through-wall radar human behavior recognition technology. It adopts a dual-stage working mode of feature extraction and classification recognition, achieving high-accuracy and high-information-utilization recognition under adverse conditions such as fuzzy radar imaging and low signal-to-noise ratio.

[0033] (2) The present invention discloses a fast multi-scale through-wall radar human behavior recognition method and device, which combines the physics background to design the intelligent training and reasoning architecture of the present invention, making the method more interpretable and faster in response.

[0034] (3) The present invention discloses a fast multi-scale through-wall radar human behavior recognition method and device, which can achieve 93.98% recognition accuracy when all modules work simultaneously. When a module fails to be deployed, the remaining modules can still maintain a high recognition accuracy of more than 90%. When the classification module is replaced with other similar lightweight networks, the overall recognition accuracy, training time, inference time, and parameter quantity of the method are improved. The input of the network is trained and verified using the feature maps output by the target detection module and micro-Doppler labeling module proposed in this patent. The comparison results prove that the proposed method can greatly improve the execution speed of the algorithm while maintaining high recognition accuracy, reduce computing and memory costs, and greatly improve the robustness and practicability of the present invention. Compared with traditional recognition inventions, the present invention is more helpful for complex scenarios such as through-wall military recognition, through-wall civilian human recognition, through-wall life signal detection, and disaster relief.

[0035] (4) The present invention discloses a fast multi-scale through-wall radar human behavior recognition method and device, whose link construction structure follows the cascade-parallel mode, and the feature extraction and classification recognition methods also have a fixed theoretical basis. This method does not require too much prior knowledge and is simple and easy to implement, ensuring the high applicability of the method.

[0036] (5) The present invention has high information utilization rate and recognition accuracy; strong robustness; a systematic perspective; fast response speed; a simple and easy working principle; low difficulty in work deployment and strong applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A schematic diagram of a fast multi-scale through-wall radar human behavior recognition method and device provided by the present invention;

[0038] Figure 2 A schematic diagram of a fast multi-scale through-wall radar human behavior recognition method and device framework provided by the present invention;

[0039] Figure 3 A schematic diagram of the working principle of the wall-penetrating radar system provided by the present invention;

[0040] Figure 4 A schematic diagram of the structure of a clutter suppression module based on adaptive background subtraction provided by the present invention;

[0041] Figure 5 The present invention provides an adaptive parameter estimation algorithm based on minimum mean square error estimation;

[0042] Figure 6 A schematic diagram of the structure of the feature extraction module provided by the present invention;

[0043] Figure 7 A schematic diagram of the classification module structure provided by the present invention;

[0044] Figure 8 This is a fast multi-scale through-wall radar human behavior recognition method and device recognition verification accuracy curve provided by the present invention.

[0045] Figure 9 This is a performance comparison of the network built under different attention modules provided by the present invention.

[0046] Figure 10 This is a performance comparison of the networks built under different target detection modules provided by the present invention.

[0047] Figure 11 This is a performance comparison of the network built under different classification modules provided by the present invention. DETAILED DESCRIPTION

[0048] The present invention is described in detail below with reference to the accompanying drawings and embodiments.

[0049] This paper proposes a method for human motion recognition based on ultra-wideband through-wall radar. This method, based on frequency-stepped radar data, combines a kernel-distance-based super-resolution Doppler estimation model, a lightweight non-anchor single-stage target detection neural network (DET), and a lightweight classification neural network (LNN) to construct a lightweight deep neural network capable of mining multi-scale feature information. This method simultaneously extracts the translational motion information and the generated μ-D features of a human body in a complex through-wall environment, and combines these two scales for judgment and recognition, ultimately providing through-wall radar human motion recognition results. Experiments have demonstrated that this method achieves high recognition accuracy. Furthermore, because its submodules are constructed based on lightweight models, the overall computational cost of the algorithm is significantly reduced, providing a fast and real-time deployment solution for through-wall radar human behavior recognition systems.

[0050] like Figure 1-2 A fast multi-scale through-wall radar human behavior recognition method and device shown includes:

[0051] Step S1: Image processing module. The through-wall radar system obtains a raw image of the human motion echo and preprocesses it. This preprocessing includes data enhancement, clutter suppression, and noise suppression to produce a processed image. After signal preprocessing, the through-wall radar system generates frequency-stepped radar data, including translational information and micro-Doppler characteristics of human motion behind the wall. This data is used for feature extraction in steps S2 and S3.

[0052] Step S2: Human body translation information feature extraction module. The human body translation information feature extraction module is implemented based on the target detection neural network. Specifically, the present invention proposes an adaptive sampling method (Adaptive Training Sample Se-lection, ATSS) based on the statistical characteristics of the target, which can improve the performance of the frameless detection method. The prior frame refers to the training sample itself. The method consists of four parts: a backbone convolutional neural network (CNN) for extracting features, a feature pyramid network (FPN) used in the network neck, an efficient attention module embedded in the connection of each part, and a detection head using a classification branch and a regression branch. First, the frequency-stepped radar data is input into the backbone CNN for extracting features. After four layers of convolution, pooling, inter-layer normalization and high-level attention weighting, a frequency-stepped radar data heat map containing some wall clutter and human motion Doppler features is obtained. This heatmap is fed into a four-layer FPN, where each layer is half the scale of the previous layer. It is further weighted using an efficient attention module, and the detection head then maps the image into two vectors containing classification and regression box information. Finally, the center point of each regression box is determined, and a curve fitting method based on Lagrange interpolation is used to annotate the frequency-stepped radar data image. The resulting annotated image represents the translational motion information of the human body.

[0053] Step S3: Micro-Doppler feature extraction module. The human body micro-Doppler feature extraction module is implemented based on the Doppler velocity estimation method improved by Gaussian kernel distance. Specifically, the module inputs the original image of the frequency-stepped radar data generated by step 1 and the frequency-stepped radar data heat map generated by step 2, uses the traditional method to perform DTOF scanning, and uses the presence of suspicious targets based on the echo amplitude to achieve parameter estimation, and then focuses on the suspicious object to analyze the target characteristics using the micro-Doppler time-frequency analysis method. In the parameter estimation process, each variable is introduced into the Gaussian kernel probability function. For the human body, the main micro-Doppler component comes from the periodic movement of the hands and feet, which usually has a certain range of movement speed. Therefore, the pixel points whose estimation results fall within this range are marked and output. The output point trace image is the micro-Doppler feature of human body movement.

[0054] Step S4: The classification module uses image stitching to generate a feature dataset based on the human translational motion information and micro-Doppler feature maps generated in steps S2 and S3. Based on this feature dataset, a lightweight convolutional neural network is used for decision recognition, ultimately providing the through-wall radar human behavior recognition results.

[0055] Preferably, the behavioral states of the human body behind the wall are divided into seven categories, including scenes where the human body moves back and forth parallel to the wall, scenes where the human body moves back and forth diagonally behind the wall, scenes where the human body rotates behind the wall, scenes where the human body squats behind the wall, scenes where the human body steps in place behind the wall, scenes where the human body stands still behind the wall, and empty scenes.

[0056] Preferably, if Figure 3 As shown in the figure, the UWB radar antenna and host computer are located on one side of the wall, and the experimenter moves on the other side of the wall. The frequency step waveform is used as the transmission signal, and its representation method is:

[0057]

[0058] Where T represents the duration of each frequency point, K is the total number of frequency points, f0 is the starting frequency, and Δf is the step size. The function rect represents the gate function and is defined as follows:

[0059]

[0060] The human body is modeled using a node model. Assuming that the human target can be divided into P strong scattering points, the received echo can be expressed as:

[0061]

[0062] Among them, a p is the echo intensity coefficient of the pth scattering point, τ p represents the echo delay of the pth scattering point, S W (t) and S N (t) represents the wall echo and noise respectively. By using the time delay characteristics of the echo and summing the point target, we can obtain the time-delay focusing imaging result of the echo.

[0063] The received radar echo is first mixed and low-pass filtered to achieve coherent demodulation. The radar echo after low-pass filtering is sampled at intervals T. Assume that there are M cycles of echo data. After the above signal processing steps, we can reconstruct the echo data into an M×K dimensional matrix S r , then S r An element S r (m,k) can be expressed as:

[0064]

[0065] Among them, S W and S N Represents the discrete reconstruction matrix of wall echo and noise, m represents the index of slow time dimension, and k represents the index of fast time dimension. rBy inverse fast Fourier transform (IFFT) in fast time dimension, we can get the range profile matrix φ r (m,τ). Specifically, one of its elements φ r (m,τ) can be expressed as:

[0066]

[0067] Preferably, if Figure 4 As shown, the through-wall radar clutter suppression method based on adaptive background subtraction processing is a weighted summation of the inputs under the current background, that is:

[0068] B(m,n)=w T (n)φ r (m,n) (6)

[0069] Where w(n)=[w0(n),w1(n),…,w L (n)] T is the weighted coefficient vector, B(m,n) is the background, φ r (m,n)=[φ r (m,n),φ r (m,n-1),…,φ r (m,nL)] T , L is the window length. Figure 4 , the weighted coefficient vector can be obtained through the LMS algorithm:

[0070]

[0071] Where μ represents the adaptive gain coefficient, which usually satisfies: 0<μ<[(L+1)P in ] -1 , P in is the input power. Since it is difficult to obtain the error e(m,n) in practical applications, the error value is estimated by the change between the current background and the previous background. From this, we can get:

[0072]

[0073] Formulas (6), (7) and (8) give the background estimate, so the target echo signal with the background eliminated can be expressed as:

[0074]

[0075] Among them, w * (n)=[1-w0(n),-w1(n),…,-w L (n)] T To load the time domain echo φ after background removal for estimation r The weight on (m,n), φr (m,n) is the time domain echo component used for background estimation when the fast-slow time points are n and m respectively, κ(m,n) is the complete time domain echo when the fast-slow time points are n and m respectively, B(m,n) is the background estimation result when the fast-slow time points are n and m respectively; w T (n)=[w0(n),w1(n),…,w L (n)] T is the time domain echo component φ loaded into the background estimation r The weights on (m,n), where T is the matrix transpose operator.

[0076] If the input vectors are independent, then Figure 5 As shown in Figure 2, the expected value of the weighted vector obtained by the LMS method will be optimal. To simplify the problem, it is assumed that the input vectors are always independent in the experiment. After extracting the envelope features, the results are displayed as a distance-time graph.

[0077] Preferably, the feature extraction module is composed of a target detection module and a micro-Doppler annotation module. Figure 6 As shown in the figure, in the object detection stage, this paper proposes an adaptive training sample selection (ATSS) method based on the statistical characteristics of the target, which can improve the performance of frameless detection methods. The prior frames mentioned below refer to the training samples themselves. The object detection module consists of four parts: a backbone convolutional neural network (CNN) for feature extraction, a feature pyramid network (FPN) at the neck of the network, an efficient attention module embedded at the junctions of each component, and a detection head using a classification branch and a regression branch.

[0078] The specific calculation process of the efficient attention module is as follows: first, the input feature map is subjected to a global average pooling operation; second, a 1-dimensional convolution operation with a convolution kernel size of is performed, and the weights of each channel are obtained through a Sigmoid activation function; finally, the weights are multiplied by the corresponding elements of the original input feature map to obtain the final output feature map. Represents the size of the adaptively learned kernel determined by the channel dimension C.

[0079] This method performs bounding box regression by using the Generalized Focal Loss loss function. The loss function is defined as:

[0080]

[0081] The predicted probabilities of the assumed values u1 and u2 are and The final prediction result is The label is y, satisfying u1≤u≤u2, |uu′|β is the scaling factor, and β is the scaling factor hyperparameter that controls the speed of weight reduction. The target detection module detects the peak of micro-Doppler oscillation in the imaging and defines the peak center point as Where i = 0, 1, ..., I-1; i′ = 0, 1, ..., I′-1 corresponds to the ordinal number of the total I peaks and I′ troughs target boxes, t i-1 , t i , t′ i′-1 , t′ i′ They are the minimum slow time and the maximum slow time area of the corresponding target box, r i-1 、r i , r′ i′-1 , r′ i′ Draw the tracking curve for the minimum distance and maximum distance areas corresponding to the i-th target box respectively:

[0082]

[0083]

[0084] where r Peak (t) and r Trough (t) represent the path curves of micro-Doppler peak and trough on the imaging, respectively, and r is defined Track (t) is the path estimation of the target center of mass on the image:

[0085]

[0086] If the target detection only contains information about one of the peaks and troughs, it is directly output as the path estimate. This module can extract macro-scale distance-time path information.

[0087] The micro-Doppler annotation module uses the improved Doppler velocity estimation of Gaussian kernel distance to more accurately mine the target micro-motion information. Specifically, for the input Z′(m,n) of the micro-Doppler annotation link, solve:

[0088]

[0089] in and is a constant, given by considering the signal-to-noise ratio and the average Doppler velocity variation, q i are the image points after attention weighting. d,i,j Defined as represents the sampled v d,i,jFor the human body, the main micro-Doppler component comes from the periodic motion of the hands and feet. Therefore, for the estimated results falling within q in the range i This module can be used to extract information about micro-scale Doppler estimation.

[0090] Preferably, if Figure 7 As shown, the classification module is built using a lightweight convolutional neural network. The information obtained by the above modules, namely the path estimate and micro-Doppler estimate of the target centroid on the image, is superimposed and fused, and then input into the classification module to produce a decision. Specifically, this module is composed of several stacked ignition modules, and is also improved by the efficient attention module. Pooling is used between ignition modules to achieve dimensionality reduction, and a residual connection scheme is introduced, thus forming the required classification and decision module. The last layer of the classification module uses a fully connected Softmax activation function, with an output dimension of 7, corresponding to the decision results for the classification and recognition of seven types of human behaviors. After the accumulated image data of each frame in the dataset is input, it is processed separately in the target detection chain and the annotation chain. The peak and trough detection results are fused to obtain the macroscopic path, which is then aggregated with the micro-Doppler annotation output to obtain a feature map. Finally, the classification module makes a decision based on the feature map. Each module of the network is built using a lightweight algorithm, and the number of parameters is far smaller than that of algorithms built using other networks with the same function.

[0091] This invention addresses the technical challenges of through-the-wall radar (TWR) dynamic target recognition, a key area of research in the field of portable ultra-wideband (UWB) radar for both military and civilian applications. It focuses on physical modeling and computer vision (CV) theory, organically combining the strengths of both. This invention proposes a fast, multi-scale through-the-wall radar human behavior recognition method and device. This method addresses the significant distortion of UWB through-the-wall radar echo signals due to wall effects, such as attenuation, refraction, and multipath. This significantly reduces the accuracy of human behavior recognition, significantly increases the computational cost of available models, and makes system deployment extremely challenging. Specifically, the method and device acquire a raw image of the through-the-wall radar human behavior echo and preprocess it. This preprocessing includes data enhancement, clutter, and noise suppression to produce a processed image. The feature extraction module, based on the frequency-stepped radar data, combines a kernel-distance-based super-resolution Doppler estimation model, a lightweight non-anchor single-stage target detection neural network (DET), and a lightweight classification neural network (LNN) to construct a lightweight deep neural network capable of mining multi-scale feature information. This method generates a feature dataset by simultaneously extracting the translational motion information and the generated μ-D features of a human body in a complex through-wall environment and combining the information from both scales. The classification module, based on this generated feature dataset, uses a lightweight convolutional neural network to perform judgment and recognition, ultimately delivering through-wall radar human behavior recognition results. Experiments have demonstrated that this method and device can achieve 93.98% recognition accuracy when all modules are operating simultaneously. Even when a module fails, the remaining modules can still maintain a high recognition accuracy of over 90%. Replacing the classification module with a similar lightweight network significantly reduces the overall recognition accuracy, training time, inference time, and parameter count of the method. The network inputs are trained and validated using the feature maps output by the target detection module and micro-Doppler annotation module proposed in this patent. Comparative results demonstrate that the proposed method significantly improves algorithm execution speed while maintaining high recognition accuracy, reducing computational and memory costs, significantly enhancing the robustness and practicality of the invention. Compared to traditional recognition methods, this invention is more applicable to complex scenarios such as through-wall military identification, through-wall civilian human identification, through-wall vital signs detection, and disaster relief.

[0092] The present invention processes, enhances, and identifies motion features at different scales. Multiple modules of the recognition system operate simultaneously, utilizing a clutter suppression method based on adaptive background subtraction, a feature extraction module, and a classification module to achieve decision-making and identification. The clutter suppression step includes adaptive background subtraction for wall clutter and noise subspace separation and suppression, highlighting the characteristics of human motion targets. This makes the present invention highly applicable to complex real-world scenarios.

[0093] Front-end data acquisition is performed using ultra-wideband radar. After data acquisition, signal processing and imaging are performed, and the imaging information matrix undergoes clutter suppression using adaptive background subtraction. Preprocessing purifies the data's useful features and provides a visual and intuitive way to interpret and understand the present invention. This helps improve system reliability in practical applications.

[0094] In order to better illustrate the purpose and advantages of the present invention, the invention is further described below with reference to the accompanying drawings, experiments and data analysis.

[0095] In the simulation experiment, the present invention divides the dynamic model of the human body behind the wall into three categories: no one behind the wall, stationary, and moving. The specific subdivisions include no one behind the wall, a person behind the wall standing toward the wall, a person behind the wall standing toward the wall, a person behind the wall standing parallel to the wall, a person behind the wall moving toward the wall, a person behind the wall moving behind the wall, and a person behind the wall in seven different states.

[0096] The verification process of the present invention includes the entire signal path from the classification module to the signal and data processing module. This step can improve the interpretability of the present invention during the training process and ensure the advantage of the present invention tending to convergence and solvability.

[0097] The verification process of the present invention trains and verifies the proposed method. The final convergence accuracy of the training process is 100%, and the convergence accuracy of the verification process is 93.98%. Figure 8 This is a graph showing the recognition verification accuracy of the method of the present invention. With the same preset parameters and hardware environment, a complete 80-round training takes only 8 minutes and 21 seconds, and inference of one frame of imaging data accumulated over 4 seconds takes only 0.9 seconds.

[0098] like Figure 9 As shown in the figure, the verification process of the present invention demonstrates the impact of the efficient attention module on the recognition accuracy of the back-end classifier, keeping the object detection and classification modules unchanged. By introducing the efficient attention module, the annotation method's sensitivity to image signal amplitude can be significantly improved while slightly increasing the overall number of parameters. This allows the Doppler estimation to focus on the useful motion signal and ignore interference from multipath effects and wall clutter. The feature map after the feature map is used before network training.

[0099] like Figure 10As shown, the verification process of the present invention provides the average precision (AP), training time, inference time, floating point operations per second (BFLOP / s), and frames per second (FPS) of the overall method when replacing the target detection module with other similar single-stage networks. In the field of through-wall radar human behavior recognition, using a lightweight detection network based on a feature pyramid does not significantly reduce AP compared to other deeper and wider detection methods, but it can improve the regression speed of the prediction box, thereby improving the overall performance of the network.

[0100] like Figure 11 As shown, the verification process of this invention shows the overall recognition accuracy, training time, inference time, and parameter count of the method when replacing the classification module with another similar lightweight network. The network inputs were trained and verified using the feature maps output by the proposed object detection module and micro-Doppler annotation module. Comparative results demonstrate that the proposed method significantly improves algorithm execution speed while maintaining a high recognition accuracy of 93.98%, while reducing computational and memory costs.

[0101] The present invention also provides a through-wall radar human behavior recognition device based on multi-link information decision-making, the device comprising:

[0102] Image processing module: obtains the original image of the human behavior echo of the through-wall radar and preprocesses the original image; the preprocessing includes data enhancement, clutter and noise suppression, and obtains the processed image;

[0103] Feature Extraction Module: This module combines a kernel-distance-based super-resolution Doppler estimation model, a lightweight non-anchor single-stage target detection neural network (DET), and a lightweight classification neural network (LNN) to construct a lightweight deep neural network capable of mining multi-scale feature information. This module simultaneously extracts the translational motion information of the human body in complex wall-penetrating environments and the resulting μ-D features, then combines these two scales to generate a feature dataset.

[0104] Classification module: It is configured as a lightweight convolutional neural network, which classifies the obtained feature data set to obtain the final recognition result.

[0105] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A fast multi-scale through-wall radar human behavior recognition method, characterized by: The following steps are involved: Step 1: Acquire the original image of the human behavior echo of the through-wall radar, and preprocess the original image to obtain frequency-stepped radar data, including the translational information and micro-Doppler characteristics of the human movement behind the wall; the preprocessing includes data enhancement, clutter and noise suppression; Step 2: Input the frequency-stepped radar data into the backbone CNN of the human translation information feature extraction module. After passing through four layers of convolution, pooling, inter-layer normalization, and efficient attention weighting, a frequency-stepped radar data heat map containing some wall clutter and human motion Doppler features is obtained. The heat map is input into a four-layer FPN, where the scale of each layer is 1 / 2 of the previous layer. The heat map is further weighted using the efficient attention module, and the image is mapped into two vectors containing classification and regression box information using the detection head. Finally, the center point of each regression box is solved, and a curve fitting method based on Lagrange interpolation is used to perform curve annotation on the frequency-stepped radar data image. The resulting annotated image is the human translation information. The human body translation information feature extraction module consists of four parts: a backbone CNN for feature extraction, a network neck using FPN, an efficient attention module embedded in the connection between each part, and a detection head using a classification branch and a regression branch; Step 3: Perform a DTOF scan on the frequency-stepped radar data generated in step 1 and the frequency-stepped radar data heat map generated in step 2, and use the echo amplitude to determine whether a suspicious target exists to perform parameter estimation. Then, focus on the suspicious object and use the micro-Doppler time-frequency analysis method to analyze the target characteristics. During the parameter estimation process, a Gaussian kernel probability function is introduced into each variable. Pixels whose estimated results fall within the human motion speed range are annotated and output. The output dot trace image is the micro-Doppler signature of human motion. Step 4: For the human body translation information and micro-Doppler feature maps generated in steps 2 and 3, a feature dataset is generated using the image stitching method. Based on the generated feature dataset, a lightweight convolutional neural network is used for judgment and recognition, and finally the through-wall radar human behavior recognition result is given.

2. The fast multi-scale through-wall radar human behavior recognition method according to claim 1, characterized in that: In step 1, the human body is modeled using a node model. Assuming that the human target is divided into P strong scattering points, the received echo is expressed as: Among them, S W (t) and S N (t) represents the wall echo and noise respectively, and t is the propagation time of the electromagnetic wave emitted by the radar in space. By using the time delay characteristics of the echo and summing the point targets, we can obtain the delayed focusing imaging result of the echo. Represents the total echo signal of the human target. The signal is split into p scattered components corresponding to p echoes according to the round-trip delay of different nodes of the human body. The amplitude gain of each echo component is a p , the time domain response is s, and the delay is τ p .

3. The fast multi-scale through-wall radar human behavior recognition method according to claim 2, characterized in that: The target echo signal Z(m,n) with background eliminated and the number of fast and slow time points n and m respectively is expressed as: Z(m,n)=κ(m,n)-B(m,n) =κ(m,n)-w T (n)φ r (m,n)# =w *T (n)φ r (m,n) Among them, w *T (n)=[1-w0(n),-w1(n),...,-w L (n)] T To load the time domain echo φ after background removal for estimation r The weight on (m, n), φ r (m, n) is the time domain echo component used for background estimation when the fast-slow time points are n and m respectively, κ(m, n) is the complete time domain echo when the fast-slow time points are n and m respectively, B(m, n) is the background estimation result when the fast-slow time points are n and m respectively; w T (n)=[w0(n),w1(n),...,w L (n)] T is the time domain echo component φ loaded into the background estimation r The weights on (m, n), where T is the matrix transpose operator.

4. The fast multi-scale through-wall radar human behavior recognition method according to claim 1, characterized in that: The specific calculation process of the efficient attention module is as follows: first, the input feature map is subjected to a global average pooling operation; second, a one-dimensional convolution operation with a convolution kernel size of 3 is performed, and the weights of each channel are obtained through a Sigmoid activation function; finally, the weights are multiplied by the corresponding elements of the original input feature map to obtain the final output feature map.

5. The fast multi-scale through-wall radar human behavior recognition method according to claim 1, characterized in that: The behavioral states of the human body behind the wall are divided into seven categories, including the scene of the human body moving back and forth parallel to the wall, the scene of the human body moving back and forth diagonally behind the wall, the scene of the human body rotating behind the wall, the scene of the human body squatting behind the wall, the scene of the human body stepping on the spot behind the wall, the scene of the human body standing still behind the wall and the empty scene.

Citation Information

Patent Citations

  • Through-the-wall radar human body behavior identification method and device based on multi-link information decision

    CN115184890A