Lightweight Ultra-Wideband Radar Gesture Recognition Method Based on Attention Mechanism

By adopting a lightweight ultra-wideband radar gesture recognition method based on attention mechanism in radar gesture recognition, combined with sliding window Doppler algorithm and feature extraction unit, the problems of low signal noise and large parameters in the prior art are solved, and higher accuracy and robustness are achieved.

CN116299427BActive Publication Date: 2025-06-17AIR FORCE UNIV PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211524482.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-30
Publication Date
2025-06-17
Estimated Expiration
2042-11-30

AI Technical Summary

Technical Problem

The existing radar gesture recognition method has low signal-to-noise signal and low information in the actual application environment, and is not suitable as the only input of the classifier. The calculation of the arrival angle increases the delay. The cascaded convolutional neural network of the memory network requires more network parameters, making it difficult for the network to converge.

Method used

The lightweight ultra-wideband radar gesture recognition method based on attention mechanism is adopted, Doppler processing is improved through the sliding window Doppler algorithm, scale telescopic unit, basic computing unit and attention unit are designed, serial pulse ultra-wideband radar gesture recognition model is established, network parameters are controlled and network convergence is accelerated.

Benefits of technology

It improves the generalization and robustness of gesture recognition methods, improves accuracy and real-timeness, reduces the amount of network parameters and accelerates network convergence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116299427B_ABST
    Figure CN116299427B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight ultra-wideband radar gesture recognition method based on an attention mechanism. Taking the micro-motion gesture echo spectrum as the input, a sliding window Doppler algorithm is proposed to generate the range-Doppler sequence of the gesture action process. Aiming at the characteristics of the near-range gesture echo Doppler spectrum with unclear edge and texture features and accompanied by target ghosts, three computing units, namely, a scale scaling unit, a basic computing unit, and an attention unit, are designed according to the lightweight principle. By jointly using channel attention and spatial attention to improve the sensitivity of the network to the target area, a serial lightweight gesture recognition model is constructed. Through experimental verification, compared with the existing gesture recognition algorithms, the lightweight gesture recognition model proposed in the present invention has better generalization and robustness, and also has advantages in accuracy and real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gesture recognition, and in particular to a lightweight ultra-wideband radar gesture recognition method based on an attention mechanism. Background Art

[0002] In recent years, with the development of intelligent interaction, gesture recognition has more application scenarios. Gesture recognition based on radar sensors has greater advantages in terms of privacy and light insensitivity. Currently, researchers at home and abroad have proposed many solutions in the field of radar gesture recognition. Kim et al. used a convolutional neural network to learn the characteristics of the time-domain echo waveform of radar gestures, thereby realizing gesture classification. Youngwook et al. extracted the micro-Doppler information in the radar gesture echo spectrogram through a DCNN network to realize gesture recognition. Skaria et al. formed a range-Doppler three-dimensional tensor using the I / Q components of the gesture echo signal. Each tensor consists of a series of range-Doppler time snapshots, and a 2D-CNN network and an LSTM network were used to extract action features for classification. Sruthy Skaria et al. used two receiving antennas that can generate orthogonal components, mapped the echo signal to three-dimensional space, and calculated the angle-of-arrival matrix as action information, and trained a DCNN to extract the features of the above information to complete the classification.

[0003] In the laboratory environment, the above methods have achieved excellent classification accuracy rates. However, in the actual application environment, the echo signal has a low signal-to-noise ratio and little information, and is not suitable as the only input for the classifier; the calculation of the angle of arrival places higher requirements on the number of receiving antennas and the design of signal algorithms, increasing the time delay; cascading a long short-term memory network and a convolutional neural network requires more network parameters for the network to converge. Summary of the Invention

[0004] In view of the above problems, the present invention aims to provide a lightweight ultra-wideband radar gesture recognition method based on an attention mechanism, which has better generalization and robustness, and also has advantages in terms of accuracy and real-time performance.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A lightweight ultra-wideband radar gesture recognition method based on an attention mechanism, characterized by including the following steps,

[0007] S1: Collect and preprocess the gesture echo signal of the ultra-wideband radar to obtain the range-Doppler map of the gesture echo signal;

[0008] S2: Use the method in step S1 to construct a gesture data set;

[0009] S3: Establish a lightweight micro gesture recognition model according to the characteristics of the gesture data set in step S2;

[0010] S4: Use the lightweight micro gesture recognition model to recognize radar gestures.

[0011] Furthermore, the specific operations of step S1 include the following steps.

[0012] S11: Conduct equivalent time sampling analysis on the gesture echo signal of the ultra-wideband radar;

[0013] S12: Convert the equivalently sampled signal through A / D conversion to obtain a digital signal;

[0014] S13: Perform clutter suppression and sliding window Doppler processing on the digital signal of the gesture echo to obtain the range-Doppler map of the gesture echo signal.

[0015] Furthermore, the specific operations of step S11 include the following steps.

[0016] S111: Based on the "stop-hop" principle, assume that there is relative motion between the ultra-wideband radar and the target. Then the distance R(t) between the radar and the target is

[0017] R(t) = R0 + v r t

[0018] where R0 is the distance at zero time, and v r = dR(t) / dt is the relative radial motion speed of the target and the radar;

[0019] Expand the moving target equation R(t) as

[0020]

[0021] In the formula, represents the time delay generated by the radar with respect to the initial position, c is the speed of light, c = 3×10 8 m / s;

[0022] When the moving distance of the target is much smaller than R0,

[0023]

[0024] S112: Assume that the radar echo signal, that is, the gesture echo signal, is S r (t), and the radar transmitted signal is S t (t), then

[0025]

[0026] where α is the attenuation factor of the echo, and γ is the stretching factor of the echo signal on the time scale,

[0027] The change in the single-pulse pulse width caused is

[0028]

[0029] In the formula, τ is the single-pulse pulse width;

[0030] S113: When the transmitted signal is a pulse signal with a period equal to T P the

[0031]

[0032] received signal is

[0033]

[0034] S114: Since

[0035]

[0036] therefore, the period of the echo signal changes to

[0037] T P ′ = T p / γ

[0038] and the change in the pulse interval ΔT P caused is

[0039]

[0040] Furthermore, the specific operations of clutter suppression and sliding window Doppler processing on the digital signal of the gesture echo in step S13 include the following steps.

[0041] S131: Assume that the radar transmitted signal is the echo signal at any moment is In the formula, f r is the echo signal frequency;

[0042] S132: Assume that the frequency of the analog signal after equivalent sampling is f′, which is equivalent to the frequency of the gesture echo signal, and regard the discrete digital signal after A / D conversion as a continuous signal. Then the digital signal after A / D conversion of the first analog signal after equivalent sampling is

[0043]

[0044] S133: After A / D sampling with a pulse interval of ΔT, the moving distance of the hand is ΔR = vΔT

[0045] The corresponding time interval is

[0046]

[0047] Then, the digital signal after the second equivalent sampled analog signal is converted by A / D

[0048]

[0049] The phase change relative to the first digital echo signal is

[0050]

[0051] S134: Further analysis of the digital signal after the second equivalent sampled analog signal is converted by A / D gives

[0052]

[0053] Wherein, is defined as the single-step frequency difference of the sliding window Doppler algorithm, Initial phase;

[0054] S135: From the digital signal after the first equivalent sampled analog signal is converted by A / D and the digital signal after the second equivalent sampled analog signal is converted by A / D, it can be inferred that the discrete signal after sampling is

[0055]

[0056]

[0057] In the formula, S r1 (k) and S r2 (k) are the discrete sampling sequences of S r1 (t) and S r2 (t) respectively;

[0058] S136: Subtract adjacent two signals once to complete the elimination of zero-frequency or low-speed clutter, that is

[0059] S rfj (k) = S r2 (k) - S r1 (k)

[0060] In the formula, S rf1 (k) is the result after subtracting the echo sequences of adjacent sampling points, and is also the constituent sequence of the data matrix;

[0061] S137: Select N consecutive pulses through a sliding window, repeat the above steps S131 - S136 for Doppler processing, the sliding step size is m, m is less than or equal to N, perform a Fourier transform on the data matrix every m sampling periods, and finally obtain the range-Doppler map of the gesture echo signal.

[0062] Further, the gesture data set described in step S2 includes six micro gestures, namely, a finger snap, a left palm wave, a forward rubbing of the thumb and index finger, a backward rubbing of the thumb and index finger, a finger undulation, and a gathering of the five fingers from a flat state to the fingertips.

[0063] Further, the specific operations of step S3 include the following steps

[0064] S301: Construct three operation units for feature extraction from the range-Doppler map of the gesture echo signal. The three operation units are a scale stretching unit, a basic calculation unit, and an attention unit;

[0065] S302: Based on the three operation units constructed in step S301, establish a serial pulse ultra-wideband radar gesture recognition model.

[0066] Further, the operations of the scale stretching unit for feature extraction include the following steps

[0067] Step a1: Convolve the input of the scale stretching unit through two convolutional branches. The convolutional branches sequentially include 1×1Conv, 3×3DW, and 1×1Conv. The 1×1Conv in the left branch contains batch normalization BN of the data and a non-linear activation function ReLu;

[0068] Step a2: Concatenate the outputs of the two convolutional branches along the channels to complete channel expansion;

[0069] Step a3: Shuffle the expanded channels as the output of the feature extraction of the scale stretching unit.

[0070] Further, the operations of the basic calculation unit for feature extraction include the following steps

[0071] Step b1: Divide the input of the basic calculation unit along the channel dimension. One branch performs feature calculation through pointwise convolution and then concatenates the features with the other branch;

[0072] Step b2: Perform inter-channel information exchange on the concatenated features through channel shuffling;

[0073] Step b3: Repeat step b1 and step b2;

[0074] Step b4: Add the feature calculation result of step b3 to the input through residual connection as the output of the feature extraction of the basic calculation unit.

[0075] Further, the operations of the attention unit for feature extraction include the following steps

[0076] Step c1: Adjust the global attention dimension of the input of the attention unit by controlling the execution dimension of global pooling, multiply the obtained attention coefficients along the corresponding dimension to the feature map, and use the feature map as the input for obtaining the attention in the next dimension, so as to obtain a feature map with joint attention;

[0077] Step c2: Add the feature with joint attention obtained through residual connection to the original input, and output it after passing through the activation function as the output of the feature extraction of the attention unit.

[0078] Furthermore, the serial pulse ultra-wideband radar gesture recognition model includes three classification network stages. Each classification network stage sequentially includes a scale expansion unit, a basic calculation unit, and an attention unit. The number of repetitions of the loop structure in the basic calculation units of the three stages in the classification network is different.

[0079] The beneficial effects of the present invention are as follows:

[0080] 1. When processing the gesture echo signal, the gesture recognition method in the present invention improves the Doppler processing, proposes a sliding window Doppler algorithm, makes the coherent processing time adjustable, and can obtain more range Doppler maps (RDM) within the same time.

[0081] 2. The gesture recognition method in the present invention designs a scale expansion unit and a basic calculation unit, uses depthwise separable convolution, channel shuffle, and residual connection to calculate and design the operation unit for feature extraction. In addition, an attention unit is designed to obtain joint attention coefficients by fusing channel attention and spatial attention, strengthen the weight of the target area, and reduce the influence of target ghosts in the range Doppler map on the recognition result.

[0082] 3. The present invention establishes a serial pulse ultra-wideband radar gesture recognition model based on the scale expansion unit, the basic calculation unit, and the attention unit. By controlling the execution position and number of the three feature calculation units, while controlling the network parameter quantity, the network convergence is accelerated. It is verified that the gesture recognition method in the present invention has better generalization and robustness, and also has advantages in accuracy and real-time performance. Description of the Drawings

[0083] Figure 1 It is a schematic diagram of equivalent sampling timing in the present invention.

[0084] Figure 2 It is a schematic diagram of the sampling receiving module and the echo signal in the present invention.

[0085] Figure 3 It is a schematic diagram of traditional Doppler processing.

[0086] Figure 4 This is the timing diagram of sliding window Doppler processing in the present invention.

[0087] Figure 5 This is the schematic diagram of performing Fourier transform on the data matrix in the present invention.

[0088] Figure 6 This is the schematic diagram of gesture acquisition in the present invention.

[0089] Figure 7 This is the structural diagram of the scale expansion and contraction unit in the present invention.

[0090] Figure 8 This is the structural diagram of the basic operation unit in the present invention.

[0091] Figure 9 This is the structural diagram of the attention unit in the present invention.

[0092] Figure 10 This is the composition diagram of the classification network stage in the present invention.

[0093] Figure 11 This is the schematic diagram of the structure of the serial pulse ultra-wideband radar gesture recognition model in the present invention.

[0094] Figure 12 This is the visual structural diagram of the scale expansion and contraction unit in the present invention.

[0095] Figure 13 This is the visual structural diagram of the basic operation unit in the present invention.

[0096] Figure 14 This is the visual structural diagram of the attention unit in the present invention.

[0097] Figure 15 This is the accuracy curves of three networks in the ablation experiment of the attention unit in the present invention.

[0098] Figure 16 This is the accuracy curves of five networks in the ablation experiment of the basic operation unit in the present invention.

[0099] Figure 17 This is the accuracy curves of six networks in the model comparison in the present invention.

[0100] Figure 18 This is the average recognition accuracy of 7 participants in the present invention.

[0101] Figure 19 This is the accuracy of 10-fold cross-validation of 7 participants in the present invention. Detailed implementation manners

[0102] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0103] Example 1:

[0104] A lightweight ultra-wideband radar gesture recognition method based on an attention mechanism, comprising the following steps:

[0105] S1: Collect and preprocess the gesture echo signal of the ultra-wideband radar to obtain the range-Doppler map of the gesture echo signal;

[0106] S2: Use the method in step S1 to construct a gesture data set;

[0107] S3: Use the gesture data set in step S2 to establish a lightweight micro-gesture recognition model;

[0108] S4: Use the lightweight micro-gesture recognition model to recognize radar gestures.

[0109] Specifically, the specific operations of step S1 include the following steps:

[0110] S11: Perform equivalent time sampling analysis on the gesture echo signal of the ultra-wideband radar;

[0111] In the present invention, the signal transmitted by the ultra-wideband radar platform adopts a nanosecond-level coherent pulse train signal, where the carrier frequency of the pulse signal is 3.4 GHz, the pulse repetition frequency is 10 MHz, the pulse width is 1 ns, the bandwidth is 1 GHz, and the antenna beam width is 10° to 20°. According to the Nyquist sampling theorem, in order to ensure that the signal obtained by real-time sampling is not distorted, the repetition frequency of the sampling pulse must be greater than or equal to 2 times the highest frequency component of the signal. When the carrier frequency is in the GHz range, it is relatively difficult to achieve and the hardware cost is relatively high. Therefore, the present invention adopts the transform sampling method, and its sampling timing diagram is shown in the appendix Figure 1 as shown.

[0112] In the appendix Figure 1 , δt is the time resolution, and the corresponding range resolution is

[0113]

[0114] The measured time interval difference is

[0115] t D =(N - 1)δt (2)

[0116] where N is the number of sampling times, c is the speed of light, c = 3×10 8 m / s;

[0117] This is equivalent to magnifying the received signal on the time axis, where the magnification factor is defined as

[0118]

[0119] According to Equation (3), it can be obtained that within every β periods, the frequency of the analog signal after equivalent sampling is approximately equal to the frequency of the original signal. Therefore, the equivalent sampling signals of β echo pulses are approximated as the echo signal within a certain unit time, and this unit time is much larger than the repetition period of the echo pulses. Since the gesture movement is slow, the movement within 1 ns and 10 μs can be regarded as stationary, where 1 ns is the pulse width and 10 s is the pulse repetition period. Therefore, based on the above analysis, the present invention designs a differential frequency sampling and receiving module to realize the reception and sampling of the hand radio frequency echo signal, which is also called equivalent time sampling, as shown in the appendix Figure 2 shown, in the appendix Figure 2 In it, (a) is the sampling and receiving module, and (b) is the echo pulse.

[0120] Perform envelope calibration sampling analysis on the echo signal, which specifically includes the following steps.

[0121] S111: Based on the "stop-jump" principle, assume that there is relative motion between the ultra-wideband radar and the target. Then the distance R(t) between the radar and the target is

[0122] R(t) = R0 + v r t (4)

[0123] where R0 is the distance at zero moment, and v r = dR(t) / dt is the relative radial motion speed of the target and the radar;

[0124] Expand the moving target equation R(t) as

[0125]

[0126] In the formula,

[0127] When the moving distance of the target is much smaller than R0, the above formula can be approximated as

[0128]

[0129] S112: Assume that the radar echo signal, that is, the gesture echo signal, is S r (t), and the radar transmission signal is S t (t), then

[0130]

[0131] where α is the attenuation factor of the echo, and γ is the stretching factor of the echo signal on the time scale.

[0132]

[0133] Therefore, it can be obtained that the change amount of the single-pulse pulse width caused by it is

[0134]

[0135] Where τ is the single-pulse width;

[0136] S113: When the transmitted signal is a pulse signal with a period equal to T P of,

[0137]

[0138] the received signal is

[0139]

[0140] S114: Since

[0141]

[0142] it can be seen that the period of the echo signal has been changed to

[0143] T P = T P / γ (13)

[0144] and a pulse interval change ΔT P is

[0145]

[0146] Compared with the single-period pulse width change in Equation (9), the change in the repetition period of the pulse train is much more obvious. And T P is larger, the change is more obvious. Therefore, in the pulse ultra-wideband radar system, the speed measurement can be achieved by measuring the change in the pulse repetition period of the echo signal after equivalent sampling.

[0147] The gesture echo is amplified by β times after equivalent sampling. The echo contains the position information and speed information during the gesture movement process, and different gestures can be distinguished accordingly.

[0148] Furthermore, step S12: The signal after equivalent sampling is subjected to A / D conversion to obtain a digital signal;

[0149] Furthermore, S13: The digital signal of the gesture echo is subjected to clutter suppression and sliding window Doppler processing to obtain the range-Doppler map of the gesture echo signal.

[0150] More specifically, the clutter suppression of the echo signal is to subtract the zero-frequency component and the low-frequency component in the digital signal through a first canceller, which can play a more obvious role in suppressing the background clutter, and then perform Doppler processing on the signal after the first cancellation.

[0151] Doppler processing performs spectral analysis on the data at the same range cell in each data matrix, that is, a discrete Fourier transform is performed row by row along the slow-time dimension in the data matrix, and the process is as shown in the appendix Figure 3 as follows. In the classical Doppler processing shown in the appendix Figure 3 as follows, when the number of pulses in the data matrix (the same as the number of sampling times) is N, a Fourier transform can be performed only after accumulating N echo pulses.

[0152] In the present invention, the above classical Doppler processing is improved, and a sliding-window Doppler algorithm is proposed. By using a sliding window to select N consecutive pulses for Doppler processing, the sliding step size is m, where m is less than or equal to N. A Fourier transform can be performed on the data matrix every m sampling periods, thereby shortening the interval time of each frame of range-Doppler map. The sliding-window Doppler processing can obtain at most N times as many range-Doppler maps as the former within the filling time of a data matrix, and its processing timing is as shown in the appendix Figure 4 as follows.

[0153] By changing the selection rule of the coherent processing time, the sliding-window Doppler processing adds an intermediate process of the action, which can ensure that the current matrix contains the information of N - m sampling pulses in the previous data matrix, retains the continuous historical information of the gesture, and is of great significance to the subsequent recognition algorithm.

[0154] The specific operations of the Doppler effect in the sliding-window Doppler algorithm based on equivalent sampling in the present invention include the following steps

[0155] S131: Assume that the radar transmitted signal is and the echo signal at any moment is

[0156]

[0157] where f r is the echo signal frequency;

[0158] S132: Assume that the frequency of the analog signal after equivalent sampling is f′, which is equivalent to the frequency of the gesture echo signal. Regarding the discrete digital signal after A / D conversion as a continuous signal, the digital signal after A / D conversion of the first analog signal after equivalent sampling is

[0159]

[0160] S133: After the A / D sampling pulse interval ΔT, the moving distance of the hand is

[0161] ΔR = vΔT (17)

[0162] The corresponding time interval is

[0163]

[0164] Then, the digital signal obtained after the second equivalent sampled analog signal is A / D converted

[0165]

[0166] The phase change relative to the first digital echo signal is

[0167]

[0168] S134: Further analysis of the digital signal obtained after the second equivalent sampled analog signal is A / D converted yields

[0169]

[0170] wherein, in the above formula

[0171]

[0172] is defined as the single-step frequency difference of the sliding window Doppler algorithm, is the initial phase;

[0173] S135: From the digital signal obtained after the first equivalent sampled analog signal is A / D converted and the digital signal obtained after the second equivalent sampled analog signal is A / D converted, it can be inferred that the discrete signal after sampling is

[0174]

[0175] In the formula, S r1 (k) and S r2 (k) are the discretized sampling sequences of S r1 (t) and S r2 (t) respectively. Therefore, the conclusions drawn in the present invention for S r1 (t) and S r2 (t) are also applicable to S r1 (k) and S r2 (k).

[0176] S136: Subtract adjacent two signals once to complete the elimination of zero-frequency or low-speed clutter, that is

[0177] S rf1 (k) = S r2 (k) - S r1 (k)

[0178] In the formula, S rf1 (k) is the result after subtracting the echo sequences of adjacent sampling points, and is also the constituent sequence of the data matrix;

[0179] S137: Select N consecutive pulses through a sliding window, and repeat the above steps S131 - S136 for Doppler processing. The sliding step size is m, where m is less than or equal to N. Perform a Fourier transform on the data matrix every m sampling periods, and finally obtain the range-Doppler map of the gesture echo signal.

[0180] From the above analysis, it can be seen that there is a phase difference between two adjacent pulses in the sliding window Doppler processing. This is due to the Doppler effect caused by the target movement and can be converted into a Doppler frequency shift, which is the single-step frequency difference. However, this digital signal is obtained after the analog echo signal undergoes equivalent sampling. Therefore, the frequency difference between individual digital echoes in Equation (22) is the Doppler frequency shift reflected based on the change in the pulse repetition period of multiple analog echo signals, rather than the Doppler frequency shift between individual analog echo pulses. Rearrange these sampled signals and then perform a Fourier transform along the range dimension, which can amplify the Doppler effect to the change amount of the echo pulse repetition period. Performing a Fourier transform on it approximately obtains its moving speed. Therefore, good results can be obtained in this paper without pulse accumulation. The principle of performing Doppler processing on the data matrix composed of digital signals is as shown in the appendix Figure 5 as follows.

[0181] Further, step S2: Use the method in step S1 to construct a gesture data set;

[0182] The described gesture data set includes six micro gestures, namely, finger snapping (gesture 1), palm waving to the left (gesture 2), thumb and index finger rubbing forward (gesture 3), thumb and index finger rubbing backward (gesture 4), finger wiggling (gesture 5), and five fingers gathering from flat to fingertips (gesture 6). There are a total of six experimental personnel who complete the above actions within 10 cm - 20 cm from the forward view of the radar antenna. Obtain the range-Doppler maps of the six types of gestures according to the sliding window Doppler algorithm. Each type of gesture contains 600 feature maps, consisting of 100 maps from each participant. The data set has a total of 3600 maps. Divide the training and test data sets in a ratio of 4:1. Each type of gesture has 480 training samples and 120 test samples. The schematic diagrams of each gesture are as shown in the appendix Figure 6 as follows.

[0183] The meaning of each gesture is shown in Table 1 below.

[0184] Table 1 Meanings of Six Gestures

[0185]

[0186] Further, step S3: Use the gesture data set in step S2 to establish a lightweight micro gesture recognition model; its specific operations include the following steps

[0187] S301: Construct three operation units for feature extraction from the range-Doppler map of gesture echo signals;

[0188] There are significant differences between the Doppler range map of near-object echoes and natural images, and lightweight convolutional networks need to be designed in combination with the number of parameters and memory. Based on the above considerations, three operation units, namely, a scale expansion unit, a basic calculation unit, and an attention unit, are proposed in the present invention.

[0189] The structure of the scale expansion unit is as shown in the appendix Figure 7 In the figure, Conv represents a convolution operation, DW represents a depthwise convolution, BN represents batch normalization of data, ReLu represents a non-linear activation function, and Concat represents a channel concatenation operation.

[0190] The operation process of the scale expansion unit is as follows:

[0191] F′ = shuffle{conccat[DS(Conv1 <f>),DS(Conv2<F>)]} (24)

[0192] Among them,

[0193] Conv1(·) = ReLu[BN(1×1Conv<·>)]

[0194] Conv2(·) = BN(1×1Conv<·>)

[0195] DS(·) = 1×1Conv(DW<·>) (25)

[0196] When the input is H×W×C, the number of parameters of the depthwise separable convolution can be calculated as follows:

[0197] Params D.S = K×K×C + 1×1×C×N (26)

[0198] The number of parameters of the conventional convolution is

[0199] Params std = K×K×C×N (27)

[0200] Comparison of the number of parameters of the two convolutions

[0201]

[0202] In the scale expansion unit, size compression is achieved by controlling the stride of the convolution kernel, and channel dimension expansion is achieved by adjusting the number of convolution kernels. The input is input into two branches for calculation respectively, and the outputs of the two branches are concatenated along the channel to complete channel expansion. Finally, channel shuffle is performed to avoid bringing the influence of information occlusion between channels into the next operation unit.

[0203] The structure of the basic calculation unit is as shown in the appendix Figure 8 shown, Figure 8 where Add represents the feature addition operation. Its operation process can be described as follows:

[0204] F′ = Add{F, shuffle[concoat(conv<1 / 2F, 1 / 2F>)]} (29)

[0205] In the basic operation unit, the input is split along the channel dimension. One branch performs feature calculation through pointwise convolution, and the other branch only performs feature concatenation and exchanges information between channels through channel shuffle. The operation structure within the dashed box can be repeated. At the end of the basic operation unit calculation, the input and the feature calculation result are added together as the output through residual connection to reduce the degradation degree of the weight matrix during the forward propagation process.

[0206] The structure of the attention unit is as shown in the appendix Figure 9 As shown Figure 9 where Dense represents the fully connected layer and Sigmoid represents the activation function. Its operation process can be described as follows:

[0207] F′ = Mul{Add[Dense(AvgPool <f>, MaxPool <f>)], F}

[0208] F″ = Mul{Conv[Concat(AvgPool<F′>, MaxPool<F′>)], F′}

[0209] F″′ = ReLu[Add(F″, F)] (30)

[0210] In the attention unit, the dimension of global attention is adjusted by controlling the execution dimension of global pooling. The obtained attention coefficients are multiplied along the corresponding dimension and assigned to the feature map, and the feature map is used as the input for obtaining attention in the next dimension, so as to obtain a feature map with joint attention. Finally, the feature with joint attention obtained is added to the original input through residual connection, and output after passing through the activation function.

[0211] In the above three operation units, depthwise separable convolution is used to implement feature calculation, generating the same number of feature maps with fewer parameters, reducing network parameter expansion and memory consumption; channel shuffle operation is used to solve the problem of non-communication of information between groups after channel segmentation; global pooling is used to obtain the global receptive field along the channel dimension and the spatial dimension, and feature recalibration is performed based on this to reduce the impact of target ghosting caused by reflection noise on the recognition result, thereby realizing adaptive feature optimization; residual connection is used to increase the information entropy of the weight matrix, perform feature correction, and weaken the feature weakening caused by the unclear edges and texture features of the Doppler map; the activation function is used to increase the non-linear relationship between network parameters.

[0212] S302: Based on the three operation units constructed in step S301, a serial pulse ultra-wideband radar gesture recognition model is established. The serial pulse ultra-wideband radar gesture recognition model includes three classification network stages. Each classification network stage sequentially includes a scale expansion unit, a basic calculation unit, and an attention unit, as shown in the appendix Figure 10 As shown, the number of repetitions of the loop structure in the basic calculation units of the three classification network stages is different.

[0213] After feature extraction is completed in the third stage, the spatial dimension scale of the feature is 1×1. Pointwise convolution is used for dimension elevation, and the dimension is adjusted to the same as the number of categories through a fully connected layer. Finally, the classification result is obtained through the softmax function. The appendix Figure 11 is a schematic diagram of the structure of the serial pulse ultra-wideband radar gesture recognition model.

[0214] Experimental verification:

[0215] 1. Radar platform parameter settings

[0216] The parameter settings of the radar platform used in the present invention are shown in Table 2 below, and the network parameter settings are shown in Table 3 below.

[0217] Table 2 Parameter Settings of Gesture Recognition Algorithm

[0218]

[0219] Table 3 Network Parameter Settings

[0220]

[0221] In the first stage of the model, the parameter settings of the scale expansion and compression unit are shown in Table 4 below; its internal structure visualization is as shown in the appendix Figure 12 as follows

[0222] Table 4 Parameter Settings of Scale Expansion and Compression Unit

[0223]

[0224] The parameter settings of the basic operation unit are shown in Table 5 below, and its internal structure visualization is as shown in the appendix Figure 13 as follows

[0225] Table 5 Parameter Settings of Basic Unit

[0226]

[0227] The parameter settings of the attention unit are shown in Table 6 below, and its internal structure visualization is as shown in the appendix Figure 14 as follows

[0228] Table 6 Parameter Settings of Attention Unit

[0229]

[0230] It can be seen that in the scale expansion and compression unit, the channel dimension expansion is only performed once at the beginning of the input, and the scale compression is only performed once. In the basic operation unit and the attention unit, the channel dimension remains unchanged. That is, during the internal calculation of each stage, the channel dimension does not change

[0231] 2. Radar Gesture Recognition Experiment

[0232] To verify the necessity of the signal processing algorithm, ablation experiments were conducted. The experimental settings are as follows

[0233] Experiment 1: Disable cancellation once in signal processing

[0234] Experiment 2: Disable the sliding window Doppler processing algorithm in data processing; the results are shown in Table 7 below

[0235] Table 7 Accuracy of Each Fold and Average Accuracy

[0236]

[0237] From the experimental results in Table 7, it can be concluded that noise, clutter, and data form have a significant impact on gesture recognition results, verifying the effectiveness of the signal processing algorithm in the present invention and providing a relatively stable and high-quality data source for the recognition algorithm.

[0238] Since the lightweight network ShuffleNet V2 is also designed in a stage-based structure, this paper selects the ShuffleNetV2 network for ablation experiments on the attention unit and the basic operation unit. The performance of the network is tested using the publicly available dataset cifar10, and the performance of each operation unit is measured by parameters such as recognition accuracy, number of parameters, convergence loss value, and the number of dataset traversals (epoch) when convergence is achieved.

[0239] The ablation experiment results of the attention unit are shown in Table 8 below.

[0240] Table 8 Ablation Experiment of Attention Unit

[0241]

[0242] The three networks shown in Table 8 above are the ShuffleNet V2 prototype network, the ShuffleNet V2 pruned network named LShuffleV2, and the ShuffleNetV2 pruned network with the attention unit added, named CBAM-LShuffleV2. It can be seen that compared with the prototype network, the computational amount of the pruned network with the attention unit added increases by about 1 / 2, but the convergence point is advanced by 8 epochs.

[0243] The training curves of the above three networks are shown in the appendix Figure 15 as follows. In the figure, the red curve represents the ShuffleNet V2 prototype network, the yellow curve represents the ShuffleNet V2 pruned network, and the blue curve represents the ShuffleNetV2 pruned network with the attention mechanism added. It can be seen more intuitively from Figure 15 that the convergence point of CBAM-LShuffleV2 appears earliest, which is 13.3% earlier than that of ShuffleNetV2. That is, CBAM-LShuffleV2 can find the global optimal point faster and reach a stable accuracy. This fully demonstrates the positive significance of the global receptive field for network training.

[0244] Table 9 below shows the ablation experiment data of the basic operation unit.

[0245] Table 9 Ablation Experiment of Basic Operation Unit

[0246]

[0247] The CBAM-bsi1 network and the CBAM-bsi21 network in Table 9 above represent the ShuffleNetV2 with attention units added and part of the structure replaced by basic operation units. Here, 2 indicates that the repetition times of the basic operation units are 2, and the same applies to the number 1.

[0248] The training curves of the five networks are shown in the appendix Figure 16 as follows.

[0249] In the appendix Figure 16 , the silver curve represents CBAM-bsi1, and the green curve represents CBAM-bsi21. It can be seen from the above experiments that the basic operation units advance the convergence point while reducing the number of parameters, but the decrease in the number of parameters leads to a slight decrease in accuracy. Moreover, from the experimental data, it can be obtained that the more basic operation units are not necessarily better.

[0250] Table 10 below shows the comparison test results between the model proposed in the present invention and the above-mentioned networks.

[0251] Comparison experiments of each model in Table 10

[0252]

[0253] CBAM-3Ubsi121 is the model proposed in the present invention, which has the same recognition accuracy as the prototype ShuffleNetV2. Although the number of parameters has increased by 45%, the convergence point has advanced by 5 epochs, and it still belongs to a lightweight network.

[0254] The accuracy curves of each model are shown in the appendix Figure 17 as follows. In the training curve in the appendix Figure 17 , the model CBAM-3Ubsi121 of the present invention maintains the highest accuracy before the convergence point arrives. During the process of rising to the highest accuracy after the convergence point arrives, its climbing speed is the fastest. This also verifies the effectiveness of the model proposed in the present invention.

[0255] Using the self-built dataset of the present invention, the recognition accuracy of the model for gestures is tested by the 10-fold cross-validation method, and VGG-16, MobileNet V2 and CBAM-3Ubsi121 are set for comparison experiments. Table 11 below lists the recognition accuracies of the three networks for 6 gestures.

[0256] Accuracy of each fold and average accuracy (%) of the three networks for gestures in Table 11

[0257]

[0258] From the results in Table 11, it can be seen that the CBAM-3Ubsi121 network has a higher accuracy rate in 10-fold cross-validation, and its distribution is more stable, with stronger robustness.

[0259] To further verify the robustness of the algorithm, the gesture echo data of 6 participants (p1-p6) were preprocessed in the same way and then input into the VGG-16, MobileNetV2, and CBAM-3Ubsi121 networks for training and verification respectively. Additionally, a volunteer (p7) was invited to provide gesture samples as unknown data to test the network performance. Attached Figure 18 is the experimental result of 10-fold cross-validation for each type of gesture of 7 volunteers.

[0260] As shown in the attachment Figure 18 As shown, blue represents the recognition result of VGG-16, green represents the recognition result of MobileNetV2, and red represents the recognition result of CBAM-3Ubsi121, abbreviated as CB3UB121. It can be seen from the figure that for the gestures of the six volunteers with prior training in the training set, the performance of the three networks remains at a relatively good level. However, in the recognition results of the gestures of the seventh volunteer, the recognition rate of VGG-16 dropped to 89.85%, and the recognition rate of MobileNetV2 decreased to 91.12%, which is much lower than the average accuracy rate of the gesture tests of the first six volunteers. Among them, the average recognition accuracy rate of the first six volunteers of VGG-16 is 94.12%, the average recognition accuracy rate of the first six volunteers of MobileNetV2 is 95.81%, and the recognition rate of the CBAM-3Ubsi121 network reached 96.79%, without a large gap compared with the gesture recognition rate of the first six volunteers. The average recognition accuracy rate of the CBAM-3Ubsi121 for the gestures of the first six volunteers is 97.10%. This verifies the robustness of the attention unit and its strong recognition ability for unknown data.

[0261] Using the gesture dataset after adding the seventh volunteer, the performance of the three networks was tested. The experimental results are shown in the attachment Figure 19 As shown in the figure. The blue curve in the figure represents the recognition result of VGG-16, the green curve represents the recognition result of MobileNetV2, and the red curve represents the recognition result of CBAM-3Ubsi121. The dotted lines of each color represent the average recognition rate of each network for seven volunteers. It can be seen from the results in the figure that due to the addition of unknown data, the cross-validation results of each fold are affected to varying degrees. The average accuracy of the 10-fold cross-validation of the VGG-16 network is 92.29%, the average accuracy of the 10-fold cross-validation of MobileNetV2 is 93.59%, while the CBAM-3Ubsi121 network still has a relatively high recognition rate, maintaining at 96.76%, and the results are relatively stable, proving the robustness and effectiveness of the network.

[0262] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.< / f> < / f> < / f>

Claims

1. A lightweight ultra-wideband radar gesture recognition method based on an attention mechanism, characterized in that, Including the following steps, S1: Collect and preprocess the gesture echo signal of the ultra-wideband radar to obtain the range-Doppler map of the gesture echo signal; S2: Use the method in step S1 to construct a gesture data set; S3: Establish a lightweight micro gesture recognition model according to the characteristics of the gesture data set in step S2; S4: Use the lightweight micro gesture recognition model to recognize the radar gesture; Among them, the specific operations of step S1 include the following steps, S11: Conduct equivalent time sampling analysis on the gesture echo signal of the ultra-wideband radar; S12: Convert the equivalently sampled signal through A / D conversion to obtain a digital signal; S13: Perform clutter suppression and sliding window Doppler processing on the digital signal of the gesture echo to obtain the range-Doppler map of the gesture echo signal; The specific operations of step S3 include the following steps, S301: Construct three operation units for feature extraction from the range-Doppler map of the gesture echo signal. The three operation units are the scale stretching unit, the basic calculation unit, and the attention unit; S302: Based on the three operation units constructed in step S301, establish a serial pulse ultra-wideband radar gesture recognition model; Among them, the specific operations of performing clutter suppression and sliding window Doppler processing on the digital signal of the gesture echo in step S13 include the following steps, S131: Set the radar transmission signal as The echo signal at any moment is In the formula, f r is the echo signal frequency; S132: Assume that the frequency of the analog signal after equivalent sampling is f′, which is equivalent to the frequency of the gesture echo signal, and regard the discrete digital signal after A / D conversion as a continuous signal. Then the digital signal after A / D conversion of the first analog signal after equivalent sampling is S133: After the A / D sampling pulse interval ΔT, the moving distance of the hand is ΔR = vΔT The corresponding time interval is Then the digital signal after A / D conversion of the second analog signal after equivalent sampling The phase change relative to the first digital echo signal is S134: Further analyze the digital signal after A / D conversion of the second analog signal after equivalent sampling to obtain Among them, is defined as the single-step frequency difference of the sliding window Doppler algorithm, is the initial phase; S135: From the digital signal after A / D conversion of the first analog signal after equivalent sampling and the digital signal after A / D conversion of the second analog signal after equivalent sampling, it can be inferred that the discrete signal after sampling is Where, S r1 (k) and S r2 (k) are the discretized sampling sequences of S r1 (t) and S r2 (t), respectively; S136: Subtract adjacent two signals once to complete the elimination of zero-frequency or low-speed clutter, that is S rf1 S(k) = S r2 S(k) - S r1 S(k) Where S rf1( k) is the result of subtracting the echo sequences of adjacent sampling points, and is also the constituent sequence of the data matrix; S137: Select N consecutive pulses through a sliding window, repeat the above steps S131 - S136 for Doppler processing, the sliding step size is m, m is less than or equal to N, and perform a Fourier transform on the data matrix every m sampling periods, and finally obtain the range-Doppler map of the gesture echo signal.

2. The lightweight ultra-wideband radar gesture recognition method based on an attention mechanism according to claim 1, characterized in that, The specific operations of step S11 include the following steps, S111: Based on the "stop-jump" principle, assume that there is relative motion between the ultra-wideband radar and the target, then the distance R(t) between the radar and the target is R(t) = R0 + v r t where R0 is the distance at zero time, and v r = dR(t) / dt is the relative radial motion speed between the target and the radar; Expand the moving target equation R(t) as In the formula, represents the time delay generated by the radar with respect to the initial position, c is the speed of light, c = 3×10 8 m / s; When the moving distance of the target is much smaller than R0, S112: Let the radar echo signal, i.e., the gesture echo signal, be S r (t), and the radar transmit signal be S r (t), then where α is the attenuation factor of the echo, and γ is the stretching factor on the time scale of the echo signal Then the change amount of the single-pulse pulse width caused is In the formula, τ is the single-pulse pulse width; S113: When the transmitted signal is a pulse signal with a period equal to T P then The received signal is S114: Since Therefore, the period of the echo signal changes to T′ P = T P / γ and the variation ΔT of the pulse interval caused P is 3. The lightweight ultra-wideband radar gesture recognition method based on the attention mechanism according to claim 1, wherein: The gesture dataset described in step S2 includes six micro gestures, namely, a finger snap, a left palm wave, a forward rubbing of the thumb and index finger, a backward rubbing of the thumb and index finger, a finger undulation, and a gathering of the five fingers from a flat state to the fingertips.

4. The lightweight ultra-wideband radar gesture recognition method based on the attention mechanism according to claim 1, wherein, The operations of the scale expansion unit for feature extraction include the following steps. Step a1: Convolve the input of the scale expansion unit through two convolutional branches. The convolutional branches sequentially include 1×1Conv, 3×3DW, and 1×1Conv. The 1×1Conv in the left branch contains batch normalization (BN) of the data and a non-linear activation function ReLu. Step a2: Concatenate the outputs of the two convolutional branches along the channels to complete channel expansion. Step a3: Shuffle the expanded channels as the output of the feature extraction of the scale expansion unit.

5. The lightweight ultra-wideband radar gesture recognition method based on the attention mechanism according to claim 1, wherein, The operations of the basic computing unit for feature extraction include the following steps. Step b1: Split the input of the basic computing unit along the channel dimension. One branch performs feature calculation through pointwise convolution and then concatenates the features with the other branch. Step b2: Perform inter-channel information exchange on the concatenated features through channel shuffling. Step b3: Repeat step b1 and step b2. Step b4: Add the feature calculation result of step b3 to the input through residual connection as the output of the feature extraction of the basic computing unit.

6. The lightweight ultra-wideband radar gesture recognition method based on the attention mechanism according to claim 1, characterized in that, The operations of the attention unit for feature extraction include the following steps. Step c1: Adjust the global attention dimension of the input of the attention unit by controlling the execution dimension of global pooling. Multiply the obtained attention coefficients along the corresponding dimension to the feature map and use the feature map as the input for obtaining the attention in the next dimension, thereby obtaining a feature map with joint attention. Step c2: Add the feature with joint attention obtained to the original input through residual connection, and output after passing through the activation function as the output of the feature extraction of the attention unit.

7. The lightweight ultra-wideband radar gesture recognition method based on the attention mechanism according to claim 1, characterized in that, The serial pulse ultra-wideband radar gesture recognition model includes three classification network stages. Each classification network stage sequentially includes a scale expansion unit, a basic computing unit, and an attention unit. The number of repetitions of the loop structure in the basic computing units of the three classification network stages is different.

Citation Information

Patent Citations

  • Gesture recognition detection method based on time-frequency spectrum and distance Doppler spectrum

    CN110309690A

  • Gesture recognition method and related device

    CN113064483A