Millimeter wave radar gesture recognition method based on Trans-CNN model

By adopting a Trans-CNN model-based method in millimeter-wave radar gesture recognition technology, the problem of large amount of feature spectrum map data and inconspicuous features in the prior art is solved, and the gesture feature extraction and strong generalization ability are achieved, reaching a recognition accuracy of 98.4%, and fast and high-precision real-time gesture recognition is achieved.

CN120183029APending Publication Date: 2025-06-20CHINA JILIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510090590.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the existing millimeter-wave radar gesture recognition technology, the feature spectrum map has large data volume, lack of features and high information redundancy, resulting in insufficient feature extraction, poor generalization capabilities of the model, high complexity, and slow recognition speed, and the inability to realize real-time gesture prediction.

Method used

The millimeter wave radar gesture recognition method based on the Trans-CNN model is adopted to analyze and process the gesture ADC data, obtain micro Doppler features, and build a point cloud data set, and use a multi-head self-attention mechanism and a one-dimensional convolutional neural network to build a gesture recognition model Trans-CNN, and input a three-dimensional mixed feature tensor for training and recognition.

Benefits of technology

It realizes sufficient extraction of gesture features and obvious features, strong generalization ability and low complexity, and can quickly and accurately recognize gestures in complex environments, with a recognition accuracy of 98.4%, achieving fast and high-precision real-time gesture recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183029A_ABST
    Figure CN120183029A_ABST
Patent Text Reader

Abstract

The invention discloses a millimeter wave radar gesture recognition method based on a Trans-CNN model. The implementation process comprises the following steps: acquiring original data of different people, backgrounds and gestures through a millimeter wave radar; processing the original data to obtain six characteristics of a time frame, an X coordinate, a Y coordinate, a speed, a distance and an azimuth angle to form a point cloud data set; performing unified processing on the point cloud data to generate a 30 * 45 * 5 feature tensor; transmitting the point cloud data into the model for learning; and connecting the millimeter wave radar to collect gesture data, processing to obtain point cloud data, and inputting the gesture recognition model to obtain a gesture result. According to the method, the gesture features are fully extracted through the point cloud data, the problems of large data volume and insufficient feature extraction of a current feature map are solved, and meanwhile, the problem of long network training time is solved by introducing a multi-head self-attention mechanism. The method is applied to the fields of automobile automatic driving, smart home and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of millimeter-wave radar signal processing, artificial intelligence, etc., and particularly relates to a millimeter-wave radar gesture recognition method based on a Trans-CNN model. Background Art

[0002] In the field of intelligent interaction, gesture recognition is an important biometric technology, which is very friendly and convenient for disabled people and the elderly. Using millimeter-wave radar to achieve non-contact perception and recognition of gestures can well avoid privacy leakage and the limitation of lighting conditions compared with traditional camera-based gesture recognition technology, making gesture recognition more flexible and reliable; in public places, such as hospitals, it can fundamentally reduce the spread of pathogens. Millimeter-wave radar has the characteristics of high frequency and high resolution, and can provide accurate information expression. And it is relatively insensitive to common interference factors such as occlusion, rain and snow, and has high robustness. Therefore, millimeter-wave radar has obvious advantages in gesture recognition technology.

[0003] However, at present, for millimeter-wave radar gesture recognition, mainly the ADC data of gestures is collected offline by the radar, and the data is analyzed and processed by MATLAB to obtain feature spectrograms such as distance-time, speed-time, azimuth-time or distance-speed, distance-angle, etc., and a data set is constructed and transmitted to models such as CNN, LSTM or pure attention for training and learning. On the one hand, the gesture ADC data volume is large, the construction of the data set is cumbersome, the information redundancy of the feature spectrogram is high, and the features are not obvious. On the other hand, the gesture recognition model has insufficient gesture feature extraction, poor generalization ability, high complexity, slow recognition speed, cannot achieve real-time gesture prediction, and is difficult to be deployed to embedded devices. Summary of the Invention

[0004] In order to solve the problems of large data volume, unclear features, high information redundancy, insufficient feature extraction, and poor model generalization ability of the current data set of the feature spectrogram, the present invention proposes a millimeter-wave radar gesture recognition method based on a Trans-CNN model. This method can not only greatly reduce the data set, make the gesture features of the point cloud data obvious, but also accurately and quickly recognize gestures in a complex environment. The model can fully extract gesture features, has strong generalization ability, and low complexity.

[0005] The technical solution for achieving the purpose of the present invention is as follows:

[0006] A millimeter-wave radar gesture recognition method based on a Trans-CNN model includes the following steps:

[0007] (1) Horizontally place the millimeter-wave radar and collect the ADC data of gestures of different people, in real application environments, and with different postures;

[0008] (2) Analyze and process the gesture ADC data, obtain the micro-Doppler features of the gesture actions, and establish a point cloud data set of the gesture actions;

[0009] (3) Uniformly process the micro-Doppler features of each gesture action sample in the point cloud data set, perform slice deletion and zero-padding to obtain a 30*45*5 three-dimensional hybrid feature tensor;

[0010] (4) Use the multi-head self-attention mechanism and the one-dimensional convolutional neural network to construct a gesture recognition model Trans-CNN. This model receives the three-dimensional hybrid feature tensor as input, optimizes the model weights during the training and testing processes, and finally obtains the trained Trans-CNN model;

[0011] (5) Connect to the ADC data of the gesture actions collected by the millimeter-wave radar in real time. After steps (2) and (3), obtain the three-dimensional hybrid feature tensor with the micro-Doppler features of the gesture actions; input this feature tensor into the pre-trained model for gesture recognition; if a preset gesture action is predicted, prompt to correctly input the gesture preset in the model; otherwise, continue to execute step (5);

[0012] (6) Send the recognized gesture action to the computer.

[0013] Further:

[0014] For the acquisition of the gesture ADC data described in step (1), connect the millimeter-wave radar to the PC through the USB interface. The millimeter-wave radar is an arbitrary 2-transmit 4-receive linear frequency modulation continuous wave radar. Use the upper computer software to send the configuration file through the serial port, adopt the TDM-MIMO scheme to transmit the frequency modulation continuous wave, after gesture reflection, obtain the echo data, mix it with the transmitted signal to obtain the intermediate frequency IF signal, and perform ADC sampling to obtain the ADC data.

[0015] For the analysis and processing of the gesture ADC data described in step (2) to obtain the micro-Doppler features of the gesture actions, the following steps are included:

[0016] 1) Decode the received LVDS data into a complex form, recombine it according to the number of antennas, the number of Chirps, and the number of sampling points, and perform 1D-FFT transformation on all sampling points of each Chirp to obtain the distance features of the gesture actions;

[0017] 2) For the data with the obtained distance features, use the vector mean cancellation algorithm to filter out the echo data generated by the static object background in the gesture data to obtain the gesture action data after filtering out the static objects and the background;

[0018] 3) For the gesture motion data obtained by filtering static objects and the background, perform 2D-FFT transformation on each sampling point of all Chirps in the velocity dimension, and the velocity characteristics of the gesture motion can be obtained;

[0019] 4) After obtaining the velocity characteristics, use the Constant False Alarm Rate (CA-CFAR) algorithm to estimate the dynamic clutter in the data and remove the clutter within the threshold;

[0020] 5) For the data obtained with range and velocity characteristics, perform Angler-FFT transformation in the receiving antenna dimension to obtain the azimuth angle characteristics of the gesture motion.

[0021] The establishment of the point cloud data set of gesture motion in step (2) is to design a peak grouping algorithm for the data that has already obtained the micro-Doppler characteristics, group the internal target points of each gesture sample after filtering the dynamic clutter, retain the target points that meet the requirements, and delete the others. Finally, the point cloud data set of the gesture can be obtained; each gesture data has different time frames (frame) according to the gesture, but they are all between 20 and 30 frames. At the data transmission rate of 115,200 bps in serial communication, the data acquisition and transmission time for each gesture is 2 - 3 s.

[0022] The decoding of the received LVDS data into complex form in step 1) is to sequentially extract the data of four receiving antennas corresponding to each transmitting antenna, splice them according to the number of antennas, the number of Chirps, and the number of sampling points, and finally obtain the three-dimensional matrix data of 8 (antennas) * 480 (Chirps) * 256 (sampling points). Perform 1D-FFT transformation on all sampling points of each Chirp of each antenna of this matrix to obtain the range characteristics of the gesture motion. The expression is:

[0023]

[0024] In the formula: B is the bandwidth; T c is the pulse width; f IF is the frequency of the intermediate frequency signal; c is the speed of light.

[0025] For the data obtained with range characteristics in step 2), filter the echo data generated by the static object background in the gesture data by means of the vector mean cancellation algorithm. The expression is:

[0026]

[0027] F1D_StaticOut(N c , N s , N t_r ) = F1D(n c , :, n t_r ) - avg(3)

[0028] Where: F1D(N c , N s , N r_t ) is the distance gesture data obtained by 1D-FFT calculation; N c is the number of Chirps, N s is the ADC sampling rate of each Chirp, N t_r is the product of the number of transmitting and receiving antennas; n c is the cyclic factor of the Chirp number; n t_r is the cyclic factor of the antenna number; avg is the vector mean of all sampling points of each Chirp; F1D_StaticOut is the gesture data after filtering out static objects and background echoes.

[0029] In step 3), the gesture action data after filtering out static objects and background is subjected to 2D-FFT transformation for each sampling point of all Chirps in the velocity dimension, and its expression is:

[0030]

[0031] Δω = 2πΔf d T c (5)

[0032] Where: v is the velocity of the target; λ is the wavelength of FMCW; Δω is the phase change between two Chirps; Δf d = -2vf0 / c is the Doppler frequency shift caused by the gesture movement.

[0033] In step 4), after obtaining the velocity characteristics, the mean constant false alarm rate algorithm (CA-CFAR) is used to estimate the dynamic clutter in the data and eliminate the clutter within the threshold, which adaptively adjusts the decision threshold according to the dynamic clutter in the echo signal; CFAR detection first assumes that the dynamic clutter is known, and then uses the reference cells in the data to estimate parameters such as clutter noise, so as to obtain the threshold based on the estimated density; if there is a value in the detected data that exceeds the threshold, it is considered that there is a target:

[0034]

[0035] Where: H is the threshold of the cell to be measured; K is the threshold of clutter estimation; N is the length of the distance dimension or velocity dimension; M is the length of the guard cell; x i / x j is the input data of the distance dimension or velocity dimension.

[0036] In step 5), the Angler-FFT transformation is performed in the receiving antenna dimension to obtain the horizontal angle characteristics of the gesture action. First, the FFT calculation is performed in the distance dimension, and then the FFT calculation of the target angle is performed in the antenna dimension. Its expression is:

[0037]

[0038] Where: S(θ) is the azimuth angle characteristic distribution of the target; X(m) is the complex signal received by the m-th antenna; d is the antenna element spacing; λ is the radar signal wavelength; θ is the azimuth angle; m is the antenna index.

[0039] The construction of the gesture recognition model Trans-CNN described in 4) above includes:

[0040] 1) The Trans-CNN model is mainly composed of three parts: an attention module, a serial one-dimensional convolutional module, and a fully connected layer;

[0041] 2) The attention module mainly consists of two parts: temporal position encoding and the calculation of the dependencies between sequence targets;

[0042] 3) The serial one-dimensional convolutional module is formed by connecting two one-dimensional convolutional layers in series through an average pooling layer;

[0043] 4) Perform a global average pooling operation on the output, and finally use a fully connected layer to output the probability of each gesture using softmax;

[0044] 5) Input the 30*45*5 three-dimensional hybrid feature tensor obtained in step (3) into the constructed Trans-CNN model, and train and validate the model until it converges.

[0045] The temporal position encoding described in step 2) is to flatten the data distributively for each frame in the time dimension using the Time Distributed function; use the Positional Encoding function to add non-linear encoding values to the sequence using the positional encoding method.

[0046] The calculation of the dependencies between targets described in step 2) is to continuously learn and calculate the dependencies between different positions of each target point using the attention mechanism based on scaled dot product. During the construction of the model, the number of neurons, the number of attention heads, the probability of Dropout, etc. in each layer can be adjusted according to your own needs.

[0047] A millimeter-wave radar gesture recognition system based on the Trans-CNN model provided by the present invention includes a microprocessor, an mmWave radar development board, a computer, and a program that can run in real time on the computer. Through this system, real-time gesture recognition of the millimeter-wave radar can be achieved.

[0048] The beneficial effects of the present invention are:

[0049] The present invention uses a millimeter-wave radar that emits frequency-modulated continuous waves to collect the echo signals reflected by gesture movements, superimposes them with the transmitted signals to obtain intermediate-frequency (IF) signals, and processes them through algorithms such as 3D-FFT, non-coherent superposition, and peak grouping to obtain the micro-Doppler characteristics of gesture movements. Six different micro-Doppler characteristics with complementary physical meanings are generated for the same gesture of the same person. Gesture point cloud datasets are established by collecting gestures in different postures from different people, and the datasets are unified and input into a neural network to detect the target gestures. Based on the different micro-Doppler characteristics of each gesture, the correlation between target points and the extraction of features between different gestures are further carried out, so as to quickly, accurately, and real-time identify 10 gesture movements in the dataset, and transmit the recognition results to the computer host. The host can control external devices through different gesture information, thereby realizing corresponding functions.

[0050] For the IF signal of gesture movements in the present invention, Fourier transform is performed in the distance dimension through ADC data, and the vector mean cancellation algorithm is used. Two-dimensional Fourier transform is performed in the velocity dimension, and angular Fourier transform is performed in the receiving antenna dimension. Finally, peak grouping processing is carried out to obtain point cloud data with different physical meanings of time frames, horizontal and vertical coordinates, distance, velocity, and azimuth angle. Compared with traditional spectrum diagram data such as distance-time, Doppler-time, and distance-Doppler, using point cloud as the input data of the model overcomes the problems of high information redundancy, unclear features, insufficient feature extraction, large data volume, long model training time, and slow gesture recognition speed in the spectrum diagram. Therefore, in the present invention, point cloud data is used to express gesture movements, enabling sufficient and obvious gesture feature extraction, which is beneficial to the rapid convergence of the model and lays the foundation for achieving high-accuracy gesture recognition.

[0051] The Trans-CNN network constructed in the present invention first uses the multi-head self-attention mechanism to easily capture the dependencies in time-series data and calculate the correlation degree between each time step and other time steps. Since the multi-head self-attention mechanism overly focuses on important information in the sequence and ignores the role of unimportant information in the sequence, the present invention introduces a one-dimensional convolutional network after the multi-head self-attention. By setting multiple convolutional layers and different convolutional kernel sizes, the feature extraction is further enriched from multiple segments, increasing the accuracy of gesture recognition. The residual idea is used to fuse the features obtained in each layer, overcoming the problems of gradient disappearance and gradient explosion in neural network training. Through the combination of the multi-head self-attention network and the one-dimensional convolutional network, the gesture movement features contained in the point cloud data are fully extracted, thereby improving the accuracy of gesture movement recognition. Experiments show that this method can accurately identify 10 gesture movements, with an accuracy rate of 98.4%, achieving fast and high-precision real-time gesture recognition. Description of the Drawings

[0052] Figure 1 It is a diagram of the gesture recognition system of the present invention;

[0053] Figure 2 It is a diagram of the gesture actions and accumulated point clouds of the present invention;

[0054] Figure 3 It is a flowchart of the gesture recognition system of the present invention;

[0055] Figure 4 It is a flowchart of obtaining point cloud data of the present invention;

[0056] Figure 5 It is a comparison diagram before and after static clutter filtering of the present invention;

[0057] Figure 6 It is a schematic diagram of the CA-CFAR principle of the present invention;

[0058] Figure 7 It is an effect diagram before and after peak grouping of the present invention;

[0059] Figure 8 It is a schematic diagram of the Trans-CNN structure of the present invention;

[0060] Figure 9 It is a schematic diagram of the calculation process of the single-head attention mechanism of the present invention;

[0061] Figure 10 It is a diagram of different palm shapes of the present invention;

[0062] Figure 11 It is a curve graph of the changes in Acc and Loss of the Trans-CNN model of the present invention;

[0063] Figure 12 It is a radar diagram of ADT6101 of the present invention;

[0064] Figure 13 It is a curve graph of the change in Acc between IWR1642 and ADT6101 of the present invention;

[0065] Figure 14 It is a confusion matrix diagram of four gestures of the ADT6101 radar of the present invention. Specific embodiments

[0066] The present invention will be further described below in conjunction with the accompanying drawings and embodiments, but it is not a limitation to the content of the present invention.

[0067] Embodiment:

[0068] Gesture recognition system:

[0069] In the gesture recognition system of the present invention, the 77GHz millimeter-wave radar IWR1642 of Texas Instruments and a laptop computer are used. The gesture recognition system mentioned in the present invention is as Figure 1 shown.

[0070] In the gesture recognition system, the millimeter-wave radar uses two transmitting antennas and four receiving antennas. Therefore, the ADC data only contains the feature of azimuth angle and does not contain the feature of elevation angle. Therefore, the dependence of the 10 gesture actions designed in the present invention on the elevation angle is negligible. All gesture actions are completed about 20 - 70 cm directly in front of the radar, including waving the hand upward, waving the hand downward, waving the hand to the left, waving the hand to the right, rotating clockwise, rotating counterclockwise, drawing a Z, drawing an S, drawing an X, and drawing a √. As Figure 2 shows the actual gesture actions and the cumulative point cloud map, where the color from light to dark represents the movement direction of the gesture.

[0071] When collecting gesture actions, place the radar horizontally and perform gesture actions about 20 cm directly in front of the radar. Parameters such as the chirp number and frame period of the millimeter-wave radar have been pre-designed. See Table 1.

[0072] Table 1 Parameter Configuration of Millimeter-Wave Radar

[0073]

[0074] Gesture Recognition Method:

[0075] The gesture recognition method in the present invention includes: processing ADC data to obtain point cloud data, establishing a gesture data set, establishing a gesture recognition model, and recognizing real-time gestures. The implementation flowchart is as Figure 3 shown.

[0076] Execute S110, place the millimeter-wave radar horizontally, and collect gesture ADC data of different people, in real application environments, and with different postures.

[0077] Execute S120, analyze and process the gesture ADC data, obtain the micro-Doppler features of the gesture actions, and establish a point cloud data set of the gesture actions.

[0078] Execute S130, uniformly process the micro-Doppler features of each gesture action sample in the point cloud data set to obtain a 30*45*5 three-dimensional hybrid feature tensor.

[0079] Execute S140, use the multi-head self-attention mechanism and a one-dimensional convolutional neural network to construct a gesture recognition model Trans-CNN. This model receives the three-dimensional hybrid feature tensor as input, optimizes the model weights during the training and testing processes, and finally obtains a trained Trans-CNN model.

[0080] Execute S150 to collect the ADC data of the gesture movement in real time through the connected millimeter-wave radar, and obtain the three-dimensional hybrid feature tensor with the micro-Doppler feature of the gesture movement through S120 and S130. Input this feature tensor into the model for gesture recognition. If the preset gesture movement is predicted, prompt to correctly input the gesture preset in the model; otherwise, continue to execute S150.

[0081] Execute S160 to send the recognized gesture movement to the computer.

[0082] For S110, the radar is placed horizontally to collect the ADC data of the gesture movement. The specific collection implementation method is to use the 77GHz millimeter-wave radar IWR1642 of Texas Instruments. Connect the radar to the PC through the USB interface. The radar is configured as 2 transmit and 4 receive, and the other parameters have been mentioned in the previous gesture recognition system. Use the host computer software to set the baud rate of the serial port to 921600bps, set one stop bit, eight data bits, and no parity bit, and send the configuration file to the radar board through the serial port using the software.

[0083] For S120, process the collected ADC data through algorithms such as 3D-FFT, vector mean cancellation algorithm, non-coherent superposition, and peak classification extraction to obtain point cloud data, and send it to the computer side through the serial port for storage, training, and recognition. Among them, the radar uses the TDM-MIMO scheme to transmit frequency-modulated continuous waves.

[0084] The transmitted signal of the radar is:

[0085] s τ (t) = A τ cos[π(2f0t + St 2 ) + φ0] (1)

[0086] Where: A τ is the FMCW amplitude; f0 is the center frequency of the carrier; S = B / T c is the FMCW slope; B is the bandwidth; T c is the pulse width; φ0 is the initial phase, generally taken as 0.

[0087] After transmitting the signal and passing through the delay t d = 2(R0 + vt) / c and the Doppler frequency shift Δf d = -2vf0 / c caused by the gesture movement, the received reflected signal is:

[0088] s r (t) = A r cosπ[2f0 + S(t - t d )](t - t d ) (2)

[0089] f r = S(t - t d ) + Δf d (3)

[0090] where: A r is the amplitude of the received echo signal; f r is the frequency of the received signal.

[0091] The transmitted signal and the echo signal are mixed and filtered by a low - frequency filter to obtain an intermediate - frequency (IF) signal.

[0092]

[0093] where: A = A r * A τ is the amplitude of the IF signal.

[0094] In collecting each gesture ADC data, the number of chirps is a fixed value, and each chirp is set with 256 sampling points. For each kind of gesture action, the acquisition time is different. Therefore, in the present invention, an infinite - frame continuous transmission is designed until there is no gesture action within 0.5 s (5 frames), that is, the current gesture acquisition is completed, and at the same time the radar stops sending signals. By real - time monitoring whether the gesture ends, the integrity of each gesture can be perfectly guaranteed. Different people have different gesture times, resulting in different time frames for gestures. After processing, the number of target points per frame of the point - cloud data will also be different. After all gesture data is collected, it is found that all gesture samples can be collected within 30 frames. Therefore, when inputting data into the subsequent model, the data is unified according to the maximum sampling time frame. Each time a gesture is collected, a csv file is obtained. By repeatedly collecting, multiple sample data of multiple gestures are obtained. These sample data are saved separately in the corresponding folder paths according to the gesture types for subsequent processing. The first row of each csv file is the title, which are the time frame (f), X - axis coordinate (x), Y - axis coordinate (y), distance (range), velocity (velocity), and azimuth (azimuth) respectively. Each subsequent row represents a target point. The static point - cloud of the gesture is as Figure 2 shown.

[0095] Regarding the processing of ADC data in S120, the flow chart for extracting micro - Doppler features and establishing a point - cloud database is as Figure 4 shown. The specific implementation process is as follows:

[0096] S210: Decode the received ADC data in LVDS format into a complex form, recombine it according to the number of chirps, the number of sampling points, and the number of antennas, and perform a 1D-FFT transform on all sampling points (range dimension) of each chirp to obtain the range feature of the gesture action.

[0097] S220: For the gesture data with the obtained range feature, filter out the clutter signals generated by static objects, background, etc. in the gesture data by means of the vector mean cancellation algorithm. Obtain the gesture action data after filtering out the clutter signals.

[0098] S230: For the data with the obtained clutter signal filtered out, perform a 2D-FFT transform on each sampling point (velocity dimension) of all chirps to obtain the velocity feature of the gesture action.

[0099] S240: For the gesture data with the obtained velocity feature, use the mean constant false alarm detection algorithm (CA-CFAR) to filter out the dynamic interference generated by irrelevant personnel, the body, and environmental noise in the gesture data, and obtain the gesture action data after filtering out the dynamic interference.

[0100] S250: For the gesture action data after filtering out the dynamic interference, perform an Angler-FFT (3D-FFT) transform in the receiving antenna dimension to obtain the azimuth feature of the gesture action.

[0101] S260: Corresponding the obtained five features of range, velocity, azimuth, X-axis, and Y-axis to the detected gesture targets one by one according to the time frame, and output them as a target point. The data format formed by each gesture is a three-dimensional matrix composed of the time frame * target point * feature of each gesture. In order to fully display the effect of dynamic gestures, a peak grouping algorithm is further designed to group the target points inside each gesture sample after filtering out the dynamic clutter. The target points that meet the requirements are retained, and the others are deleted. Obtain the final point cloud data.

[0102] Regarding the splitting and combining of ADC data into a complex form in S210, sequentially extract the data of the four receiving antennas corresponding to each transmitting antenna, splice the chirps of each frame according to the receiving antenna, and finally obtain a three-dimensional matrix data of 960 (chirps) * 256 (sampling points) * 8 (antennas). Perform a 1D-FFT transform on all sampling points (range dimension) of each chirp of each antenna to obtain the matrix F1D(N c ,N s ,N r_t ) containing range information after the transform, where N c is the number of chirps, N s is the number of sampling points of each chirp, N t_rIt is the product of the number of transmitting and receiving antennas. Each data in the matrix represents the echo signal intensity of a certain chirp of a certain antenna at a certain sampling point. The distance calculation expression is:

[0103]

[0104] In the formula: c is the speed of light.

[0105] For the data containing distance characteristics obtained in S220, the echo data generated by the static object background in the gesture data is filtered by means of the vector mean cancellation algorithm. First, for each chirp, the intensities of all sampling points are added together and the average value is calculated. The expression is:

[0106]

[0107] Secondly, subtract the average value from all the data of F1D(N c ,N s ,N r_t ) to obtain the gesture data with the static object and background echo filtered. The static filtering effect is as Figure 5 shown. The expression is:

[0108] F1D_StaticOut(N c ,N s ,N t_r ) = F1D(n c ,n s ,n t_r ) - avg(6)

[0109] In the formula, avg is the vector mean of all sampling points of each chirp, and F1D_StaticOut is the gesture data with the static object and background echo filtered.

[0110] Perform 2D-FFT transformation on the velocity dimension for S230. Perform 2D-FFT transformation on all chirps (i.e., the velocity dimension) of each sampling point in the F1D_StaticOut matrix to obtain the transformed matrix F2D(N c ,N s ,N r_t ). The velocity calculation expression is:

[0111]

[0112] In the formula: λ is the wavelength of LFMCW; Δω is the phase change between two Chirps.

[0113] For S240, CA-CFAR is used to filter dynamic interference, and the principle is as Figure 6As shown below. The specific process is as follows: After 2D-FFT, the Cell Averaging Constant False Alarm Rate (CA-CFAR) algorithm is adopted in the range dimension and the velocity dimension respectively to estimate the non-target clutter in the environment and eliminate the clutter within the threshold. Its core idea is to estimate the threshold of the result of 2D-FFT and compare it with this threshold. If it is greater than this threshold, it is considered a target. Otherwise, it is clutter.

[0114]

[0115] Where: H obj is the threshold of the unit to be measured; K is the threshold of clutter estimation; N is the length of the range dimension or the velocity dimension; M is the length of the guard cell; x in the range dimension i is all sampling points of each Chirp, and x in the velocity dimension i is all Chirps of the same sampling point.

[0116] For the results obtained by S250 after CFAR detection of each day's receiving antenna, an Angler-FFT transform is performed in the third dimension, i.e., the antenna dimension, to obtain the azimuth angle feature of the gesture action, and its expression is:

[0117]

[0118] Where: ω is the phase difference between the corresponding data in the matrix after 2D-FFT corresponding to the receiving antenna; D Rx is the distance between the receiving antennas.

[0119] Each transmitting antenna has four receiving antennas. Angle FFT processing is performed on the data received by the receiving antennas corresponding to each transmitting antenna respectively to obtain the angle data of all targets, and finally the average value is taken to obtain the azimuth angle.

[0120] In S130, the 6 micro-Doppler features in each sample are unified to obtain a 30*45*5 three-dimensional hybrid feature tensor. The specific processing method: The three-dimensional data of time frame * target point * feature obtained in the specific implementation S260 is uniformly processed into three-dimensional data with a format of 30*45*5 through the designed peak grouping algorithm. The processing process is: (1) For the data after CFAR in the range dimension, set the peak grouping threshold H p in the range dimension. When there is an output in the range dimension CFAR and the peak is greater than H p , record the distance index of this target point. (2) For the data after CFAR in the velocity dimension, set the same threshold H p in the velocity dimension, and check whether there is an output in the velocity dimension CFAR at the distance index of the target point retrieved in step 1 and the peak is greater than H p, if any, record the target speed index. (3) For each frame of data, retrieve all the points that match the target speed index as the final point cloud of the current frame for output. As Figure 7 shown.

[0121] In the figure, the colors from light to dark represent the motion direction of the gesture (from the 0th frame to the last frame). The plus sign indicates points less than the threshold H p (points that do not meet the requirements). The pentagram indicates points outside the gesture motion range. The dot indicates points that meet the requirements. 6-a represents the cumulative point cloud before peak grouping, 6-b represents the cumulative point cloud after using the peak grouping algorithm, and 6-c represents the cumulative point cloud after deleting points outside the threshold.

[0122] For the situation where there are certain differences in the same gesture collected by volunteers, and not all gesture samples are 30 frames with 45 target points per frame, the present invention performs zero-padding processing on those with less than 30 frames or less than 45 target points per frame, and performs slicing and deletion processing on those with more than 30 frames or more than 45 target points per frame. All gesture samples can be uniformly processed into 30 frames per sample, 45 target points per frame, and 5 features per target point. Finally, a three-dimensional mixed feature tensor of 30*45*5 is formed. Table 2 shows the reorganized point cloud data. Each row represents a target point.

[0123] Table 2 Point cloud data of upward waving

[0124]

[0125] Execute S140 based on the composite model Trans-CNN of the multi-head self-attention mechanism and the one-dimensional convolutional neural network. Input the three-dimensional mixed feature tensor in the second step into the gesture recognition model Trans-CNN, and train and validate the network. Through experimental verification, the accuracy rate of this model reaches 98.5%, and the training and convergence speed of the model is very fast, realizing fast and high-precision real-time gesture action recognition. The accuracy rate is higher than the methods that use distance-time spectrograms, speed-time spectrograms, azimuth spectrograms, etc. as databases and use LSTM models, CNN models, etc. as recognition models to predict gestures.

[0126] The specific structure of the Trans-CNN model is as Figure 8 shown, and the implementation process is as follows:

[0127] Step 1: Copy the three-dimensional mixed feature tensor input(30*45*5) and send it to the input layer for data format matching. According to the size of the gesture database, adjust the number of gesture samples included in each training. In the present invention, the batch size used is 16 (batch_size = 16).

[0128] Step 2: Pass the output tensor of the input layer through the Time Distributed layer to flatten the data distributively for each frame in the time dimension using Flatten, ensuring that the data for each time frame is fully involved in the calculation of the attention mechanism. Map the flattened data to 128 dimensions through the linear transformation of the fully connected layer. The learning through the linear transformation helps to extract features in the sequence. Use the Positional Encoding layer to add fixed encoding to the data so that the attention layer can learn the temporal relationship between targets in each frame sequence.

[0129] Step 3: Copy the data inputs with positional encoding added and perform normalization processing.

[0130]

[0131] Where: X i is the feature at the i-th position of each frame; is the mean of X i ; is the standard deviation of X i ; is a constant.

[0132] Send it to the multi-head self-attention layer. The result Out_MH obtained through linear transformation, calculation of attention scores, weighted summation, and multi-head merging is connected with inputs through residual connection and then normalized to obtain the output tensor of the attention module (MSA_inputs).

[0133] Regarding the network result of the multi-head self-attention layer mentioned in Step 3 as Figure 8 shown in -b. Specific calculation process:

[0134] Step 3-1: The calculation in the multi-head self-attention mechanism mainly depends on the query vector Q, key-value vector K, and value vector V. Use the input sample data to train the three weight matrices (W q , W k , W v ) of the linear layer, and map the input feature tensor to the query vector space, key-value vector space, and value vector space respectively through the three weight matrices. The specific expressions are as follows:

[0135]

[0136] Where x is the input data.

[0137] Step 3-2: Take out the query vector and perform a dot product with the key-value vector. Transpose the key-value vector K and multiply it with Q to obtain the correlation between Q and K, and calculate the weight coefficient through softmax normalization. Its expression is:

[0138]

[0139] In the formula, d k is the scaling factor.

[0140] Step 3-3: Apply the weight coefficient calculated in Step 3-2 to the value vector V corresponding to 3-1 to improve the generalization ability of the model, obtain the attention to the gesture sample data, and thus obtain the output data of a single attention head. Its expression is:

[0141]

[0142] In the formula: h i is the output of the i-th attention head, Q is the query vector; K is the key-value vector; V is the value vector; d k is the length of K.

[0143] When calculating the output data of the self-attention mechanism, the Dropout function is used to randomly discard 50% of the data to prevent overfitting of the network. The calculation process of the single-head self-attention mechanism is as Figure 9 shown.

[0144] Step 3-4: In the present invention, 16 attention heads are designed. The calculations of Q, K, and V in each attention head are independent of each other. Therefore, each head has its own query, key, and value vectors. Different feature expressions are learned in different feature subspaces to better capture the relationships and dependencies between different features in the input data. Therefore, the expressions of the query, key, and value vectors of the remaining attention heads in different feature subspaces are calculated using the calculation methods of the first three steps. And they are concatenated on the feature dimension. Its expression is:

[0145] Out_MH = MultiHead(Q, K, V) = Concat(h1, h2, …, h i )(14)

[0146] In the formula, Out_MH is the concatenation of the 16 attention heads on the feature dimension. h i is the number of attention heads.

[0147] Step 4: Copy MSA_inputs and input them into the CNN module. After one-dimensional convolution, the PreLU activation function, and the Dropout layer, the calculated result is obtained.

[0148] The network structure of the CNN module for Step 4 is as Figure 8 -c shown. The specific calculation process:

[0149] Step 4-1: Input the copied MA_inputs into the 1D convolutional layer for convolution operation. In the present invention, the kernel size in the first feature extraction module is set to 1, the number of kernels is set to 256, and the stride and padding are default values. Its expression is:

[0150]

[0151] In the formula: f(·) is the activation function; x ij is the input data and output data; i, j are the positions of the target points in the data; n is the number of the one-dimensional convolutional layer; w n is the weight matrix of the convolutional layer; b is the bias; Down(·) is the average pooling function; x i,j is the feature after pooling; x i,j-1 is the feature before pooling.

[0152] Step 4-2: Introduce non-linearity to the result of the one-dimensional convolution calculation through the PReLU activation function to capture the relationship between features. After reducing the result of the calculation by half through Dropout, the result of the first convolutional layer is obtained through one-dimensional average pooling. Then, through conv1d (kernel size set to 1, number of kernels set to 128), PReLU activation function, and Dropout, features in different time and space are obtained.

[0153] Step 5: The output result of the CNN module is passed through conv1d (kernel size set to 1, number of kernels set to 128), Dropout, and normalization, and then connected with MA_inputs through residual connection, and a global average pooling operation is performed. The mean value of all time features is processed to generate one-dimensional feature data. Finally, the SoftMax function is set through the fully connected layer to obtain the probability sizes of 10 gestures.

[0154] During the process of building the model, the number of neurons in each layer, the number of attention heads, the probability of Dropout, etc. can be adjusted according to your own needs.

[0155] For the gesture recognition system and gesture recognition method proposed in the present invention, experimental verification was carried out, and the achieved effects are as follows:

[0156] Specific example 1:

[0157] The gesture recognition method proposed by the present invention includes a total of 10 gestures, namely waving upward, waving downward, waving left, waving right, rotating clockwise, rotating counterclockwise, drawing an S, drawing a Z, drawing a tick, and drawing a cross. When collecting gesture actions, only need to place the radar horizontally, about 20 cm in front of the radar to perform gesture actions. The present invention invited a total of 5 experimental personnel (3 males and 2 females) to construct a gesture point cloud dataset. Each experimental personnel changed the shape of the hand to achieve slight changes in the same gesture, enhancing the universality of the gesture. Two palm shapes were used in this experiment. As Figure 10 shown. For the same gesture, 100 samples were collected from each person, and finally 10 gestures totaled 10,000 sample data were formed.

[0158] The model was trained on an I9-12900H + GPU3060 + Win11 laptop. The input data of the model was uniformly processed into 30*45*5. The number of training epochs of the model, Epoch, was set to 200 rounds. The EarlyStopping function was used. When the loss value of the model on the validation set did not decrease within 30 consecutive epochs, the training would stop and the best model would be saved. The loss function used the categorical cross-entropy function (categorical_crossentropy) which is very widely used for multi-object classification applications. The optimizer used the Adam optimizer with an adaptive learning rate. The learning rate was adjusted dynamically. When the loss value Loss did not decrease within 10 epochs on the validation set, training would continue by reducing to 0.2 times the current learning rate. The initial value was set to 0.001. The training set and validation set of the model were divided according to 70% and 30%. Num_heads and Batch_size were set to 16.

[0159] During the training and testing process, the accuracy (Acc) and loss value (Loss) were used to evaluate the model. The change curves are as Figure 11 shown. The overall accuracy of all gestures reached 98.5%.

[0160] To test the performance of the model, 1500 sample data were collected again through the IWR1642 radar, 150 samples for each gesture. The TRANS-CNN model was verified, and a confusion matrix was drawn on the validation set. At the same time, the recall rate, precision, and F1-Score were used to perform performance statistics for each gesture. The detailed results are shown in Table 3.

[0161] Table 3 Model performance corresponding to 10 gestures

[0162]

[0163]

[0164] The gestures in the table represent waving up (up), waving down (down), waving left (left), waving right (right), rotating clockwise (cw), rotating counterclockwise (ccw), drawing Z, drawing S, drawing a cross (×), and drawing a tick (√) respectively.

[0165] Among them, the recognition accuracy of 8 gestures, namely up, down, left, ccw, drawing S, and drawing √, in the confusion matrix reaches over 97%. It can be seen that the TRANS-CNN model has a very good effect on gesture recognition. However, the model is not very good at recognizing the right, Cw, and × gestures, but the accuracy also reaches over 96%. The recognition accuracy of the Z gesture is 94%, and the recognition effect needs to be improved. The average recognition accuracy of the ten gestures reaches 98.4%.

[0166] Specific Example 2:

[0167] Taking the feature spectrogram as the input data of the model, the size of each gesture data is 300 - 400k, and the entire dataset will be very large. Experiments show that this kind of data makes the model have many training parameters and a slow model convergence speed. While using point cloud data as the input data of the model, the data size of each gesture is only 4 - 9k, and the entire dataset is very small. During the training process, the training parameters will be greatly reduced, and the convergence speed will also increase significantly. It only takes 5 - 10 minutes to complete the training of all data, and only 4 - 5s for a single iteration. Table 4 compares the other four models with point cloud data as the input in this study and the other three models with feature spectrogram as the input data. Through comparison, it can be seen that on the basis of the same gesture types, the 3DCNN + Transformer model and the Transformer model have increased by 0.14%, indicating that the addition of the convolutional network model can extract richer feature information, but the model complexity is relatively high. The two-stream metric learning prototype network recognizes gestures through two-way four-layer CNN. Compared with the Transformer model and the 3DCNN + Transformer model, the accuracy has increased by 1.3% and 1.16% respectively, and the training time for each Epoch is 4.56s. The model has a fast convergence speed and high accuracy, but the input data volume is large, and it is difficult to achieve real-time gesture recognition. While the proposed Trans-CNN network in this study has a recognition accuracy of 98.5%. Compared with the RFT model based on pure attention mechanism, although the number of gesture types has been reduced to 10, the accuracy has increased by 3.02%, and the gesture recognition is rapid. Through comprehensive comparison, the model proposed in this study using point cloud as the input data has certain improvements in terms of accuracy, recognition speed, and real-time performance. It has good application value.

[0168] Table 4 Comparison of the training results of each model

[0169]

[0170] Specific Example 3:

[0171] Using the radar RF board ADT6101 (such as Figure 12 ), the radar is set to transmit two signals and receive two signals, and the frequency-modulated continuous wave is transmitted using the TDM-MIMO mode to collect gesture point cloud data. Two volunteers were invited for the experiment, and data of four gestures, namely up, down, left, and right, were collected respectively. There were 120 groups for each gesture, and a total of 480 groups of point cloud data were obtained. Subsequently, the dataset was divided in a ratio of 7:3 for the training and testing of the TRANS-CNN model. The model has a fast convergence speed, and the specific test results are listed in Table 5. Using the ADT6101 radar, although the gesture recognition accuracy decreased slightly, it still remained above 97%, indicating a relatively high gesture recognition accuracy. This shows that the model can achieve significant performance on different radar devices.

[0172] To more comprehensively evaluate the performance of the model on the test set, the curves of the IWR1642 and ADT6101 radar data against the model test accuracy (Acc) were plotted (see details in Figure 13 ). The results shown in the figure indicate that for both the IWR1642 and ADT6101 radars, the model shows a fast convergence speed and high accuracy on the test set. Therefore, the model can quickly and accurately recognize gestures on different radar devices. The Trans-CNN model was trained, and the accuracy on the test set was 97.1%.

[0173] The confusion matrix of the ADT6101 radar for recognizing four gestures was plotted on the test set as shown in Figure 14 . It can be seen from the confusion matrix that the recognition accuracy of both the up and down gestures reached above 97%, and the recognition accuracy of the left and right gestures was also close to 97%. Therefore, the gesture recognition method based on the Trans-CNN network model proposed in the present invention can achieve a relatively high recognition accuracy on any simple millimeter-wave radar, having good application value.

Claims

1. A millimeter wave radar gesture recognition method based on the Trans-CNN model, characterized by: The steps include: (1) Place the millimeter-wave radar horizontally to collect gesture ADC data from different people, real application environments, and different postures; (2) Analyze and process the gesture ADC data, obtain the micro-Doppler characteristics of the gesture, and establish a point cloud dataset of the gesture; (3) The micro-Doppler features of each gesture action sample in the point cloud dataset are unified, and sliced, deleted, and zeroed to obtain a 30*45*5 three-dimensional mixed feature tensor; (4) A gesture recognition model Trans-CNN is constructed using a multi-head self-attention mechanism and a one-dimensional convolutional neural network. The model receives a three-dimensional mixed feature tensor as input, optimizes the model weights during training and testing, and finally obtains a trained Trans-CNN model. (5) Collecting ADC data of gesture actions in real time by connecting to a millimeter-wave radar, and obtaining a three-dimensional mixed feature tensor having micro-Doppler characteristics of gesture actions through steps (2) and (3); inputting the feature tensor into a pre-trained model for gesture recognition; if a preset gesture action is predicted, prompting the user to correctly input the gesture preset by the model; otherwise, continuing to step (5); (6) Send the recognized gesture action to the computer.

2. The millimeter wave radar gesture recognition method based on the Trans-CNN model according to claim 1 is characterized in that: The gesture ADC data collection in step (1) is to connect the millimeter wave radar to the PC through a USB interface. The millimeter wave radar is any 2-transmit 4-receive linear frequency modulated continuous wave radar. The host computer software is used to send the configuration file through the serial port. The frequency modulated continuous wave is transmitted using the TDM-MIMO scheme. After gesture reflection, the echo data is obtained, which is mixed with the transmission signal to obtain the intermediate frequency IF signal. The ADC data can be obtained by ADC sampling.

3. The millimeter wave radar gesture recognition method based on the Trans-CNN model according to claim 1, characterized in that: The step (2) of analyzing and processing the gesture ADC data to obtain the micro-Doppler characteristics of the gesture action includes the following steps: 1) Decode the received LVDS data into complex form, reassemble it according to the number of antennas, chirp number and sampling points, perform 1D-FFT transformation on all sampling points of each chirp, and obtain the distance feature of the gesture action; 2) For the data obtained from the distance feature, the echo data generated by the static object background in the gesture data is filtered out by means of a vector mean elimination algorithm, so as to obtain the gesture action data after filtering out the static object and the background; 3) For the gesture action data obtained by filtering out static objects and background, perform 2D-FFT transformation on each sampling point of all Chirps in the speed dimension to obtain the speed characteristics of the gesture action; 4) After obtaining the velocity characteristics, the CA-CFAR algorithm is used to estimate the dynamic clutter in the data and remove the clutter within the threshold; 5) For the data of distance and speed characteristics, perform Angler-FFT transformation in the receiving antenna dimension to obtain the azimuth characteristics of the gesture action.

4. The millimeter wave radar gesture recognition method based on the Trans-CNN model according to claim 1, characterized in that: The step (2) of establishing the point cloud data set of the gesture action is to design a peak grouping algorithm for the data of the acquired micro-Doppler features, group the internal target points of each gesture sample after the dynamic clutter is filtered out, retain the target points that meet the requirements, and delete the target points that meet the requirements, and finally obtain the point cloud data set of the gesture; the time frame of each gesture data varies according to the gesture, but is between 20-30 frames. At the data transmission rate of the serial communication of 115200bps, the data collection and transmission time of each gesture is 2-3s.

5. The millimeter wave radar gesture recognition method based on the Trans-CNN model according to claim 3 is characterized in that: In step 1), the received LVDS data is decoded into a complex form by sequentially taking out the data of the four receiving antennas corresponding to each transmitting antenna, splicing them according to the number of antennas, the number of Chirps and the number of sampling points, and finally obtaining 8 (antenna) * 480 (Chirp) * 256 (sampling points) three-dimensional matrix data; 1D-FFT transformation is performed on all sampling points of each Chirp of each antenna of this matrix to obtain the distance feature of the gesture action, and the expression is: Where: B is the bandwidth; T c is the pulse width; f IF is the frequency of the intermediate frequency signal; c is the speed of light.

6. The millimeter wave radar gesture recognition method based on the Trans-CNN model according to claim 3 is characterized in that: In step 2), the distance feature data is obtained by filtering out the echo data generated by the static object background in the gesture data using the vector mean elimination algorithm, and the expression is: F1D_StaticOut(N c ,N s ,N t_r )=F1D(n c ,:,n t_r )-avg (3) Where: F1D(N c ,N s ,N r_t ) is the distance gesture data obtained by 1D-FFT calculation; N c is the number of chirps, N s The ADC sampling rate for each Chirp, N t_r is the product of the number of transmitting and receiving antennas; n c is the cyclic factor of the Chirp number; n t_r is the cyclic factor of the number of antennas; avg is the vector mean of all sampling points of each Chirp; F1D_StaticOut is the gesture data after filtering out static objects and background echoes.

7. The millimeter wave radar gesture recognition method based on the Trans-CNN model according to claim 3 is characterized in that: Step 3) The gesture action data obtained by filtering out static objects and background is subjected to 2D-FFT transformation in the velocity dimension for each sampling point of all Chirps, and the expression is: Δω=2πΔf d T c (5) Where: v is the speed of the target; λ is the wavelength of FMCW; Δω is the phase change of the two chirps; Δf d =-2vf0 / c is the Doppler frequency shift caused by gesture movement.

8. The millimeter wave radar gesture recognition method based on the Trans-CNN model according to claim 3 is characterized in that: Step 4) After obtaining the velocity feature, the CA-CFAR algorithm is used to estimate the dynamic clutter in the data and remove the clutter within the threshold, which is to adaptively adjust the decision threshold according to the dynamic clutter in the echo signal; CFAR detection first assumes that the dynamic clutter is known, and then uses the reference unit in the data to estimate the parameters such as clutter noise, so as to obtain a threshold based on the estimated density; if there is a value exceeding the threshold in the detected data, it is considered that there is a target: Where: H is the threshold of the unit to be tested; K is the threshold of clutter estimation; N is the length of the distance dimension or speed dimension; M is the length of the protection unit; x i / x j It is the distance dimension or speed dimension input data.

9. The millimeter wave radar gesture recognition method based on the Trans-CNN model according to claim 3 is characterized in that: Step 5) performs Angler-FFT transformation in the receiving antenna dimension to obtain the horizontal angle feature of the gesture action, which is to first perform FFT calculation in the distance dimension and then perform FFT calculation of the target angle in the antenna dimension. The expression is: Where: S(θ) is the azimuth characteristic distribution of the target; X(m) is the complex signal received by the mth antenna; d is the antenna element spacing; λ is the wavelength of the radar signal; θ is the azimuth angle; m is the antenna index.

10. The millimeter wave radar gesture recognition method based on the Trans-CNN model according to claim 1, characterized in that: The construction process of the gesture recognition model Trans-CNN described in 4) includes: 1) The Trans-CNN model mainly consists of three parts: attention module, serial one-dimensional convolution module and fully connected layer; 2) The attention module mainly consists of two parts: temporal position encoding and calculation of dependencies between sequence targets; 3) The serial one-dimensional convolution module consists of two one-dimensional convolution layers connected in series through an average pooling layer; 4) Perform a global average pooling operation on the output, and finally use a fully connected layer to output the probability of each gesture using softmax; 5) The 30*45*5 three-dimensional mixed feature tensor obtained in step (3) is input into the constructed Trans-CNN model, and the model is trained and verified until the model converges.