Radar gesture recognition method based on convolutional neural network

By using convolutional neural networks, adaptive feature fusion convolution modules and improved CBAM modules in radar gesture recognition, the problems of long running time and many parameters in the prior art are solved, and the recognition effect of high precision and low complexity is achieved.

CN119992649AActive Publication Date: 2025-05-13HEBEI UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510044668.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-12
Publication Date
2025-05-13
Estimated Expiration
2045-01-12

AI Technical Summary

Technical Problem

While the existing radar-based gesture recognition method achieves high recognition accuracy, it has a long running time and many model parameters, making it not suitable for landing applications.

Method used

The radar gesture recognition method based on convolutional neural network is adopted, combining the adaptive feature fusion convolution module (AFFM) and the improved CBAM module to extract and fuse features of different receptive fields to reduce background interference and improve recognition accuracy.

Benefits of technology

The recognition accuracy rate is achieved by more than 99%, while reducing the running time and model parameters, making the method more suitable for real-time applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992649A_ABST
    Figure CN119992649A_ABST
Patent Text Reader

Abstract

The invention relates to a radar gesture recognition method based on a convolutional neural network, and the method comprises the following steps: obtaining a radar gesture signal, and carrying out the preprocessing of the radar gesture signal, and obtaining a radar gesture data set composed of images which can be inputted into the convolutional neural network; constructing a convolutional neural network, wherein the convolutional neural network comprises an adaptive feature fusion convolution module AFFM, a plurality of adaptive mean pooling and a plurality of 3 * 3 convolution layers, and an improved CBAM module; a radar gesture feature map is input to an AFFM, the radar gesture feature map is input into an improved CBAM module after being processed by a 3 * 3 convolutional layer and self-adaptive mean pooling, the input of the improved CBAM module is the output of superior self-adaptive mean pooling, and after the output of the improved CBAM module and the output of superior self-adaptive mean pooling are subjected to element-by-element multiplication operation, the output of the improved CBAM module and the output of superior self-adaptive mean pooling are output. And carrying out adaptive mean pooling and linear layer processing to obtain a radar gesture recognition result. According to the method, the recognition accuracy is effectively improved under the condition of not increasing the parameter quantity and the running time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of radar gesture recognition, and in particular to a radar gesture recognition method based on a convolutional neural network. Background Art

[0002] With the rapid development of smart wearable devices and mobile devices, human-computer interaction (HCI) has become a hot topic in the past decade. Although traditional contact physical devices have high measurement accuracy, they have the disadvantage of poor convenience. Non-contact devices can bring more freedom and convenience to users. As one of the intuitive and effective methods in non-contact HCI, hand gesture recognition (HGR) stands out.

[0003] The mainstream HGR methods today can be divided into two types: one is optical-based HGR, and the other is radar-based HGR. Optical-based HGR methods include cameras, infrared and other methods. The measurement accuracy of these methods is easily affected by external environments such as light and temperature, and the camera-based method needs to collect user images before identifying gestures, which may cause user privacy leakage. The radar-based HGR method can well protect user privacy and is not affected by the environment. It is low-cost, low-power, and has a fast processing speed. It is suitable for gesture recognition in various environments.

[0004] Existing radar-based HGR methods also have defects. Some scholars use three-dimensional convolutional neural networks to realize gesture recognition, but the number of parameters is too large and the running time is long; some scholars use LSTM or Transformer architecture, but this method requires a multi-layer architecture to improve the accuracy, resulting in a long running time; the hybrid model complicates the model, increasing the running time and the number of parameters. In short, if the current method wants to experiment with high recognition accuracy, its running time is too long and the model parameters are too many, which is not conducive to the implementation of gesture recognition. Therefore, it is urgent to provide a radar gesture recognition method that can achieve high-precision recognition while ensuring lower running time and fewer model parameters. Summary of the invention

[0005] In view of the shortcomings of the prior art, the technical problem to be solved by the present invention is to provide a radar gesture recognition method based on convolutional neural network. The recognition method is equipped with an adaptive feature fusion convolution module and an improved CBAM module, and its recognition accuracy reaches more than 99%, and the running time and parameter quantity are relatively small.

[0006] The technical solution adopted by the present invention to solve the technical problem is:

[0007] A radar gesture recognition method based on a convolutional neural network, the recognition method comprising the following contents:

[0008] Acquire radar gesture signals and perform preprocessing to obtain a radar gesture dataset consisting of images that can be input into a convolutional neural network;

[0009] Constructing a convolutional neural network, wherein the convolutional neural network includes an adaptive feature fusion convolution module AFFM, multiple adaptive mean pooling and multiple 3×3 convolution layers, and an improved CBAM module;

[0010] The radar gesture feature map is input into the adaptive feature fusion convolution module AFFM, and then input into the improved CBAM module after a 3×3 convolution layer and adaptive mean pooling. The input of the improved CBAM module is the output of the upper adaptive mean pooling. The output of the improved CBAM module is multiplied element by element with the output of the upper adaptive mean pooling, and then processed by an adaptive mean pooling and linear layer to obtain the radar gesture recognition result.

[0011] The adaptive feature fusion convolution module AFFM includes an instance normalization layer. After the input radar gesture feature map is processed by the instance normalization layer, it is divided into five branches for processing, wherein the first branch is a 1×1 convolution operation and a ReLU function, the second branch is a 3×3 convolution operation and a ReLU function, the third branch is a 5×5 convolution operation and a ReLU function, the fourth branch is a 1×1 convolution operation, a ReLU function and a maximum pooling, and the fifth branch is a 1×1 convolution operation and a Sigmoid function; the output features of the first four branches are processed by feature splicing along the channel dimension and then element-by-element multiplication with the output of the fifth branch; and then the final fusion features are output after being processed by a batch normalization layer and an adaptive mean pooling layer;

[0012] The improved CBAM module includes a spatial attention mechanism and an improved channel attention mechanism. The improved channel attention mechanism introduces a learnable parameter α, and multiplies the output features of the maximum pooling and average pooling in the channel attention mechanism by 1-α and α respectively, and then adds and sums them. The summation result is integrated with features using a 1×1 grouped convolution, where the number of groups is the number of input channels, and finally processed by a Sigmoid function to obtain the channel attention C;

[0013] The convolutional neural network is trained with the radar gesture data set, and gesture recognition is performed with the trained convolutional neural network.

[0014] Furthermore, the radar gesture dataset is obtained after image transformation of the public gesture dataset in the IEEE data port. The public gesture dataset in the IEEE data port contains 12 gestures consisting of a total of 4609 gesture records. The length of these gestures ranges from 5 to 81 frames. Each frame contains information of up to 80 detection points, and each detection point contains corresponding distance, speed, x-coordinate, y-coordinate and signal amplitude; the image transformation is to transform the data into the format of number of frames*detection points=80*80, and for frames and detection points less than 80, 0 is used to fill.

[0015] Furthermore, during the training process, random initialization is used for parameter initialization, and the loss function is the cross entropy loss function. In the network training, small batch gradient descent and Adam optimization algorithm are used to adjust the network parameters, where the batch size is 8.

[0016] Furthermore, the image that can be input into the convolutional neural network is high-dimensional channel image data, and the number of channels is greater than 3.

[0017] Furthermore, the recognition accuracy of the recognition method is not less than 99.0%, and the time complexity is within 2.2s while ensuring high recognition accuracy, and the model parameter amount Params is within 4.0M.

[0018] The present invention also protects a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of the identification method can be implemented.

[0019] Compared with the prior art, the present invention has the following beneficial effects:

[0020] The convolutional neural network constructed by the method of the present invention uses an adaptive feature fusion convolution module (AFFM) in the input layer to extract and fuse features of different receptive fields, capture and extract common features of the same gestures at different speeds and positions, reduce the influence of background interference, and eliminate the influence caused by gesture differences. An improved CBAM module is used before the output layer to improve the recognition accuracy while reducing the number of model parameters, and radar gesture signals can be recognized with high efficiency and accuracy.

[0021] The model in the present invention is significantly better than existing methods such as LSTM, Transformer, and hybrid models in terms of running time. In terms of parameter quantity, it is significantly less than that of three-dimensional convolutional neural networks and hybrid models. The present invention can effectively increase the recognition accuracy without increasing the parameter quantity and running time.

[0022] The outstanding essential features of the present invention are:

[0023] a) An innovative adaptive feature fusion convolution module (AFFM) is proposed, which comprehensively utilizes the multi-scale fusion strategy and residual weights to capture and extract the common features of the same gestures at different speeds and positions by extracting and fusing the features of different receptive fields, thereby improving the recognition accuracy of the model.

[0024] b) The overall network of the present invention significantly improves the accuracy of gesture recognition, reaching more than 99% recognition accuracy for 12 gestures, and has lower time complexity and space complexity than most existing methods. A radar gesture recognition model based on deep learning is used to classify and recognize radar gesture data that has been imaged. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A schematic diagram of the structure of a convolutional neural network according to an embodiment of the present invention.

[0026] Figure 2 Schematic diagram of the structure of the adaptive feature fusion convolution module AFFM.

[0027] Figure 3 Schematic diagram of the structure of a traditional CBAM module.

[0028] Figure 4 Schematic diagram of the structure of the improved channel attention mechanism in the present invention.

[0029] Figure 5 Schematic diagram of 12 gestures in the public dataset. DETAILED DESCRIPTION

[0030] The present invention is further explained below in conjunction with the embodiments and drawings, but this is not intended to limit the scope of protection of the present application.

[0031] The task definition of radar gesture recognition in the present invention is:

[0032] Radar gesture recognition is essentially a classification problem. After the collected radar gesture signal is processed using a signal processing method, the network input X is obtained. The convolutional neural network in the present invention can essentially be regarded as a function F(.),

[0033] Y=F(X) (1)

[0034] Through F, the input X can be classified to obtain the category Y of the gesture, thereby recognizing the gesture.

[0035] The overall structure of the convolutional neural network in the present invention is as follows: Figure 1 As shown, the convolutional neural network includes an adaptive feature fusion convolution module AFFM, multiple adaptive mean pooling and multiple 3×3 convolution layers, and an improved CBAM module;

[0036] The radar gesture feature map is input into the adaptive feature fusion convolution module AFFM, and then input into the improved CBAM module after a 3×3 convolution layer and adaptive mean pooling. The input of the improved CBAM module is the output of the upper adaptive mean pooling. The output of the improved CBAM module is multiplied element-by-element with the output of the upper adaptive mean pooling, and then processed by an adaptive mean pooling and linear layer to obtain the radar gesture recognition result.

[0037] AFFM preprocesses the input feature map to extract gesture information and reduce the size of the feature map while overcoming gesture differences and background interference. It uses the improved CBAM module to select feature maps before the fully connected layer, highlighting important information that is beneficial to classification while suppressing unimportant information, thereby improving the accuracy of gesture recognition, with fewer parameters and faster operation.

[0038] Specifically, the structure of AFFM is as follows: Figure 2 As shown in the figure, including the instance normalization layer, the input radar gesture feature map is processed by the instance normalization layer and then divided into five branches for processing, wherein the first branch is a 1×1 convolution operation and a ReLU function, the second branch is a 3×3 convolution operation and a ReLU function, the third branch is a 5×5 convolution operation and a ReLU function, the fourth branch is a 1×1 convolution operation, a ReLU function and a maximum pooling, and the fifth branch is a 1×1 convolution operation and a Sigmoid function; the output features of the first four branches are processed by feature concatenation along the channel dimension and then element-by-element multiplication with the output of the fifth branch; after that, the final fusion features are output after being processed by a batch normalization layer and an adaptive mean pooling layer.

[0039] Due to the different habits of gesture performers, the position and speed of gestures vary. Even the same gesture made by the same person in different situations may have significant differences in speed, position, etc. This difference is manifested as data duplication, offset or missing on the input image X. The adaptive feature fusion convolution module AFFM can effectively reduce the interference caused by gesture differences and external interference information.

[0040] Assuming that the processed radar signal is X, we first transform X using the instance normalization layer to pull the data with distribution deviation back to the standard distribution and eliminate the differences introduced by the different positions of the gesture. Then, we use multiple branches to extract features. These branches have different convolution kernels (including 1×1, 3×3 and 5×5). We use multiple branches to capture the feature representation of the input data under different receptive fields. Figure 2 After extracting features, the four branches on the left are spliced ​​along the channel dimension to obtain the spliced ​​feature X f .X fThe complementary information extracted from different branches is integrated to further improve the accuracy of HGR, and this multi-branch structure reduces the parameters of the model. Figure 2 The rightmost 1×1 convolution is used to transform the input data and process it using the Sigmoid activation function to generate the residual weight X w After and X f Perform element-by-element multiplication to obtain the weighted fusion feature X z The batch normalization layer is then used to transform the data to reduce the time complexity of the model, and the adaptive mean pooling is used to reduce the input size to 1 / 2 of the original size, while retaining sufficient feature information and reducing the time complexity of the model, and outputting the final fusion features.

[0041] The spatial channel attention mechanism CBAM includes the channel attention mechanism and the spatial attention mechanism, and its block diagram is as follows Figure 3 As shown. The spatial channel attention mechanism is divided into two stages: the input feature format is N*H*W, where N is the number of channels, and H and W are the height and width of the matrix respectively. The channel attention mechanism compresses the input in the spatial dimension (compresses H*W to 1), that is, the input features are obtained by parallel maximum pooling and average pooling, and then sent to the shared fully connected layer to integrate the features and sum them up, and then processed using the Sigmoid activation function to generate the channel attention C, and C is multiplied by the input element by element as the input Z1 of the spatial attention mechanism; the spatial attention mechanism compresses the input Z1 in the channel dimension (compresses N to 1), and according to the mean pooling and maximum pooling, the features are processed into two one-channel feature maps and then spliced, and the convolution layer is used to extract the features and then the Sigmoid activation function is used to process the spatial attention S, and S is multiplied by Z1 element by element to obtain the final output feature map Z.

[0042] This embodiment improves the channel attention mechanism. The output features of the maximum pooling and average pooling in the channel attention mechanism are multiplied by 1-α and α respectively, and then added and summed. The summation result is integrated with features using 1×1 grouped convolution. The number of groups is the number of input channels. Finally, the channel attention C is obtained by Sigmoid function processing. The specific structure of the improved channel attention mechanism is shown in Figure 4 The learnable parameter α is automatically determined during the training process.

[0043] The calculation process of the improved channel attention mechanism is: perform maximum pooling and average pooling operations on the input features to generate two feature maps Z max With Z avg , use the learnable parameter α to perform weighted summation on the two feature maps, then use 1*1 group convolution to integrate the features and then use the Sigmoid activation function to generate the channel attention C. The calculation process is expressed as:

[0044] C = σ{conv1[α.Z avg +(1-α).Z max ]} (2)

[0045] Among them, σ represents the Sigmoid operation, and conv1 represents 1*1 group convolution.

[0046] The spatial attention mechanism operates as follows Figure 3 As shown in (b), this is the existing structure. The specific processing process is: apply mean pooling and maximum pooling to the channel dimension to obtain two one-channel feature maps Z a With Z m , concatenate the two feature maps in the channel dimension to obtain a two-channel feature map Z2. Apply a two-dimensional convolution operation to Z2 and then use the Sigmoid activation function to obtain the final spatial attention S. The calculation formula is as follows:

[0047] SS=σ{conv2[Z avg ,Z max ]} (3)

[0048] Among them, conv2 represents a two-dimensional convolution operation, [] represents the concatenation of two feature maps according to the channel dimension, and σ represents the Sigmoid operation.

[0049] The improved CBAM module is used to assign weights to feature maps, highlighting important information in feature maps and suppressing unimportant information, thereby effectively improving the recognition accuracy of the model.

[0050] In the present invention, the radar gesture signal is preprocessed to obtain a radar gesture data set composed of images that can be input into a convolutional neural network. The input feature X is a radar gesture feature map, which can be image data of any channel, or high-dimensional channel image data (higher than 3 channels). The input format is batch*number of channels*height*width.

[0051] Example 1

[0052] 1) Obtain multiple gesture records in the radar gesture dataset, divide all gesture records into training set and test set, and preprocess each gesture record to obtain a multi-channel image.

[0053] 2) Construct a convolutional neural network for gesture recognition, which includes an adaptive feature fusion convolution module, an improved CBAM module and three convolution layers, three ReLU activation functions, three batch normalization layers, three adaptive mean pooling layers and a linear layer.

[0054] 3) Input the training set into the neural network model constructed in step 2) for training.

[0055] 4) After preprocessing the radar gesture records to be recognized, the trained neural network model is used to perform gesture recognition to obtain the gesture classification results.

[0056] The pretreatment process in 1) and 4) is specifically as follows:

[0057] The two-dimensional point cloud data of the radar is read into a high-dimensional matrix form, and divided into subsets according to the distance, speed, x-coordinate, y-coordinate, signal strength and other information in the point cloud information, and each type of information is processed into a matrix format of the number of frames × the number of sampling points. Then, the matrices containing various types of information are spliced ​​to obtain multi-channel image data.

[0058] The model structure of the convolutional neural network in step 2) is specifically:

[0059] The adaptive feature fusion convolution module is the first structure of the network, and the subsequent connections are: the first layer of convolution, the first layer of adaptive mean pooling layer, the second layer of convolution, the third layer of convolution, the second layer of adaptive mean pooling layer, the improved CBAM module, the third layer of adaptive mean pooling layer, and the linear layer; the output of the adaptive feature fusion convolution module is connected to the input of the first layer of convolution, the output of the first layer of convolution module is connected to the input of the first layer of adaptive mean pooling layer after the ReLU activation function and the batch normalization layer, the output of the first layer of adaptive mean pooling layer is connected to the input of the second layer of convolution, and the output of the second layer of convolution is connected to the input of the second layer of convolution. The output is connected to the input of the third convolution layer after passing through the ReLU activation function and the batch normalization layer. The output of the third convolution layer is connected to the input of the second adaptive mean pooling layer after passing through the ReLU activation function and the batch normalization layer. The output of the second mean pooling layer is connected to the input of the improved CBAM module. The output of the improved CBAM module and the output of the second mean pooling layer are weighted and connected to the input of the third adaptive mean pooling layer. The output of the third mean pooling layer is connected to the input of the linear layer. The output of the linear layer is processed by Softmax to obtain a probability distribution, thereby obtaining the classification result.

[0060] The adaptive feature fusion convolution module includes an instance normalization layer, three layers of 1×1 convolution, one layer of 3×3 convolution, one layer of 5×5 convolution, a maximum pooling layer, a batch normalization layer and an adaptive mean pooling layer. The output of the instance normalization layer is connected to the input of all the convolutions mentioned above. The first layer of 1×1 convolution is followed by a ReLU activation function and is used as part of the splicing feature; the 3×3 convolution is followed by a ReLU activation function and is used as part of the splicing feature; the 5×5 convolution is followed by a ReLU activation function and is used as part of the splicing feature; the second layer of 1×1 convolution is followed by a ReLU activation function, and the output of the activation function is used as the input of the maximum pooling layer, and the output of the maximum pooling layer is used as part of the splicing feature. The third layer of 1×1 convolution is used as a residual weight, and a Sigmoid activation function is used after the convolution layer to assign weights to the splicing features to obtain weighted fusion features. The weighted fused features are used as the input of the batch normalization layer, and the batch normalization layer is connected to the adaptive mean pooling layer. The result of the adaptive mean pooling layer is the result of the adaptive feature fusion convolution module.

[0061] In the convolutional neural network, the first convolution layer contains a 3×3 two-dimensional convolution, followed by a ReLU activation function and a batch normalization layer. The padding of the convolution is (1,1), the stride is (1,1), and the input and output channels are 128 and 256 respectively.

[0062] The second convolution layer consists of a 3×3 2D convolution, followed by a ReLU activation function and a batch normalization layer. The padding of the convolution is (1,1), the stride is (1,1), and the input and output channels are 256 and 512 respectively.

[0063] The third convolution layer consists of a 3×3 two-dimensional convolution, followed by a ReLU activation function and a batch normalization layer. The convolution has a padding of (0,0), a stride of (1,1), and input and output channels of 512 and 512 respectively.

[0064] The first and second layers of adaptive mean pooling both reduce the size of the feature map to half of its original size, and the third layer of mean pooling reduces the size of the feature map to 1×1.

[0065] The improved CBAM module includes a maximum pooling layer at the spatial level, a maximum pooling layer at the channel level, a mean pooling layer at the spatial level, a mean pooling layer at the channel level, a 1×1 convolution layer, and a 1×1 group convolution layer. The input feature map is passed through the maximum pooling layer at the spatial level and the mean pooling layer at the spatial level to obtain two feature maps. The two feature maps are weighted and summed using the trainable parameter α to obtain a new feature map. This feature map is used as the input of the 1×1 group convolution. The output of the 1×1 group convolution is transformed by the Sigmoid transformation to obtain the channel attention C. The channel attention mechanism is passed through the mean pooling layer at the channel level and the maximum pooling layer at the channel level. The two feature maps are concatenated to obtain a two-channel input feature map. This feature map is used as the input of the 1×1 convolution layer, and the output is transformed by the Sigmoid activation function to obtain the final spatial attention S.

[0066] Example 2

[0067] The radar gesture recognition method based on the convolutional neural network in this embodiment recognizes and classifies radar gestures processed into five-channel image data. The specific steps are:

[0068] The dataset used is the public gesture dataset in the IEEE data port. The dataset contains a total of 4609 gesture records consisting of 12 types of gestures, such as Figure 5 As shown in the figure, they are arm left, arm right, hand away, hand close, arm raised, arm down, palm up, palm down, hand left, hand right, horizontal fist and vertical fist. Each gesture corresponds to several data files stored in CSV format. The length of these gestures ranges from 5 to 81 frames. Each frame contains point cloud data information of up to 80 detection points. Each detection point contains the corresponding distance, speed, x-coordinate, y-coordinate and signal amplitude. The number of channels is 5, that is, the size is 80*5*80. The point cloud image data of the dataset is transformed, and the data is organized into the format of frame number*detection point (80*80) according to different information types. For frames and detection points less than 80, 0 is used to fill in, so as to obtain the radar gesture dataset.

[0069] The radar gesture dataset was used to train the above convolutional neural network. During the training process, the random initialization method was used to initialize the network parameters. The loss function used in the training was the cross entropy loss function. The mini-batch gradient descent and Adam optimization algorithms were used to adjust the network parameters, with a batch size of 8. The radar gesture dataset was randomly divided into a training set, a validation set, and a test set, with a ratio of 6:2:2.

[0070] The training process is carried out on a personal computer, and the specific configuration is shown in Table 1.

[0071] Table 1 Experimental environment

[0072] Parameter name Parameter Value operating system Windows 10 GPU Graphics Card NVIDIARTX1050Ti(2G) CUDA Version 11.4 PyTorch Version 1.12 Programming Environment Python 3.8

[0073] During the training process, three indicators are set: recognition accuracy Acc, time complexity Time, and model parameter quantity Params. The calculation method of Acc is:

[0074]

[0075] N T is the number of correctly identified, and N is the total number of test sets. The Time calculation rule starts with reading the file and ends with loading the model to process the file and get the recognition result. To avoid contingency, when calculating the time complexity of the model, each model is run 1000 times and the average value is used as the time complexity, in seconds (s). Params is the number of floating-point numbers required for storage, expressed in millions (M).

[0076] The internal processing flow of the convolutional neural network is as follows: the input X is sent to the constructed convolutional neural network, and the input is first weighted feature fused using AFFM, and the size of the input feature is reduced to obtain feature Y. A convolution operation with a convolution kernel size of 3*3 is performed on Y, and adaptive mean pooling is used to obtain the intermediate feature map Y1.

[0077] Two convolution operations with a kernel size of 3*3 are performed on Y1, and adaptive mean pooling is used to obtain the intermediate feature map Y2. After applying the improved CBAM module to the feature map Y2, it is multiplied with Y2 to obtain Y3, which highlights the useful information in the feature map and suppresses useless information and background noise.

[0078] Adaptive mean pooling is used to downsample Y3, which reduces the size of the feature map without changing the feature information in the feature map. The downsampled Y3 is flattened to obtain Y4, that is, the original format of batch*channels*height*width is changed to batch*(channels*height*width).

[0079] Using the linear layer, the weights and biases between different features of Y4 are calculated. The linear layer can output the probability of each category to obtain the classification result Z.

[0080] In order to verify the effectiveness of the model proposed in the present invention, this embodiment adopts the following comparative models:

[0081] a) DCS-CTN: radar gesture classification using 3DCNN and Transformer.

[0082] b)DFDRN: Recognize radar gestures using dual-stream fusion of 2DCNN and 3DCNN.

[0083] c) S3D: A model proposed for video recognition, suitable for frame data, and improved by adding attention blocks.

[0084] d) 3DCNN+LSTM: After using 3DCNN to initially extract features, LSTM is used for gesture recognition.

[0085] e) CMFF-HGR: Gesture recognition using residual neural network with multi-stream fusion.

[0086] The experimental environment for training the comparison model is the same as that in Table 1. The results obtained from the comparison experiment are shown in Table 2, which records the specific information of the three indicators of the experimental setting. The recognition accuracy is expressed in percentage, the time complexity is in seconds, and the number of model parameters is expressed in millions. By comparison, it can be seen that compared with other models, the model proposed in the present invention has higher recognition accuracy, lower time complexity than most models, and lower network parameter quantity than most models. The model has higher recognition accuracy and lower time and space complexity, and is more suitable for the practical application of real-time radar gesture recognition system.

[0087] Table 2 Model comparison experimental results

[0088] Network Model Acc(%) Time(s) Params(M) DCS-CTN 97.8 2.30 6.23 DFDRN 96.7 2.71 25.64 S3D 95.6 2.38 20.44 3DCNN+LSTM 95.5 2.49 5.63 CMFF-HGR 94.0 0.92 58.64 The present invention 99.3 2.16 3.87

[0089] In order to verify the effectiveness of the module proposed in the present invention, the present invention conducted an ablation experiment on the above-mentioned radar gesture dataset, and the results are shown in Table 3. In the ablation experiment, in order to verify the effectiveness of the AFFM and improved CBAM modules, ablation experiments were performed on these two modules respectively. The experiment includes 4 groups in total, namely, AFFM with improved CBAM module; AFFM with non-improved CBAM module (that is, conventional CBAM module); without AFFM, with improved CBAM module and without AFFM, without improved CBAM module. When there is no AFFM, the convolutional layer is used instead, and when there is no improved CBAM module, the spatial channel attention mechanism (CBAM module) is used. It can be seen from the experimental results that the proposed AFFM and improved CBAM modules have effectively improved the recognition accuracy, and the improved CBAM module has reduced the time and parameter amount compared with before the improvement, which is more conducive to the realization of gesture recognition.

[0090] Table 3 Ablation experiment results

[0091]

[0092] Any matters not described in the present invention are applicable to the prior art.

Claims

1. A radar gesture recognition method based on convolutional neural network, characterized in that: The identification method includes the following contents: Acquire radar gesture signals and perform preprocessing to obtain a radar gesture dataset consisting of images that can be input into a convolutional neural network; Constructing a convolutional neural network, wherein the convolutional neural network includes an adaptive feature fusion convolution module AFFM, multiple adaptive mean pooling and multiple 3×3 convolution layers, and an improved CBAM module; The radar gesture feature map is input into the adaptive feature fusion convolution module AFFM, and then input into the improved CBAM module after a 3×3 convolution layer and adaptive mean pooling. The input of the improved CBAM module is the output of the upper adaptive mean pooling. The output of the improved CBAM module is multiplied element by element with the output of the upper adaptive mean pooling, and then processed by an adaptive mean pooling and linear layer to obtain the radar gesture recognition result. The adaptive feature fusion convolution module AFFM includes an instance normalization layer. After the input radar gesture feature map is processed by the instance normalization layer, it is divided into five branches for processing, wherein the first branch is a 1×1 convolution operation and a ReLU function, the second branch is a 3×3 convolution operation and a ReLU function, the third branch is a 5×5 convolution operation and a ReLU function, the fourth branch is a 1×1 convolution operation, a ReLU function and a maximum pooling, and the fifth branch is a 1×1 convolution operation and a Sigmoid function; the output features of the first four branches are processed by feature splicing along the channel dimension and then element-by-element multiplication with the output of the fifth branch; and then the final fusion features are output after being processed by a batch normalization layer and an adaptive mean pooling layer; The improved CBAM module includes a spatial attention mechanism and an improved channel attention mechanism. The improved channel attention mechanism introduces a learnable parameter α, and multiplies the output features of the maximum pooling and average pooling in the channel attention mechanism by 1-α and α respectively, and then adds and sums them. The summation result is integrated with features using a 1×1 grouped convolution, where the number of groups is the number of input channels, and finally processed by a Sigmoid function to obtain the channel attention C; The convolutional neural network is trained with the radar gesture data set, and gesture recognition is performed with the trained convolutional neural network.

2. The identification method according to claim 1, characterized in that: The radar gesture dataset is obtained after image transformation from the public gesture dataset in the IEEE data port. The public gesture dataset in the IEEE data port contains 12 gestures consisting of a total of 4609 gesture records. The length of these gestures ranges from 5 to 81 frames. Each frame contains information of up to 80 detection points, and each detection point contains the corresponding distance, speed, x-coordinate, y-coordinate and signal amplitude. The image transformation is to transform the data into the format of number of frames * detection points = 80 * 80. For frames and detection points less than 80, 0 is used to fill them.

3. The identification method according to claim 1, characterized in that: During the training process, random initialization is used for parameter initialization, and the loss function is the cross entropy loss function. In network training, small batch gradient descent and Adam optimization algorithm are used to adjust the network parameters, and the batch size is 8.

4. The identification method according to claim 1, characterized in that: The image that can be input into the convolutional neural network is high-dimensional channel image data, and the number of channels is greater than 3.

5. The identification method according to claim 1, characterized in that: The recognition accuracy of the recognition method is not less than 99.0%. At the same time, the time complexity is within 2.2s while ensuring high recognition accuracy, and the model parameter amount Params is within 4.0M.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the identification method described in claims 1-5 can be implemented.

Citation Information

Patent Citations

  • Active defense detection method based on face key point watermark

    CN117474741A

  • Satellite remote sensing image fire point segmentation method and system based on deep learning

    CN117765411A

  • Camouflage target detection algorithm based on edge refinement and enhancement network

    CN118298282A

  • Method and system for identifying identity of shielded person in complex scene of mine

    CN118711214A

  • A human-machine-interface system

    EP3388981A1