Gesture recognition method based on WIFI-CSI multi-scale feature fusion

By introducing multi-scale feature fusion and lightweight CBAM methods in WiFi-CSI gesture recognition, the problem of insufficient single feature capture capability is solved, and high-precision, lightweight and robust gesture recognition effects are achieved.

CN119939406AActive Publication Date: 2025-05-06TIANJIN UNIV OF TECH & EDUCATION (TEACHER DEV CENT OF CHINA VOCATIONAL TRAINING & GUIDANCE)
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411743021.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-05-06
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

In the existing WiFi-CSI signal deep learning method, a single architecture neural network can only capture a single feature, resulting in gesture recognition tasks in resource-limited and complex scenarios that cannot meet the practical application needs.

Method used

A multi-scale feature fusion gesture recognition method based on WIFI-CSI is adopted. Through the multi-scale feature fusion module, depth-pooling module and gated loop unit module in the feature extraction module, multi-scale features are extracted and fused, combined with lightweight CBAM and parallel network structure, the representation ability and robustness of the gesture recognition model are improved.

Benefits of technology

High-precision, lightweight and robust gesture recognition is achieved, which improves the recognition accuracy in complex scenarios and reduces the calculation complexity and parameter quantity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939406A_ABST
    Figure CN119939406A_ABST
Patent Text Reader

Abstract

The invention discloses a gesture recognition method based on WIFI-CSI multi-scale feature fusion, the gesture recognition method is based on a gesture recognition model, and the gesture recognition model comprises a feature extraction module and a gesture recognition module, the feature extraction module comprises a multi-scale feature fusion module 1, a multi-scale feature fusion module 2, a depth-pooling module and a gating circulation unit module, the input of the multi-scale feature fusion module 1 is a sample xinput 1, and each sample is three-dimensional data formed by a velocity spectrum in a human body coordinate system in a time frame dimension. The output of the gesture recognition module is the predicted gesture activity category, the gesture recognition model improves the training efficiency and the prediction accuracy, and the accuracy is 84.30 + / -0.31%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of gesture recognition technology in the field of wireless sensing, and in particular relates to a gesture recognition method based on WIFI-CSI multi-scale feature fusion. Background Art

[0002] In recent years, with the rapid development of communication and perception integration, more and more researchers have invested in the field of wireless perception. Gesture recognition is a hot topic in the field of wireless perception. It is crucial to the development of new human-computer interaction modes and has brought great changes to human life. WiFi perception technology is an emerging technology. It is used to conduct in-depth research in the field of gesture recognition. WiFi perception technology perceives more subtle changes in human activities through the multipath effect of signal propagation. Activity changes are manifested as changes in channel state information (CSI).

[0003] However, most deep learning methods using WiFi-CSI signals use a single-architecture neural network, which can only capture a single feature and has a large number of parameters. For gesture recognition tasks in complex scenarios with limited resources, it cannot meet the actual application requirements. Therefore, a lightweight gesture recognition method based on WiFi-CSI multi-scale feature fusion is urgently needed. Summary of the invention

[0004] In view of the deficiencies in the prior art, an object of the present invention is to provide a gesture recognition model, which can realize the recognition of gesture activity categories and has the advantages of high precision, light weight and strong robustness.

[0005] Another object of the present invention is to provide a method for training a gesture recognition model.

[0006] Another object of the present invention is to provide a gesture recognition method based on WIFI-CSI multi-scale feature fusion.

[0007] The present invention is achieved through the following technical solutions.

[0008] A gesture recognition model includes: a feature extraction module and a gesture recognition module, wherein the feature extraction module includes: a multi-scale feature fusion module 1, a multi-scale feature fusion module 2, a depth-pooling module and a gated recurrent unit module, and the input of the multi-scale feature fusion module 1 is a sample x input1 , the output is the feature vector X1, and the multi-scale feature fusion module 1 is used to extract each sample x input1The input of the multi-scale feature fusion module 2 is the feature vector X1, and the output is the feature vector X2. The multi-scale feature fusion module 2 is used to extract the features in the feature vector X1 and perform multi-scale feature fusion on it. The input of the depth-pooling module is the feature vector X2, and the output is the feature vector X3. The depth-pooling module is used to extract the features of the feature vector X2 and perform dimensionality reduction. The input of the gated recurrent unit module is the feature vector X3, and the output is the feature vector H. The gated recurrent unit module is used to extract the time features of the feature vector X3 to obtain the feature vector H.

[0009] The input of the gesture recognition module is the feature vector H, and the output is the predicted gesture activity category;

[0010] Each sample is three-dimensional data composed of the velocity spectrum in the time frame dimension in the human body coordinate system.

[0011] In the above technical solution, the multi-scale feature fusion module 1 is a fusion structure, which includes: an upper branch structure, a lower branch structure and a summation module. The upper branch structure includes: a hole convolution module, a first batch normalization layer, a PReLU activation function and a first CBAM connected in sequence. The lower branch structure includes: a first depthwise separable convolution module, a second batch normalization layer, a ReLU activation function and a second CBAM connected in sequence. The summation module is used to add the output of the first CBAM and the output of the second CBAM element by element to obtain a feature vector.

[0012] In the above technical solution, the multi-scale feature fusion module 2 includes: a filling module and the fusion structure, the filling module is used to perform a filling operation of padding=1, the filling module is connected to the hollow convolution module in the fusion structure, and the filling module and the first depth-separable convolution module of the multi-scale feature fusion module 2 simultaneously receive the feature vector X1.

[0013] In the above technical solution, the depth-pooling module includes: a second depth-separable convolution module, a padding module, a batch normalization layer, a ReLU activation function, a maximum pooling layer and a first fully connected layer connected in sequence, the padding module is used to perform a padding operation of padding=1, the feature vector X2 is used as the input of the second depth-separable convolution module, and the feature vector X3 is output by the first fully connected layer.

[0014] In the above technical solution, the second depth-separable convolution module includes: two-dimensional depth convolution and two-dimensional point convolution.

[0015] In the above technical solution, the first fully connected layer includes: a first linear transformation layer, a ReLU function and a first Dropout layer, a second linear transformation layer and a ReLU function connected in sequence.

[0016] In the above technical solution, the gesture recognition module includes: a second Dropout layer, a second fully connected layer and a normalized Softmax function connected in sequence.

[0017] A training method for a gesture recognition model includes: taking each sample in a training set as a sample x input1 The gesture recognition model is input with the labels to be trained, and the predicted gesture activity category of each sample is obtained. The gesture recognition model is provided with a loss function calculation module, and the loss function calculation module uses the cross entropy loss function to calculate the loss value. The gesture recognition model optimizes its parameters by using the gradient descent back propagation method, and the training is terminated until the maximum number of iterations is reached, thereby obtaining a trained gesture recognition model.

[0018] A gesture recognition method based on WIFI-CSI multi-scale feature fusion includes: inputting samples of a test set into a trained gesture recognition model to identify gesture activity categories, and obtaining a predicted gesture activity category for each sample.

[0019] Compared with the prior art, the present invention has the following beneficial effects:

[0020] 1. The lightweight CBAM is introduced into the gesture recognition model of the present invention, so that the gesture recognition model can extract key information of channels and spaces, reduce unnecessary features, and enhance the representation ability of the gesture recognition model of the present invention;

[0021] 2. The gesture recognition model of the present invention introduces an element-by-element (corresponding element) addition method to fuse features, thereby increasing the robustness of the gesture recognition model and improving the generalization ability of the gesture recognition model.

[0022] 3. The gesture recognition model of the present invention introduces a parallel network structure to capture different scale features. It can be seen from floating point operations (Flops) and parameter quantities (Params) that the training efficiency of the gesture recognition model of the present invention is improved.

[0023] 4. The gesture recognition method of the present invention has a high accuracy rate in recognizing complex gesture activities, and enhances the robustness of the gesture recognition model while maintaining its lightweight. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 A framework diagram of the invented gesture recognition model;

[0025] Figure 2 A schematic diagram of a fusion structure in the invented gesture recognition model;

[0026] Figure 3 This is a graph showing the prediction results of the gesture recognition model of the present invention on the test set. DETAILED DESCRIPTION

[0027] A gesture recognition method based on WIFI-CSI multi-scale feature fusion of the present invention is described in detail below with reference to the accompanying drawings and embodiments.

[0028] Example 1

[0029] like Figure 1 As shown, a gesture recognition model includes: a feature extraction module and a gesture recognition module, wherein the feature extraction module includes: a multi-scale feature fusion module 1 (D-CBAM-1), a multi-scale feature fusion module 2 (D-CBAM-2), a depth-pooling module and a gated recurrent unit module (GRU, Shuai, G., Yuefei, H., Shuo, Z., Jingcheng, H., Guangqian, W., Meixin, Z., & Qingsheng, L. (2020) Short-term runoff prediction with GRU and LSTM networks without requiring time step optimization during sample generation, Journal of hydrology, 589.), and the input of the multi-scale feature fusion module 1 (D-CBAM-1) is the sample x input1 , the output is the feature vector X1, and the multi-scale feature fusion module 1 is used to extract each sample x input1 The input of the multi-scale feature fusion module 2 (D-CBAM-2) is the feature vector X1, and the output is the feature vector X2. The multi-scale feature fusion module 2 is used to extract the features in the feature vector X1 and perform multi-scale feature fusion on it. The input of the deep-pooling module (DS-Conv Block) is the feature vector X2, and the output is the feature vector X3. The deep-pooling module is used to extract the features of the feature vector X2 and perform dimensionality reduction. The input of the gated recurrent unit module is the feature vector X3, and the output is the feature vector H. The gated recurrent unit module is used to extract the time features of the feature vector X3 to obtain the feature vector H.

[0030] The input of the gesture recognition module is the feature vector H, and the output is the predicted gesture activity category.

[0031] Each sample is three-dimensional data composed of a velocity spectrum (BVP) in a human body coordinate system (the velocity spectrum in a human body coordinate system is two-dimensional data) in a time frame dimension.

[0032] Example 2

[0033] A gesture recognition model, based on Example 1, the multi-scale feature fusion module 1 is as follows Figure 2 The fusion structure shown in the figure includes: an upper branch structure, a lower branch structure and an addition module. The upper branch structure includes: a sequentially connected dilated convolution module (Yali, P., Lu, Z., Shigang, L., Xiaojun, W., Yu, Z., & Xili, W. (2019) Dilated Residual Networks with Symmetric Skip Connection for Image Denoising., Neurocomputing, 345: 67-76.), the first batch normalization layer (Batch Nomalization Layer), a PReLU activation function and the first CBAM (Sanghyun, W., Jongchan, P., Joon-Young, L., & In So, K. (2018) CBAM: Convolutional Block Attention Module., European Conference on Computer Vision, abs / 1807.06521:3-19.), the lower branch structure includes: a first depth-wise separable convolutional module (M.Humayun, K., Md.Ali, H., & Wonjae, S. (2022) CSI-DeepNet: A Lightweight Deep Convolutional Neural Network Based Hand Gesture Recognition System Using Wi-Fi CSI Signal, IEEE Access, 10: 114787-114801.), a second batch normalization layer (BatchNomalization Layer), a ReLU activation function and a second CBAM, and the summation module is used to sum the output of the first CBAM and the output of the second CBAM element by element to obtain a feature vector.

[0034] The dilated convolution module of the multi-scale feature fusion module 1 and the first depth-separable convolution module simultaneously receive the sample x input1 , sample x input1 At the same time, it is input into the hole convolution module and the first depth-separable convolution module of the multi-scale feature fusion module 1, and the feature vector x is output by the first CBAM. d1 , the feature vector x is output by the second CBAM ds1 ; The summation module adds the feature vector x d1 and the eigenvector x ds1Perform element-by-element addition to obtain the feature vector X1; the convolution kernel size of the hole convolution module of the multi-scale feature fusion module 1 is 6×6, the sliding step size is 2, and the number of convolution kernels is 8; the first depth-separable convolution module of the multi-scale feature fusion module 1 includes: a two-dimensional depth convolution module and a two-dimensional point convolution module, the convolution kernel size of the two-dimensional depth convolution module is 6×6, the sliding step size is 2, and the number of convolution kernels is 8, and the convolution kernel size of the two-dimensional point convolution module is 1×1, the sliding step size is 1, and the number of convolution kernels is 8.

[0035] The multi-scale feature fusion module 2 includes: a filling module and a Figure 2 In the fusion structure shown, the padding module is used to perform a padding operation of padding = 1. The padding module is connected to the hole convolution module in the fusion structure. The padding module and the first depth-separable convolution module of the multi-scale feature fusion module 2 simultaneously receive the feature vector X1. The feature vector X1 is used as the input of the padding module and the first depth-separable convolution module, and the first CBAM of the multi-scale feature fusion module 2 outputs the feature vector x d2 , the feature vector x is output by the second CBAM of the multi-scale feature fusion module 2 ds2 , the sum module adds the feature vector x d2 and the eigenvector x ds2 Perform element-by-element addition to obtain the feature vector X2. The convolution kernel size of the dilated convolution module of the multi-scale feature fusion module 2 is 3×3, the sliding step size is 1, dilation=2, and the number of convolution kernels is 16. The first depth-separable convolution module of the multi-scale feature fusion module 2 includes: two-dimensional depth convolution and two-dimensional point convolution. The convolution kernel size of the two-dimensional depth convolution is 3×3, the sliding step size is 1, and the number of convolution kernels is 16. The convolution kernel size of the two-dimensional point convolution is 1×1, the sliding step size is 1, and the number of convolution kernels is 16.

[0036] The depth-pooling module includes: a second depth-separable convolution module, a padding module, a batch normalization layer (Batch Nomalization), a ReLU activation function, a maximum pooling layer (MaxPool) and a first fully connected layer (FC1) connected in sequence. The padding module is used to perform a padding operation of padding=1. The feature vector X2 is used as the input of the second depth-separable convolution module, and the feature vector X3 is output by the first fully connected layer (FC1). Among them, the second depth-separable convolution module includes: a two-dimensional depth convolution and a two-dimensional point convolution. The number of convolution kernels of the two-dimensional depth convolution is 16, the convolution kernel size is 3×3, and the sliding step size is 1. The number of convolution kernels of the two-dimensional point convolution is 16, the convolution kernel size is 1×1, and the sliding step size is 1. The first fully connected layer (FC1) includes: a first linear transformation layer, a ReLU function and a first Dropout layer, a second linear transformation layer and a ReLU function connected in sequence, and the drop rate of the first Dropout layer is 0.5.

[0037] The gesture recognition module includes: the second Dropout layer, the second fully connected layer (FC2), and the normalized Softmax function (the normalized Softmax function outputs the probability of each gesture activity category, and the gesture activity category with the highest probability is used as the predicted gesture activity category). The second fully connected layer (FC2) maps the high-dimensional feature vector to the category space, avoiding redundant calculations and unnecessary computational overhead.

[0038] Example 3

[0039] Widar3.0 data set (from: Yue, Z., Yi, Z., Kun, Q., Guidong, Z., Yunhao, L., Chenshu, W., & Zheng, Y. (2021) Widar3.0: Zero-Effort Cross-Domain Gesture Recognition with Wi-Fi, IEEE Transactions on Pattern Analysis and Machine Widar (Widar) is a gesture activity dataset based on WiFi-CSI, consisting of 43,000 original samples, divided into 22 gesture activity categories; data of 7 gesture activity categories in the Widar dataset are selected, with a total of 27,600 samples as dataset D, among which the 7 gesture activity categories are "push and pull", "swipe", "clap", "slide", "draw N", "draw Z" and "draw O", and each sample is a three-dimensional data composed of the velocity spectrum (BVP) in the human coordinate system (the velocity spectrum in the human coordinate system is two-dimensional data) in the time frame dimension; the dimension of each sample in dataset D is 22×20×20, 22 is the time frame of the sample, the length and width of the velocity spectrum (BVP) in the human coordinate system are 20, and each sample contains a label, that is, the gesture activity category (true value) of the sample.

[0040] The samples in the dataset D are divided into a training set and a test set in a ratio of 8:2, and the gesture recognition module makes predictions in seven gesture activity categories.

[0041] Example 4

[0042] A method for training a gesture recognition model, comprising:

[0043] Each sample in the training set in Example 3 is taken as sample x input1The and labels are input into the gesture recognition model of Example 2 for training to obtain the predicted gesture activity category of each sample, wherein a batch of samples for training is 16, and a loss function calculation module is provided in the gesture recognition model. The loss function calculation module adopts the cross entropy loss function (Bing, L., Wei, C., Wei, W., Le, Z., Zhenghua, C., & Min, W. (2021) Two-Stream Convolution Augmented Transformer for Human Activity Recognition, AAAI Conference on Artificial Intelligence, 35.1: 286-293.) to calculate the loss value, and the gesture recognition model uses the gradient descent back propagation method (Lin, W., Yi, Z., & Tao, C. (2015) Back Propagation Neural Network with Adaptive Differential Evolution Algorithm for Time Series Forecasting, Expert Systems with Applications, 42.2:855-863.) The parameters are optimized until the maximum number of iterations epcohs=200 is reached, and the training is terminated to obtain a trained gesture recognition model.

[0044] Example 5

[0045] A gesture recognition method based on WIFI-CSI multi-scale feature fusion includes: inputting samples of a test set into the gesture recognition model trained in Example 4 to identify gesture activity categories, and obtaining a predicted gesture activity category for each sample.

[0046] Example 6

[0047] A gesture recognition method based on WIFI-CSI multi-scale feature fusion is basically the same as Example 5, with the only difference being that: in the gesture recognition model adopted in this embodiment, the multi-scale feature fusion module 1 and the multi-scale feature fusion module 2 do not have an upper branch structure and a summation module, and the multi-scale feature fusion module 2 does not have a filling module.

[0048] Example 7

[0049] A gesture recognition method based on WIFI-CSI multi-scale feature fusion is basically the same as Example 5, with the only difference being that: in the gesture recognition model adopted in this embodiment, the multi-scale feature fusion module 1 and the multi-scale feature fusion module 2 do not have a lower branch structure and a summation module.

[0050] Example 8

[0051] A gesture recognition method based on WIFI-CSI multi-scale feature fusion is basically the same as Example 5, with the only difference being that in the gesture recognition model adopted in this embodiment, the multi-scale feature fusion module 1 and the multi-scale feature fusion module 2 do not have the first CBAM and the second CBAM.

[0052] The recognition accuracy of the gesture recognition methods of Examples 5 to 8 is compared, as shown in Table 1. As can be seen from Table 1, the gesture recognition method of the present invention has the highest recognition accuracy for the gesture activity of the test set, which is 84.30±0.31%. Therefore, the gesture recognition method of the present invention has the best gesture activity recognition effect. Figure 3 As shown, the gesture recognition method of the present invention has a recognition accuracy rate of more than 80% for each gesture activity category in the test set.

[0053] Table 1

[0054] Example Accuracy(%) Example 6 83.44±1.13 Example 7 71.19±0.60 Example 8 84.22±0.68 Example 5 84.30±0.31

[0055] Example 9

[0056] The gesture activity categories of the test set are identified using MLP, CNN-5, CNN+GRU, LSTM, GRU, ViT and ResNet18 (Jianfei, Y., Xinyan, C., Han, Z., Chris Xiaoxuan, L., Dazhuo, W., Sumei, S., & Lihua, X. (2023) SenseFi: A Library and Benchmark on Deep-Learning-Empowered WiFi Human Sensing, PATTERNS, 4.3: 100703-100703.) to obtain the gesture activity category of each sample.

[0057] Using the accuracy (Acc) for evaluating prediction ability, the floating-point operations (Flops) for evaluating computational complexity, and the parameter quantity (Params) for measuring GPU memory requirements, Example 5 is compared with MLP, CNN-5, CNN+GRU, LSTM, GRU, ViT, and ResNet18. The comprehensive evaluation results are shown in Table 2.

[0058] Table 2

[0059]

[0060]

[0061] It can be seen from Table 2 that the gesture recognition method of the present invention has an accuracy rate of 84.30±0.31% when the number of parameters is small and the calculation complexity is low.

[0062] The present invention is described above by way of example. It should be noted that, without departing from the core of the present invention, any simple deformation, modification or other equivalent replacement that can be made by those skilled in the art without inventive effort falls within the protection scope of the present invention.

Claims

1. A gesture recognition model, characterized in that: include: Feature extraction module and gesture recognition module, wherein the feature extraction module includes: multi-scale feature fusion module 1, multi-scale feature fusion module 2, depth-pooling module and gated recurrent unit module, and the input of multi-scale feature fusion module 1 is sample x input1 , the output is the feature vector X1, and the multi-scale feature fusion module 1 is used to extract each sample X input1 The input of the multi-scale feature fusion module 2 is the feature vector X1, and the output is the feature vector X2. The multi-scale feature fusion module 2 is used to extract the features in the feature vector X1 and perform multi-scale feature fusion on it. The input of the depth-pooling module is the feature vector X2, and the output is the feature vector X3. The depth-pooling module is used to extract the features of the feature vector X2 and perform dimensionality reduction. The input of the gated recurrent unit module is the feature vector X3, and the output is the feature vector H. The gated recurrent unit module is used to extract the time features of the feature vector X3 to obtain the feature vector H. The input of the gesture recognition module is the feature vector H, and the output is the predicted gesture activity category; Each sample is three-dimensional data composed of the velocity spectrum in the time frame dimension in the human body coordinate system.

2. The gesture recognition model according to claim 1, characterized in that: The multi-scale feature fusion module 1 is a fusion structure, which includes: an upper branch structure, a lower branch structure and a summation module. The upper branch structure includes: a hole convolution module, a first batch normalization layer, a PReLU activation function and a first CBAM connected in sequence. The lower branch structure includes: a first depthwise separable convolution module, a second batch normalization layer, a ReLU activation function and a second CBAM connected in sequence. The summation module is used to add the output of the first CBAM and the output of the second CBAM element by element to obtain a feature vector.

3. The gesture recognition model according to claim 2, characterized in that: The multi-scale feature fusion module 2 includes: a filling module and the fusion structure, the filling module is used to perform a filling operation of padding=1, the filling module is connected to the hollow convolution module in the fusion structure, and the filling module and the first depth-separable convolution module of the multi-scale feature fusion module 2 simultaneously receive the feature vector X1.

4. The gesture recognition model according to claim 1, characterized in that: The depth-pooling module includes: a second depth-separable convolution module, a padding module, a batch normalization layer, a ReLU activation function, a maximum pooling layer and a first fully connected layer connected in sequence. The padding module is used to perform a padding operation of padding=1. The feature vector X2 is used as the input of the second depth-separable convolution module, and the feature vector X3 is output by the first fully connected layer.

5. The gesture recognition model according to claim 4, characterized in that: The second depth-wise separable convolution module includes: two-dimensional depth-wise convolution and two-dimensional point-wise convolution.

6. The gesture recognition model according to claim 4, characterized in that: The first fully connected layer includes: a first linear transformation layer, a ReLU function and a first Dropout layer, a second linear transformation layer and a ReLU function connected in sequence.

7. The gesture recognition model according to claim 1, characterized in that: The gesture recognition module includes: a second Dropout layer, a second fully connected layer and a normalized Softmax function connected in sequence.

8. A method for training a gesture recognition model, characterized in that: include: Take each sample in the training set as sample x input1 The gesture recognition model is input with a label into any one of claims 1 to 7 for training to obtain a predicted gesture activity category for each sample, wherein a loss function calculation module is provided in the gesture recognition model, and the loss function calculation module uses a cross entropy loss function to calculate the loss value. The gesture recognition model optimizes its parameters using a gradient descent back propagation method until the training is terminated when the maximum number of iterations is reached, thereby obtaining a trained gesture recognition model.

9. A gesture recognition method based on WIFI-CSI multi-scale feature fusion, comprising: The samples of the test set are input into the gesture recognition model trained in claim 8 to identify the gesture activity category, and the predicted gesture activity category of each sample is obtained.

Citation Information

Patent Citations

  • SAR image recognition method and device based on multi-scale feature and width learning

    CN110222700A

  • Dynamic gesture recognition method based on multi-modal data

    CN113255602A

  • Intelligent control method and device based on lightweight gesture recognition

    CN115328319A

  • Electronic component depth migration identification method based on multi-scale attention mechanism

    CN115375946A

  • Real-time WiFi signal gesture recognition method allowing user authentication

    CN115392321A