A gesture recognition method for WiFi data based on machine learning
By building a multi-task perception model and using knowledge distillation to train WiFi CSI channel data, the problem of low gesture recognition accuracy of WiFi data is solved, and higher recognition accuracy and practicality of IoT applications are achieved, and accurate response of smart home systems is supported.
Patent Information
- Application Number
- CN202211596202.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-13
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-12-13
AI Technical Summary
The accuracy of gesture recognition based on WiFi data in the prior art is low, and the existing multitasking sensing method has not been effectively applied to commercial WiFi devices, which limits its practicality and accuracy in Internet of Things applications.
Build a multi-task perception model, including a single-task model, residual adapter and depth encoder, train the multi-task perception model through knowledge distillation method, use WiFi's CSI channel data for gesture recognition, extract low-level and advanced features of the amplification matrix, and predict and classify through classifiers.
It improves the accuracy of gesture recognition based on WiFi data and the practicality of IoT applications, and uses multi-task perception models to mine more user information than single tasks, supporting the accurate response of smart home systems.
Smart Images

Figure CN116070156B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet of Things action recognition, and more specifically, to a gesture recognition method for WiFi data based on machine learning. Background Art
[0002] Gesture recognition (human gesture recognition, HGR) plays an important role in human-computer interaction and can support many emerging Internet of Things applications, such as smart home, user authentication, and healthcare. Traditional gesture recognition systems rely on cameras, wearable devices, and smartphones; these methods incur additional device costs, reduce user comfort, or raise privacy concerns; specifically, gesture recognition based on computer vision is already a relatively mature solution; however, this type of solution requires the use of a camera for image acquisition, posing a risk of user privacy leakage; at the same time, this solution can only perform gesture recognition under line-of-sight conditions, limiting the application scenarios. In addition, solutions based on wearable devices and built-in sensors of mobile phones need to be carried by users, greatly affecting user comfort.
[0003] In recent years, with the continuous development of ubiquitous network technology, radio frequency sensing has seen rapid development and received a great deal of attention; among them, sensing research based on Wi-Fi, WB, MCW radar, and meter waves has emerged in an endless stream; since radio frequency sensing data does not have direct semantic information, this technology has extremely high privacy; in addition, this solution is a non-wearable and non-invasive sensing method that has no impact on user comfort; compared with other radio frequency sensing solutions, sensing based on commercial Wi-Fi devices also has the advantages of easy deployment and low cost, and thus has received much attention.
[0004] Although significant progress has been made in the field of WiFi-based sensing, various pioneering methods are limited to single-task sensing, such as human activity recognition, indoor positioning, gait recognition, and respiration detection; there are some studies on multi-user gesture recognition based on WiFi, however, ARIL and WiHF are the only two works on multi-task sensing based on WiFi; specifically, ARIL aims to perform the joint tasks of activity recognition and indoor positioning, and is built on a general software radio peripheral device rather than a commercial WiFi device, which reduces its practicality; WiHF focuses on simultaneously achieving cross-domain gesture recognition and user recognition and can only handle the joint gesture recognition and user recognition tasks.
[0005] The development of multi-task sensing based on WiFi will provide more user information than single-task sensing, which will promote many Internet of Things applications. For example, in smart home applications, the system can provide information on "who is doing what where", which will help the smart home system respond accurately to users;
[0006] The prior art discloses a multimodal multitask model for gesture detection and gesture recognition, including a modal feature extraction module, a multimodal fusion module, and a model multitask classification module. Among them, the modal feature extraction module includes a network structure for respectively extracting features of different modal data and a shared feature layer. The modal feature extraction module is used to preprocess multimodal data and extract shared multimodal features. The multimodal fusion module includes a multimodal channel attention module and a task-related feature layer. The multimodal fusion module is connected to the modal feature extraction module, takes the shared multimodal features as the input of the multimodal channel attention module, extracts the fused task-related features, and obtains a task-related feature layer. The model multitask classification module is connected to the multimodal fusion module, takes the fused task-related features as the input, and classifies each task. Among them, during the training process, the network parameters of the modal feature extraction module, the multimodal fusion module, and the model multitask classification module are iteratively updated. However, the prior art does not perform gesture recognition for WiFi data, and the accuracy of gesture recognition is low. Summary of the Invention
[0007] In order to solve the problem in the prior art that multitask perception is not performed for WiFi data and the accuracy of gesture recognition is low, the present invention proposes a gesture recognition method for WiFi data based on machine learning. By processing the CSI channel data of WiFi, gesture recognition for WiFi data is realized, and the accuracy of gesture recognition is improved.
[0008] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0009] A gesture recognition method for WiFi data based on machine learning, including the following steps:
[0010] S1: Construct a multitask perception model. The multitask perception model includes several single-task models, a residual adapter, and a first depth encoder. The single-task model includes a shallow encoder, a second depth encoder, and a classifier. The shallow encoder is used to extract low-level features of the amplitude matrix of the task. The second depth encoder is used to extract high-level features of the task according to the low-level features. The classifier is used to predict and classify the task. The single-task model is used to predict the classification result of a single task. The residual adapter is used to extract the compensation features specific to each single task. S2: Take the CSI channel data of WiFi as an input sample and input it into the multitask perception model, and use the knowledge distillation method to train the multitask perception model to obtain a trained multitask perception model. S3: Use the trained multitask perception model to predict and classify WiFi data, and realize gesture recognition for the classified gesture recognition task.
[0011] The working principle of the present invention is as follows:
[0012] Using WiFi CSI channel data as input samples, a multi-task perception model is constructed and trained using the knowledge distillation method. The trained multi-task perception model is used to predict and classify the multiple tasks of WiFi CSI channel data, and gesture recognition is completed for the classified gesture recognition tasks.
[0013] Preferably, the training method of the multi-task perception model is as follows:
[0014] The amplification matrix of the task corresponding to the WiFi CSI channel data is used as a data set, which includes a training set and a test set. The ratio of the training set to the test set is 8:2. The training set is input into the corresponding single-task model for training, and the parameters of the single-task model are frozen after training. A residual adapter is used to extract compensation features specific to each single task. A first deep encoder is used to extract common features of all single tasks based on the low-level features of all single tasks. The combination of the corresponding compensation features and the common features is input into each classifier for prediction and classification.
[0015] Furthermore, the knowledge distillation method is as follows:
[0016] According to the tasks corresponding to the WiFi CSI channel data, the corresponding single-task models are trained. There are T tasks in the sample, one of which is a single task t. For single task t, a single-task model M is trained. t st (·;θ st ,δ st ,ξ st ), freeze the parameters corresponding to the single-task model; construct the first loss function Make the common feature F n wi,comm After linear transformation under Euclidean distance, it is close to the high-level features F produced by the single-task model. t,n st,high ; Using the second loss function The logistic regression Z obtained by the multi-task perception model t,n wi Logistic regression Z with single-task model t,n st Similar; for the input sample H n The final output is a set of probability distributions p n ={p 1,n ,p 2,n ,...,p T,n}, corresponding to all T tasks; for the results of the prediction task, define the third loss function l wi (p n ,y n); The first loss function, the second loss function, and the third loss function are combined to obtain a fourth loss function, and the parameters of the multi-task perception model are learned by minimizing the fourth loss function.
[0017] Furthermore, for the input sample H n For the t-th single task, the expression of the first loss function is as follows:
[0018]
[0019] Where, LT t (·) is the linear transformation of the t-th task, implemented as a 1*1*C*C convolution, where C is F n wi,comm The number of channels; It is a Euclidean linear transformation of high-level features.
[0020] Furthermore, the expression of the second loss function is as follows:
[0021]
[0022] Where, It is the classification result of the multi-task perception model for the current single task, and the obtained value is the possible probability of each single task; The classification result of the single-task model for the current single task; σ(·) is the softmax function, and τ is a hyperparameter for the user to adjust the distillation strength.
[0023] Furthermore, each element in the probability distribution is calculated as follows:
[0024]
[0025] Where, Yes H n The predicted label on the t-th single task, and m t ∈{1,2,...,M t}.
[0026] The third loss function expression is as follows:
[0027]
[0028] Where, ω t is the hyperparameter of the t-th single task, l t wi (·,·) is a loss function based on cross entropy.
[0029] Furthermore, the expression of the fourth loss function is as follows:
[0030]
[0031] In the formula, D is the training set, and λ is the weight.
[0032] Preferably, the amplification matrix corresponding to the CSI channel data of the said WiFi for the task is expressed as
[0033] ||H||∈R L*S*P ; in the formula, L is the number of links, L = RX * TX, RX is the number of antennas of the receiver, and TX is the number of antennas of the transmitter; S is the number of subcarriers; P is the length of sampling, P = r * t e , where r represents the sampling rate of the CSI capture tool, and t e represents the gesture execution time.
[0034] Preferably, the training method for the single-task model is as follows:
[0035] Preprocess and denoise the input data to obtain the processed input data; use a shallow encoder to extract low-level features on each link; use the low-level features as input for the deep encoder to extract high-level features; use a loss function to obtain the prediction result through a classifier.
[0036] Preferably, for T different tasks, the multi-task perception model is provided with T classifiers, and the classifier for the t-th single task is expressed as C t wi (·; ξ t wi ); the said classifier uses the softmax function.
[0037] Compared with the prior art, the beneficial effects of the present invention are:
[0038] 1. For WiFi data, it is applied to the Internet of Things to improve practicality.
[0039] 2. A multi-task perception model is constructed to classify multiple tasks, and more perception tasks than single tasks are mined through multi-task perception to improve the accuracy of gesture recognition.
[0040] 3. The multi-task imbalance problem is solved by the knowledge distillation method, and specific task compensation features are extracted through the residual adapter to improve the performance of the multi-task perception model. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a flowchart of the said gesture recognition method for WiFi data based on machine learning.
[0042] Figure 2 is a training flowchart of the said multi-task perception model.
[0043] Figure 3 It is the structural parameter diagram of the multi-task perception model in the embodiment. Detailed implementation manners
[0044] The present invention will be described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0045] Embodiment 1
[0046] In this embodiment, as Figure 1 shown, a gesture recognition method for WiFi data based on machine learning includes the following steps:
[0047] S1: Construct a multi-task perception model, the multi-task perception model includes a plurality of single-task models, a residual adapter, and a first depth encoder; the single-task model includes a shallow encoder, a second depth encoder, and a classifier; the shallow encoder is used to extract low-level features of the amplitude matrix of the task; the second depth encoder is used to extract high-level features of the task according to the low-level features; the classifier is used to perform prediction classification on the task; the single-task model is used to predict the classification result of a single task; the residual adapter is used to extract compensation features specific to each single task; S2: Input the CSI channel data of WiFi as an input sample into the multi-task perception model, and use the knowledge distillation method to train the multi-task perception model to obtain a trained multi-task perception model. S3: Use the trained multi-task perception model to perform prediction classification on the WiFi data, and for the gesture recognition task after classification, realize gesture recognition.
[0048] In this embodiment, the training method of the multi-task perception model is as follows:
[0049] Use the amplitude matrix corresponding to the task of the CSI channel data of WiFi as the data set, the data set includes a training set and a test set; the ratio of the training set to the test set is 8:2; input the training set into the corresponding single-task model for training respectively, and freeze the parameters of the single-task model after training; use the residual adapter to extract compensation features specific to each single task; use the first depth encoder to extract the common features of all single tasks according to the low-level features of all single tasks; input the combination of the corresponding compensation features and common features into each classifier for prediction classification.
[0050] In this embodiment, the knowledge distillation method is as follows:
[0051] Train the corresponding single-task model according to the task corresponding to the CSI channel data of WiFi. For single task t, train a single-task model M t st (·; θ st , δ st,ξ st ), freeze the parameters corresponding to the single-task model; construct the first loss function Make the common feature F n wi,comm After linear transformation under Euclidean distance, it is close to the high-level features F produced by the single-task model. t,n st,high ; Using the second loss function The logistic regression Z obtained by the multi-task perception model t,n wi Logistic regression Z with single-task model t,n st Similar; for the input sample H n The final output is a set of probability distributions p n ={p 1,n ,p 2,n ,...,p T,n}, corresponding to all T tasks; for the results of the prediction task, define the third loss function l wi (p n ,y n ); The first loss function, the second loss function, and the third loss function are combined to obtain a fourth loss function, and the parameters of the multi-task perception model are learned by minimizing the fourth loss function.
[0052] More specifically, for the input sample H n For the t-th single task, the expression of the first loss function is as follows:
[0053]
[0054] Where, LT t (·) is the linear transformation of the t-th single task, implemented as a 1*1*C*C convolution, where C is F n wi,comm The number of channels; It is a Euclidean linear transformation of high-level features.
[0055] More specifically, the expression of the second loss function is as follows:
[0056]
[0057] Where, It is the classification result of the multi-task perception model for the current single task, and the obtained value is the possible probability of each single task; The classification result of the single-task model for the current single task; σ(·) is the softmax function, and τ is a hyperparameter for the user to adjust the distillation strength.
[0058] More specifically, each element in the probability distribution is calculated as follows:
[0059]
[0060] Where, Yes H n The predicted label on the t-th single task, and m t ∈{1,2,...,M t}.
[0061] More specifically, the third loss function is expressed as follows:
[0062]
[0063] Where, ω t is the hyperparameter of the t-th single task, l t wi (·,·) is a loss function based on cross entropy.
[0064] More specifically, the expression of the fourth loss function is as follows:
[0065]
[0066] Where D is the training set, λ is The weight of .
[0067] Preferably, the amplification matrix of the WiFi CSI channel data corresponding task is expressed as ||H||∈R L*S*P Where L is the number of links, L = RX * TX, RX is the number of antennas of the receiver, TX is the number of antennas of the transmitter; S is the number of subcarriers; P is the sampling length, P = r * t e , where r represents the sampling rate captured by the CSI tool, t e Indicates the gesture execution time.
[0068] Preferably, the training method for the single-task model is as follows:
[0069] The input data is preprocessed and denoised to obtain processed input data; a shallow encoder is used to extract low-level features on each link; a deep encoder uses low-level features as input to extract high-level features; and a loss function is used to obtain prediction results through a classifier.
[0070] Preferably, for different T tasks, the multi-task perception model has T classifiers, and the classifier of the t-th single task is represented by C t wi (·;ξ t wi); The classifier uses the softmax function; the softmax function is built into the classifier, and what the softmax function obtains are the probabilities of each data category. The classifier will judge and compare the obtained probabilities, and regard the category with the highest probability as the classification result of this data.
[0071] Embodiment 2
[0072] In this embodiment, a gesture recognition method for WiFi data based on machine learning is as Figure 2 shown, and includes the following steps:
[0073] Using the amplitude matrix of CSI samples as the data set to implement the perception task; using a CSI capture tool based on Intel 5300 NIC to obtain relevant data by itself; or downloading relevant data sets from the Internet, such as Widar 3.0; constructing a multi-task perception model and training it; the training process is as follows:
[0074] Training the single-task model: preprocessing and denoising the input data. For the single perception task, use the corresponding shallow encoder to extract low-level features on each link The second deep encoder φ n wi (·; δ wi ) takes the low-level features as input and extracts the high-level features F of the single task t,n st,high ; The classifier C wi (·; ξ wi ) outputs the corresponding predicted classification result, and the loss function used is The high-level features are extracted from the low-level features in the single task and are only the features of the current single task.
[0075] The single task can be gesture recognition, indoor positioning, user recognition, etc. Different tasks use different data; taking gesture recognition as an example, assuming there are 5 different gestures, the classifier will obtain the probability data that the signal data belongs to 5 different categories, and the data category with the largest probability is the classification prediction result of the signal data.
[0076] Freeze the parameters of each single-task model. For the problem of unbalanced data precision, each single task introduces a residual adapter f(·; σ wi ) to extract task-specific compensation features. For the CSI sample H n The compensation feature on the t-th single task is
[0077] Using the first deep encoder φ wi (·; δ wi), with all single-task low-level features F n wi,low As input, extract the common features F of all single tasks n wi,comm , F n wi,comm =φ wi (F n wi,low ; δ wi ).
[0078] The multi-task perception model has T classifiers for different T tasks; the classifier for the tth task is C t wi (·;ξ t wi ), the input of each classifier is a combination of the corresponding task-specific compensation features and common features. For the classification results of the prediction task, the loss function is defined as: The multi-task perception model parameters are learned by minimizing the final loss function:
[0079] After obtaining the trained multi-task perception model, for newly obtained CSI samples, the corresponding single-task models are trained according to the tasks corresponding to the samples. After completing the training of the single-task models, the corresponding parameters are frozen. Different residual adapters are introduced according to the accuracy of the input data, and the combination of compensating features and common features is input into each classifier to obtain the final classification result.
[0080] Example 3
[0081] Based on the machine learning-based gesture recognition method for WiFi data described in Example 1 and Example 2; Figure 3 As shown, S represents the number of CSI subcarriers; P represents the time length of the CSI sample; L is the number of links; the convolution layer of the shallow encoder includes a one-dimensional convolution layer with a convolution kernel shape of 7*1 and a stride of 2, a one-dimensional normalization layer, and a one-dimensional maximum pooling layer with a convolution kernel shape of 3*1 and a stride of 2; the deep encoder includes two one-dimensional convolution layers with a convolution kernel shape of 3*1 and a stride of 2, four one-dimensional convolution layers with a convolution kernel shape of 3*1 and a stride of 1, and six one-dimensional normalization layers; the classifier includes a one-dimensional convolution layer with a convolution kernel shape of 3*1 and a stride of 1, a one-dimensional normalization layer, an average pooling layer, and a linear classification layer; the residual adapter includes a one-dimensional convolution layer with a convolution kernel shape of 3*1 and a stride of 2, and a one-dimensional normalization layer.
[0082] The pseudo code for training the multi-task perception model is as follows:
[0083] For the single-task shallow encoder after the t-th single-task training;
[0084] φ t st (·): single-task deep encoder after training for the t-th single task;
[0085] C t st (·): single-task classifier after training for the t-th single task;
[0086] Multi-task shallow encoder;
[0087] φ wi (·): multi-task deep encoder;
[0088] C t wi (·): classifier of the t-th single-task multi-task perception model;
[0089] f t (·): residual adapter for the t-th single task in the multi-task perception model;
[0090] Input: training set where y t,i ∈{1,2,...,M t} is x i The label on the t-th single task, and M t is the number of classifications for the t-th single task;
[0091] Output: Loss J for a randomly sampled training set.
[0092] start:
[0093]
[0094]
[0095] Finish.
[0096] Obviously, the above embodiments of the present invention are only examples for clearly explaining the present invention, and are not intended to limit the implementation methods of the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A gesture recognition method for WiFi data based on machine learning, characterized in that It includes the following steps: S1: Construct a multi-task perception model, where the multi-task perception model includes several single-task models, a residual adapter, and a first depth encoder; The single-task model includes a shallow encoder, a second depth encoder, and a classifier; The shallow encoder is used to extract low-level features of the amplitude matrix of the task; The second depth encoder is used to extract high-level features of the task based on the low-level features; The classifier is used to perform prediction classification on the task; The single-task model is used to predict the classification result of a single task; the residual adapter is used to extract compensation features specific to each single task; S2: Use the CSI channel data of WiFi as input samples and input them into the multi-task perception model. Use the knowledge distillation method to train the multi-task perception model to obtain the trained multi-task perception model; The knowledge distillation method is as follows: Train corresponding single-task models according to the tasks corresponding to the CSI channel data of WiFi. For single task t, train a single-task model , and freeze the parameters corresponding to the single-task model; Construct the first loss function , such that the common features are close to the high-level features generated by the single-task model after linear transformation under the Euclidean distance ; Using the second loss function to make the logistic regression obtained by the multi-task perception model similar to the logistic regression of the single-task model similar; The final output for the input sample is a set of probability distributions , corresponding to all T tasks; for the results of the prediction tasks, a third loss function is defined ; Integrate the first loss function, the second loss function, and the third loss function to obtain the fourth loss function. The parameters of the multi-task perception model are learned by minimizing the fourth loss function; S3: Use the trained multi-task perception model to perform prediction classification on WiFi data to achieve gesture recognition; The augmentation matrix corresponding to the CSI channel data of the described WiFi for the task is expressed as ; Where L is the number of links, , is the number of antennas of the receiver, is the number of antennas of the transmitter; S is the number of subcarriers; P is the length of the sampling, , where represents the sampling rate of the CSI capture tool, represents the gesture execution time.
2. The gesture recognition method for WiFi data based on machine learning according to claim 1, wherein The training method for the multi-task perception model is as follows: Use the amplitude matrix corresponding to the CSI channel data of WiFi for the task as the data set, and the data set includes a training set and a test set; the ratio of the training set to the test set is 8:2; Input the training set into the corresponding single-task model for training respectively, and freeze the parameters of the single-task model after training; Use the residual adapter to extract compensation features specific to each single task; Use the first depth encoder to extract common features of all single tasks based on the low-level features of all single tasks; Input the combination of the corresponding compensation features and common features into each classifier for prediction classification.
3. The gesture recognition method for WiFi data based on machine learning according to claim 2, wherein, For the t-th single task in the input sample The expression of the first loss function is as follows: In the formula, is the linear transformation of the t-th single task, implemented as a 1*1*C*C convolution, where C is the number of channels; is the Euclidean linear transformation of the high-level features.
4. A gesture recognition method for WiFi data based on machine learning according to claim 3, characterized in that, The expression of the second loss function is as follows: In the formula, , is the classification result of the multi-task perception model for the current single task, and the obtained result is the probability of each possible single task; , which is the classification result of the single-task model for the current single task; is the softmax function, is a hyperparameter for the user to adjust the distillation intensity.
5. A gesture recognition method for WiFi data based on machine learning according to claim 4, characterized in that, The calculation method for each element in the probability distribution is as follows: In the formula, is the predicted label on the t-th single task, and ; The expression of the third loss function is as follows: In the formula, is the hyperparameter of the t-th single task, is the loss function based on cross entropy.
6. A gesture recognition method for WiFi data based on machine learning according to claim 5, characterized in that, The expression of the fourth loss function is as follows: where D is the training set, is weight.
7. A gesture recognition method for WiFi data based on machine learning according to claim 1, wherein The training method for the single-task model is as follows: Preprocess and denoise the input data to obtain the processed input data; use the shallow encoder to extract low-level features on each link; the depth encoder uses the low-level features as input to extract high-level features; use the loss function to obtain the prediction result through the classifier.
8. A gesture recognition method for WiFi data based on machine learning according to claim 1, characterized in that, For different T tasks, the multi-task perception model has T classifiers, and the classifier for the t-th single task is expressed as ; The classifier uses the softmax function.
Citation Information
Patent Citations
Activity classification model construction method and system based on Wi-Fi signals and transfer learning
CN111460901A
Gesture recognition method and system based on knowledge distillation
CN114970640A