A cross-domain small sample gesture recognition method based on wi-fi signals
Patent Information
- Application Number
- CN202311152244.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-07
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2043-09-07
AI Technical Summary
[0005](1)可穿戴设备和安装摄像头的方法:这些传统的手势识别解决方案需要用户佩戴设备或者在环境中安装摄像头,这可能导致不舒适的用户体验,并且可能侵犯用户的隐私
[0066](1)本发明提出了一个全新的跨域手势识别方法,它仅仅依赖于一对Wi-Fi收发器就可以提取一应俱全的手势特征,显著降低了设备的部署成本。同时,本发明使用了元学习任务生成方案,能够生成大量的任务来模拟不同的领域变化,对于新的目标域,每个手势只需使用一个样本就可以快速适应新领域,显著降低了数据收集开销。这些优势极大地节省了人力和物力的消耗。
Smart Images

Figure CN117290714B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of gesture recognition technology, and in particular relates to a cross-domain small sample gesture recognition method based on Wi-Fi signals. Background Technology
[0002] Currently, gesture recognition is gradually permeating all aspects of human life, such as healthcare, transportation, and smart homes. Through gesture recognition, users can interact directly and in a friendly manner with IoT devices. For example, in a smart home, a simple wave of the hand can play television programs without contact. In supermarkets, customers can purchase goods without touching self-service kiosks. Most traditional gesture recognition solutions require wearable devices or cameras. While these methods are promising, they can lead to uncomfortable user experiences or privacy violations. In recent years, wireless sensing technology has emerged as a promising alternative for achieving device-free gesture recognition due to its device-free and non-invasive nature. The principle behind wireless sensing is that the wireless signal reflected by a target changes as the target moves. Therefore, by analyzing the changes in the signal, motion-related information can be inferred. Extensive research in recent years on various wireless signals such as Wi-Fi, RFID, acoustics, and mmWave has been dedicated to inferring gestures. Among these signals, Wi-Fi-based solutions are particularly attractive, as Wi-Fi infrastructure is not only inexpensive but also widely deployed around us.
[0003] With the significant success of deep learning in numerous fields, building Wi-Fi-based gesture recognition systems using deep learning has become a growing trend. While these methods are promising, they all rely on the assumption that training and test examples are identical and independently distributed. That is, they only perform well in specific domains. Once the domain changes, this assumption breaks, and the system's performance drops sharply. This is because the received signals contain a large amount of adverse environmental information unrelated to gestures, causing the test distribution to deviate from the training distribution. A simple approach to address this domain dependency challenge is to collect large amounts of data across all possible domains. Unfortunately, this is impractical in real-world scenarios due to the heavy workload of data collection. Recent attempts to address cross-domain issues have emerged; however, these solutions still have several drawbacks hindering their real-world application. First, many systems utilize adversarial learning or transfer learning to improve the model's cross-domain generalization ability. However, these methods require large amounts of target domain data to train the model, which is both time-consuming and labor-intensive. Second, some existing systems rely on multiple transceivers to extract all available features and assume that the device's location is known, which is not always true in real-world scenarios. Furthermore, all of the above methods assume that the transceiver deployment is fixed, which is usually not true. This means that gesture performance will degrade when the device deployment changes.
[0004] Based on the above analysis, the problems and shortcomings of the existing technology are as follows:
[0005] (1) Wearable devices and camera installation methods: These traditional gesture recognition solutions require users to wear devices or install cameras in the environment, which may lead to an uncomfortable user experience and may infringe on user privacy.
[0006] (2) Domain dependence of deep learning-based Wi-Fi gesture recognition systems: These systems perform well in specific domains, but their performance drops sharply once the domain changes. This is because the received signals contain a large amount of unfavorable environmental information unrelated to gestures, causing the test distribution to deviate from the training distribution.
[0007] (3) Data collection difficulties: To address the challenge of domain dependence, collecting large amounts of data covering various possible domains is one solution. However, in real-world scenarios, collecting large-scale data is a burdensome and impractical task.
[0008] (4) Disadvantages of solutions to cross-domain problems: Existing methods for solving cross-domain problems have several drawbacks. First, some methods use adversarial learning or transfer learning to improve the model's cross-domain generalization ability, but this requires a large amount of target domain data for training, which is very time-consuming and labor-intensive. Second, some systems rely on multiple transceivers to extract complete features and assume that the device location is known, which is not always true in real-world scenarios. In addition, all methods assume a fixed transceiver deployment, while in reality, device deployment may change, which can lead to a degrade in gesture performance. Summary of the Invention
[0009] To address the problems existing in the prior art, this invention provides a cross-domain few-sample gesture recognition method based on Wi-Fi signals.
[0010] This invention is implemented as follows: a cross-domain few-sample gesture recognition method based on Wi-Fi signals, the method comprising the following steps:
[0011] S1: Data preprocessing: Phase shift is eliminated using conjugate multiplication, static components and noise in the signal are eliminated using two different filters, and then Doppler spectrum is extracted using signal processing methods such as principal component analysis and short-time Fourier transform.
[0012] S2: Doppler variation pattern extraction: A multi-domain adversarial network was used to extract the Doppler variation pattern through the combined action of a feature extractor, a gesture classifier, and a domain discriminator.
[0013] S3: Prototype network initialization: In the meta-training phase, the meta-learning model is initialized using the training dataset from the source domain.
[0014] S4: Meta-learning multi-task generation: Employs a task generation scheme based on the training dataset to generate a large number of tasks to simulate different domain changes and train the basic network to adapt to the new domain;
[0015] S5: Parameter Update: Update the parameters of the prototype network using multiple tasks generated by S4, and use cross-entropy loss to evaluate the loss of each task;
[0016] S6: New Domain Adaptation: In order to achieve accurate gesture recognition in the new domain, some samples are selected from each class in the new domain for domain adaptation to further calibrate the parameters of the training dataset and complete the training of the model.
[0017] Furthermore, the S1 data preprocessing section includes:
[0018] (1) Receive the CSI data collected by the receiving device and model it as follows:
[0019]
[0020] Where H s (f) represents the sum of all static path responses, N d This represents the number of dynamic paths. (η+β) is the phase offset caused by carrier frequency offset, sampling frequency offset, and packet detection delay.
[0021] (2) Phase offset is eliminated by calculating the conjugate multiplication between antennas. For antenna selection, the coefficient ρ, which is the ratio of the amplitude to the variance of the CSI signal, is used. m Indicators for antenna selection:
[0022]
[0023] Where var and mean represent k th subcarrier m th The variance and mean of the antenna amplitude readings are calculated. Based on this calculation, the present invention selects the antenna pair with the highest and lowest ratio coefficients as the final antenna selection. The insight of the antenna selection scheme is to choose an antenna with the largest CSI variance and an antenna with the largest CSI amplitude. The reason is that a larger CSI variance generally produces a larger dynamic response, while a higher CSI amplitude generally produces a larger static response.
[0024] (3) Use high-pass and low-pass filters respectively to eliminate static components and high-frequency noise.
[0025] (4) In order to further denoise and compress the CSI data dimension, the PCA method is applied to retain the significant components caused by the target motion and the main components extracted by PCA are subjected to short-time Fourier transform to obtain the Doppler spectrum.
[0026] Furthermore, the Doppler variation mode extraction portion of S2 includes:
[0027] (1) Feature Extraction. The feature extractor focuses on extracting features of Doppler variation patterns. In the gesture feature extraction module, this invention utilizes a CNN to extract useful gesture features from wireless signals. The output features of this feature extractor can be expressed as:
[0028] Z = F d (X;θ d ) = CNN(X; θ d )
[0029] Where θ d Used to represent feature extractor F d (X;θ d ) parameters, such as weights and biases.
[0030] (2) Calculate the loss of the classifier and the multi-domain discriminator. The gesture classifier is connected to the feature extractor, helping the model obtain the most useful features from the feature extractor by maximizing the accuracy of gesture label recognition. The probability distribution of gesture label prediction is as follows:
[0031]
[0032] Where θ y This represents all the parameters of the gesture classifier, such as weights and biases. The predicted probability of each gesture sample is obtained through the gesture classifier, and then the cross-entropy function is used as the loss function, as shown below:
[0033]
[0034] Where |X s | represents the number of gesture samples. By minimizing the loss, the gesture classifier achieves the highest recognition accuracy.
[0035] A multi-domain adversarial discriminator is used to predict domain labels. By maximizing cross-entropy loss, it reduces the prediction accuracy for the domain, causing the feature extractor to try to deceive the domain discriminator. Through a maximal-minus game-like adversarial process, the feature extractor ultimately extracts Doppler change features independent of the domain. This invention employs a multi-domain discriminator, associating each domain discriminator with a gesture category, which can better adapt to specific patterns in different domains, thereby avoiding negative transitions. The predicted label for the gesture feature domain can be obtained from each domain discriminator.
[0036]
[0037] in, This represents all parameters of the discriminator in the k-th domain, such as weights and biases. It uses probability-weighted data points. Training a multi-domain discriminator allows each discriminator to focus more on relevant data points. The domain discriminator G is calculated by summing the losses of the individual domain discriminators. d Total loss:
[0038]
[0039] Where |X s | represents the number of gesture samples. This is a real domain tag.
[0040] (3) Calculate the total loss of the multi-domain adversarial module. The core idea of the adversarial domain generalization module is to minimize the gesture recognizer loss while simultaneously minimizing the domain discriminator loss; therefore, the final loss function is:
[0041] Loss = Loss y -λLossd
[0042] Here, λ is a weighting parameter. By minimizing the loss function Loss, the domain-independent Doppler variation mode is finally obtained.
[0043] Furthermore, the prototype network initialization in S3 includes a meta-training phase and a meta-testing phase. In the meta-training phase, a large number of gesture samples from the source domain are used to form a training dataset, and the prototype network is initialized using the training dataset. In the subsequent meta-testing phase, a small number of samples from the target domain are used to quickly adapt the well-trained meta-learning model.
[0044] Furthermore, in S4 and S5, a domain generalization task generation scheme based on the Widar3.0 dataset is used during the training phase to generate a large number of domain variations, and the parameters of the prototype network are retrained and updated using the generated tasks. Parameter updates include:
[0045] (1) Each task T i All are divided into support sets and query set Use the support set to train the task-specific parameters of the base network. Use query sets to evaluate task performance and iteratively update the initial parameters θ.
[0046] (2) Perform gradient adjustment on the initial model parameters θ, as shown below:
[0047]
[0048] Where α represents the step size, which is a hyperparameter that controls the learning rate of the model.
[0049] (3) Define the following meta-objective function:
[0050]
[0051] The meta-objective function is optimized using stochastic gradient descent, and the parameter θ is updated as follows:
[0052]
[0053] Here, the metastep size ξ is a hyperparameter. Through the optimization process described above, the parameter θ continuously learns knowledge from different tasks and ultimately remains sensitive to different tasks.
[0054] Furthermore, in step S6, some samples from the new domain are used for domain adaptation to further calibrate the parameters of the training dataset and complete the model training. Specifically, the optimized parameters θ are first used as the initial values of the basic model, and then the parameters θ are updated using a small number of samples from the new domain.
[0055] Another object of the present invention is to provide a Wi-Fi signal-based cross-domain few-shot gesture recognition system for implementing the aforementioned Wi-Fi signal-based cross-domain few-shot gesture recognition method, the system comprising:
[0056] Gesture data signal acquisition module: Uses commercial transceivers based on Wi-Fi signals to collect signals. All transceivers are readily available mini desktop computers, and Linux CSI Tool is installed on the devices to record CSI measurement values.
[0057] Signal preprocessing module: used for phase shift correction and elimination of static components and noise in the signal;
[0058] Multi-domain adversarial module: Through the combined action of feature extractor, gesture classifier and domain discriminator, and through a maximum-minimum game adversarial process, Doppler change features independent of the domain are extracted;
[0059] Meta-learning module: It uses meta-learning strategies to teach the basic network how to learn, while using a small number of samples from the target domain to enable the model to have rapid domain adaptation capabilities and complete accurate gesture recognition.
[0060] Another object of the present invention is to provide a computer device, the computer device including a memory and a processor, the memory storing a computer program, the computer program being executed by the processor, causing the processor to perform the steps of the cross-domain few-sample gesture recognition method based on Wi-Fi signals:
[0061] Commercial transceiver equipment based on Wi-Fi signals was used to acquire signals and record CSI measurements on the equipment. Doppler spectrum was extracted using signal processing methods such as principal component analysis and short-time Fourier transform.
[0062] Doppler variation patterns are extracted using a multi-domain adversarial network. During the meta-training phase, the meta-learning model is initialized using the training dataset from the source domain. A task generation scheme based on the training dataset is employed to generate a large number of tasks to simulate different domain variations. The parameters of the prototype network are updated using these generated tasks, while cross-entropy loss is used to evaluate the loss of each task. Finally, some samples are selected from each class in the new domain for domain adaptation to further calibrate the parameters of the training dataset, completing the model training. The trained prototype network is then used for gesture recognition.
[0063] Another object of the present invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the cross-domain few-sample gesture recognition method based on Wi-Fi signals.
[0064] Another objective of this invention is to provide an information data processing terminal for implementing the cross-domain small sample gesture recognition system based on Wi-Fi signals.
[0065] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows:
[0066] (1) This invention proposes a novel cross-domain gesture recognition method that can extract all gesture features using only a pair of Wi-Fi transceivers, significantly reducing device deployment costs. Simultaneously, this invention employs a meta-learning task generation scheme, capable of generating a large number of tasks to simulate different domain variations. For a new target domain, each gesture only requires one sample to quickly adapt to the new domain, significantly reducing data collection overhead. These advantages greatly save on manpower and material resources.
[0067] (2) This invention uses two different filters to eliminate static components and noise in the signal while applying PCA to retain only the significant components caused by target motion, which can further remove noise and compress the data dimension of CSI. In addition, through multi-domain adversarial networks, it can eliminate environmental information while retaining effective information to the maximum extent, and has good dynamic adaptability to the environment. When crossing different domains such as user, environment, location, and direction, it can achieve accurate cross-domain gesture recognition, making up for the shortcomings of traditional gesture recognition methods in terms of cross-domain accuracy decline.
[0068] (3) The gesture recognition method of this invention adopts an effective feature extraction method, namely a multi-domain adversarial network. The feature extractor focuses on extracting features of Doppler variation patterns; the gesture classifier maximizes prediction accuracy, making the feature extractor inclined to extract more features; while the multi-domain adversarial discriminator is used to predict domain labels, and by maximizing cross-entropy loss, it reduces its prediction accuracy for the domain, constantly forcing the feature extractor to try its best to deceive the domain discriminator. Through the above-mentioned minimax game adversarial process, the feature extractor finally extracts Doppler variation features that are independent of the domain. Finally, multiple experiments have proven that this feature extraction method has a good effect on improving the accuracy of gesture recognition. Attached Figure Description
[0069] Figure 1 This is a block diagram of the cross-domain small-sample gesture recognition method based on Wi-Fi signals provided by the present invention;
[0070] Figure 2 This is a schematic diagram of the antenna selection method provided by the present invention;
[0071] Figure 3These are Doppler spectrum diagrams extracted from the original signal under different conditions provided by the present invention; in the figure, (a) shows user 1 performing a "push-pull" gesture; (b) shows user 1 performing a "scan" gesture; (c) shows user 1 performing a "push-pull" gesture at a position different from (a); and (d) shows user 2 performing a "push-pull" gesture.
[0072] Figure 4 This is a schematic diagram of the multi-domain confrontation module provided by the present invention;
[0073] Figure 5 These are the environmental characteristics and sensing areas of different rooms provided by this invention;
[0074] Figure 6 This is a scenario diagram of device deployment and domain configuration in the sensing area provided by the present invention;
[0075] Figure 7 This is a gesture recognition accuracy map across five different domain factors provided by the present invention;
[0076] Figure 8 This is an ablation experiment accuracy diagram of the domain adversarial model provided by the present invention;
[0077] Figure 9 This is an ablation experiment accuracy diagram of the meta-learning model provided by this invention;
[0078] Figure 10 This is an experimental result diagram comparing the gesture recognition method provided by this invention with mainstream methods. Detailed Implementation
[0079] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0080] Example 1:
[0081] The purpose of this invention is to provide a cross-domain few-shot gesture recognition method based on Wi-Fi signals. This method eliminates irrelevant interference through a multi-domain adversarial network, maximizing the extraction of gesture-related features. Simultaneously, a meta-learning strategy is employed to drive the basic network to learn efficiently with few samples, thereby achieving accurate gesture recognition in new domains. This method can be applied to various system platforms, with the computer processor as the execution entity.
[0082] The following explanation, in conjunction with the accompanying drawings, further illustrates the following:
[0083] The method flow of the cross-domain few-sample gesture recognition method based on Wi-Fi signals is as follows: Figure 1As shown, the cross-domain few-sample gesture recognition method based on Wi-Fi signals provided in this embodiment of the invention includes the following steps:
[0084] (1) Data preprocessing: In the signal preprocessing module, we use CSI to generate Doppler change patterns caused by gestures. The Doppler spectrogram can provide high-precision channel feature information, which can be used to distinguish different gestures, thereby improving the accuracy of gesture recognition.
[0085] Step 1.1: Receive the CSI data collected by the receiving device and model it as follows:
[0086]
[0087] Where H s (f) represents the sum of all static path responses, N d The number of dynamic paths is represented by (η+β), which is the phase offset caused by carrier frequency offset, sampling frequency offset, and packet detection delay.
[0088] Step 1.2: Eliminate phase shift by calculating the conjugate multiplication between antennas. For antenna selection, the coefficient ρ, which is the ratio of the amplitude to the variance of the CSI signal, is used. m Indicators for antenna selection:
[0089]
[0090] Where var and mean represent k th subcarrier m th The variance and mean of the antenna amplitude readings. Based on this calculation, we selected the antenna pairs with the highest and lowest ratio coefficients. For example... Figure 2 As shown, 1 th The antenna has the largest amplitude but a smaller variance, while 3 th The antenna has the largest variance but the smallest amplitude. Therefore, the first and third antennas have the largest and smallest amplitude and variance scaling factors, respectively, which will be used as the final antenna selection criteria.
[0091] Step 1.3: Use two different filters to eliminate static components and noise in the signal.
[0092] Step 1.4: Extract the following using principal component analysis and short-time Fourier transform: Figure 3 The Doppler spectrum diagram shown.
[0093] (2) Doppler variation pattern extraction: This was achieved using methods such as... Figure 4 The multi-domain adversarial network shown extracts Doppler variation patterns through the combined action of a feature extractor, a gesture classifier, and a domain discriminator.
[0094] Step 2.1: Feature Extraction. In the gesture feature extraction module, this invention utilizes a CNN to extract useful gesture features from the wireless signal. The output features of this feature extractor can be represented as:
[0095] Z = F d (X;θ d ) = CNN(X; θ d )
[0096] Where θ d Parameters used to represent feature extractors, such as weights and biases.
[0097] Step 2.2: Calculate the loss of the classifier and the multi-domain discriminator. The gesture classifier is connected to the feature extractor, helping the model obtain the most useful features from the feature extractor by maximizing the accuracy of gesture label recognition. The probability distribution of gesture label predictions is as follows:
[0098]
[0099] Where θ y This represents all the parameters of the gesture classifier, such as weights and biases. The predicted probability of each gesture sample is obtained through the gesture classifier, and then the cross-entropy function is used as the loss function, as shown below:
[0100]
[0101] Where |X s | represents the number of gesture samples. By minimizing the loss, the gesture classifier achieves the highest recognition accuracy.
[0102] The purpose of a multi-adversarial domain discriminator is to eliminate differences between domains and help learn transferable features. Essentially, it's a process of adversarial learning. Feature extraction aims to confuse the domain discriminator as much as possible by extracting domain-invariant features related to gestures, while the domain discriminator's goal is to distinguish which domain a sample comes from. This invention employs a multi-domain discriminator, associating each domain discriminator with a gesture category. This allows for better adaptation to specific patterns in different domains, thus avoiding negative transfer. Predicted labels for the gesture feature domains can be obtained from each domain discriminator.
[0103]
[0104] in, This represents all parameters of the discriminator in the k-th domain, such as weights and biases. It uses probability-weighted data points. Training a multi-domain discriminator allows each discriminator to focus more on relevant data points. The domain discriminator G is calculated by summing the losses of the individual domain discriminators. d Total loss:
[0105]
[0106] Where |X s | represents the number of gesture samples. This is a real domain tag.
[0107] Step 2.3: Calculate the total loss of the multi-domain adversarial module. The core idea of the adversarial domain generalization module is to minimize the gesture recognizer loss while simultaneously minimizing the domain discriminator loss; therefore, the final loss function is:
[0108] Loss = Loss y -λLoss d
[0109] Here, λ is a weighting parameter. By minimizing the loss function Loss, the domain-independent Doppler variation mode is finally obtained.
[0110] (3) Prototype network initialization: In the meta-training stage, the meta-learning model is initialized using the training dataset from the source domain;
[0111] (4) Meta-learning Multi-task Generation: A task generation scheme based on the training dataset is adopted to generate a large number of tasks to simulate different domain changes and teach the basic network how to adapt to a new domain. Specifically, this invention utilizes the Widar3.0 public dataset, which contains multiple domain factors, including user, location, and orientation. This invention generates tasks on the training dataset, with each task corresponding to a single domain. Therefore, the final structural domain range is U1L1O1-U m L n O k By employing a task generation scheme, this invention can obtain diverse domain variations with a limited training dataset, enabling the model to effectively learn from a small number of samples and thus providing good cross-domain results.
[0112] (5) Parameter update: Update the parameters of the prototype network using the multiple tasks generated in step (4), and use cross-entropy loss to evaluate the loss of each task;
[0113] Specifically, each task T i They are all divided into support sets and query set Use the support set to train the task-specific parameters of the base network. The query set is used to evaluate task performance and the initial parameters θ are updated iteratively. Then, gradient adjustment is performed on the initial model parameters θ, as follows:
[0114]
[0115] Where α represents the step size, which is a hyperparameter that controls the learning rate of the model.
[0116] To find the parameter θ that minimizes the loss of all tasks, the following meta-objective function is defined:
[0117]
[0118] The meta-objective function is then optimized using stochastic gradient descent, and the parameter t is updated as follows:
[0119]
[0120] Here, the metastep size ξ is a hyperparameter. Through the optimization process described above, the parameter θ continuously learns knowledge from different tasks and ultimately remains sensitive to different tasks.
[0121] (6) New Domain Adaptation: To achieve accurate gesture recognition in the new domain, some samples are selected from each class in the new domain for domain adaptation to further calibrate the parameters of the training dataset and complete the model training. Specifically, the optimized parameters θ are first used as the initial values of the basic model, and then the parameters θ are updated using a small number of samples from the new domain.
[0122] Example 2:
[0123] The gesture recognition system was experimentally verified and evaluated in an indoor laboratory. Four different experiments were conducted to verify the cross-domain recognition accuracy, robustness, and advantages compared to other methods. The transceiver consisted of one transmitter and at least three receivers. All transceivers were readily available mini desktop computers (physical dimensions 170mm × 170mm) equipped with an Intel 5300 wireless network card. Linux CSI tools were installed on the devices to facilitate the recording of CSI measurements, and the devices were configured to operate in monitoring mode. Furthermore, extensive gesture recognition experiments were conducted in three indoor environments: an empty classroom with tables and chairs, a spacious hallway, and an office furnished with sofas and tables. Figure 5 It shows the general environmental characteristics and sensing areas of different rooms. Figure 6 This demonstrates a typical example of device deployment and domain configuration in a 2m x 2m square area. In the experiment, all devices were held at a height of 110 cm, allowing users of varying heights to comfortably make gestures. A total of 16 volunteers (12 men and 4 women) of varying heights (185 cm to 155 cm) and body types participated in the experiment, ranging in age from 22 to 28 years old.
[0124] The cross-domain few-sample gesture recognition method based on Wi-Fi signals provided in this embodiment of the invention includes the following steps:
[0125] (1) Data preprocessing: In the signal preprocessing module, we use CSI to generate Doppler change patterns caused by gestures. The Doppler spectrogram can provide high-precision channel feature information, which can be used to distinguish different gestures, thereby improving the accuracy of gesture recognition.
[0126] Step 1.1: Receive the CSI data collected by the receiving device and model it as follows:
[0127]
[0128] Where H s (f) represents the sum of all static path responses, N d The number of dynamic paths is represented by (η+β), which is the phase offset caused by carrier frequency offset, sampling frequency offset, and packet detection delay.
[0129] Step 1.2: Eliminate phase shift by calculating the conjugate multiplication between antennas. For antenna selection, the coefficient ρ, which is the ratio of the amplitude to the variance of the CSI signal, is used. m Indicators for antenna selection:
[0130]
[0131] Where var and mean represent k th subcarrier m th The variance and mean of the antenna amplitude readings. Based on this calculation, we selected the antenna pairs with the highest and lowest ratio coefficients. For example... Figure 2 As shown, 1 th The antenna has the largest amplitude but a smaller variance, while 3 th The antenna has the largest variance but the smallest amplitude. Therefore, the first and third antennas have the largest and smallest amplitude and variance scaling factors, respectively, which will be used as the final antenna selection criteria.
[0132] Step 1.3: Use two different filters to eliminate static components and noise in the signal.
[0133] Step 1.4: Extract the following using principal component analysis and short-time Fourier transform: Figure 3 The Doppler spectrum diagram shown.
[0134] (2) Doppler variation pattern extraction: This was achieved using methods such as... Figure 4 The multi-domain adversarial network shown extracts Doppler variation patterns through the combined action of a feature extractor, a gesture classifier, and a domain discriminator.
[0135] Step 2.1: Feature Extraction. In the gesture feature extraction module, this invention utilizes a CNN to extract useful gesture features from the wireless signal. The output features of this feature extractor can be represented as:
[0136] Z = F d (X;θ d ) = CNN(X; θ d )
[0137] Where θ d Parameters used to represent feature extractors, such as weights and biases.
[0138] Step 2.2: Calculate the loss of the classifier and the multi-domain discriminator. The gesture classifier is connected to the feature extractor, helping the model obtain the most useful features from the feature extractor by maximizing the accuracy of gesture label recognition. The probability distribution of gesture label predictions is as follows:
[0139]
[0140] Where θ y This represents all the parameters of the gesture classifier, such as weights and biases. The predicted probability of each gesture sample is obtained through the gesture classifier, and then the cross-entropy function is used as the loss function, as shown below:
[0141]
[0142] Where |X s | represents the number of gesture samples. By minimizing the loss, the gesture classifier achieves the highest recognition accuracy.
[0143] The purpose of a multi-adversarial domain discriminator is to eliminate differences between domains and help learn transferable features. Essentially, it's a process of adversarial learning. Feature extraction aims to confuse the domain discriminator as much as possible by extracting domain-invariant features related to gestures, while the domain discriminator's goal is to distinguish which domain a sample comes from. This invention employs a multi-domain discriminator, associating each domain discriminator with a gesture category. This allows for better adaptation to specific patterns in different domains, thus avoiding negative transfer. Predicted labels for the gesture feature domains can be obtained from each domain discriminator.
[0144]
[0145] in, This represents all parameters of the discriminator in the k-th domain, such as weights and biases. It uses probability-weighted data points. Training a multi-domain discriminator allows each discriminator to focus more on relevant data points. The domain discriminator G is calculated by summing the losses of the individual domain discriminators. d Total loss:
[0146]
[0147] Where |X s | represents the number of gesture samples. This is a real domain tag.
[0148] Step 2.3: Calculate the total loss of the multi-domain adversarial module. The core idea of the adversarial domain generalization module is to minimize the gesture recognizer loss while simultaneously minimizing the domain discriminator loss; therefore, the final loss function is:
[0149] Loss = Loss y -λLoss d
[0150] Here, λ is a weighting parameter. By minimizing the loss function Loss, the domain-independent Doppler variation mode is finally obtained.
[0151] (3) Prototype network initialization: In the meta-training stage, the meta-learning model is initialized using the training dataset from the source domain;
[0152] (4) Meta-learning Multi-task Generation: A task generation scheme based on the training dataset is adopted to generate a large number of tasks to simulate different domain changes and teach the basic network how to adapt to a new domain. Specifically, this invention utilizes the Widar3.0 public dataset, which contains multiple domain factors, including user, location, and orientation. This invention generates tasks on the training dataset, with each task corresponding to a single domain. Therefore, the final structural domain range is U1L1O1-U m L n O k By employing a task generation scheme, this invention can obtain diverse domain variations with a limited training dataset, enabling the model to effectively learn from a small number of samples and thus providing good cross-domain results.
[0153] (5) Parameter update: Update the parameters of the prototype network using the multiple tasks generated in step (4), and use cross-entropy loss to evaluate the loss of each task;
[0154] Specifically, each task T i They are all divided into support sets and query set Use the support set to train the task-specific parameters of the base network. The query set is used to evaluate task performance and the initial parameters θ are updated iteratively. Then, gradient adjustment is performed on the initial model parameters θ, as follows:
[0155]
[0156] Where α represents the step size, which is a hyperparameter that controls the learning rate of the model.
[0157] To find the parameter θ that minimizes the loss of all tasks, the following meta-objective function is defined:
[0158]
[0159] The meta-objective function is then optimized using stochastic gradient descent, and the parameter θ is updated as follows:
[0160]
[0161] Here, the metastep size ξ is a hyperparameter. Through the optimization process described above, the parameter θ continuously learns knowledge from different tasks and ultimately remains sensitive to different tasks.
[0162] (6) New Domain Adaptation: To achieve accurate gesture recognition in the new domain, some samples are selected from each class in the new domain for domain adaptation to further calibrate the parameters of the training dataset and complete the model training. Specifically, the optimized parameters θ are first used as the initial values of the basic model, and then the parameters θ are updated using a small number of samples from the new domain.
[0163] Example 3: Verification of gesture recognition effect:
[0164] In a specific implementation of this invention, the proposed neural network is implemented using the PyTorch deep learning framework, and the network is optimized and trained using the Adam optimizer and cosine annealing learning rate, with the number of samples per batch set to 32.
[0165] Performance evaluation was conducted using the CSI-based public gesture dataset Widar 3.0. The Widar 3.0 dataset consists of two subsets; this invention uses the first subset for experiments. This subset consists of 16 different users performing six common gestures in five locations and five directions across three different rooms. The six gestures are push / pull, sweep, clap, swipe, circle, and zigzag. Simultaneously, one transmitter and six receivers were deployed in each environment, with the transmitter broadcasting Wi-Fi packets at a rate of 1000 packets per second. Detailed experimental evaluations are as follows:
[0166] Experiment 1: The cross-domain gesture recognition accuracy was evaluated based on five different domain factors, including transceiver deployment, location, orientation, personnel, and environment. Overall performance was as follows: Figure 7As shown, when the number of target samples for each gesture is only 1, the gesture recognition accuracies across transceiver deployments, locations, and directions are 90.4%, 88.25%, and 87.4%, respectively. For cross-user and cross-environment experiments, the gesture recognition accuracy is slightly lower, indicating that environmental and user domain factors have a significant impact on feature distribution. When the number of target samples for each gesture is increased to 3, the cross-domain accuracy is greatly improved; therefore, the gesture recognition system designed in this invention has good application prospects.
[0167] Experiment 2: In this section, ablation experiments were conducted to evaluate the effectiveness of the domain adversarial model. Specifically, an evaluation experiment with a target sample size of 1 was performed under the same experimental settings, comparing the accuracy with and without the domain adversarial module. The results are as follows: Figure 8 As shown, the features extracted by the domain adversarial module have better recognition accuracy under the classification of the meta-learner. This is because the CSI signal without domain adversarial processing contains a lot of environment-related interference information, while the domain adversarial module can greatly eliminate the influence of domain information, thereby improving the accuracy of gesture recognition.
[0168] Experiment 3: In this section, ablation experiments were conducted to evaluate the effectiveness of the meta-learning model. Specifically, an evaluation experiment with a target sample size of 1 was performed under the same experimental settings, comparing the accuracy with and without the meta-learning module. The results are as follows: Figure 9 As shown, for cross-domain few-shot gesture recognition, using the meta-learning strategy is consistently superior to not using it.
[0169] Experiment 4: In this section, three common gesture recognition methods (SignFi, EI, and RF-Net) were selected for comparison with the method of this invention. Specifically, the cross-location gesture recognition accuracy was evaluated under the same experimental settings for target sample sizes of 1 and 3, respectively. The results are as follows: Figure 10 As shown, the cross-domain few-sample gesture recognition method proposed in this invention clearly has better performance. Although EI and RF-Net have cross-domain capabilities, they still cannot accurately recognize gestures when there are only a few target samples. This demonstrates the superiority of the method in this invention.
[0170] In this embodiment, a computer device includes a memory and a processor. The memory is used to store a program that supports the processor in executing the above-described cross-domain small-sample gesture recognition method based on Wi-Fi signals. The processor is configured to execute the program stored in the memory.
[0171] In this embodiment, a computer-readable storage medium stores a computer program, which, when executed by a processor, performs the steps of the above-described method for cross-domain few-sample gesture recognition based on Wi-Fi signals.
[0172] It should be noted that embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.
[0173] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A cross-domain few-sample gesture recognition method based on Wi-Fi signals, characterized in that, include: S1: Data preprocessing: Phase shift is eliminated using conjugate multiplication, static components and noise in the signal are eliminated using two different filters, and then the Doppler spectrum is extracted using principal component analysis and short-time Fourier transform. S2: Doppler variation pattern extraction: A multi-domain adversarial network was used to extract the Doppler variation pattern through the combined action of a feature extractor, a gesture classifier, and a multi-adversarial domain discriminator. S3: Prototype network initialization: In the meta-training phase, the meta-learning model is initialized using the training dataset from the source domain. S4: Meta-learning multi-task generation: Employs a task generation scheme based on the training dataset to generate a large number of tasks to simulate different domain changes and train the prototype network to adapt to the new domain. S5: Parameter Update: Update the parameters of the prototype network using multiple tasks generated by S4, and use cross-entropy loss to evaluate the loss of each task; S6: New Domain Adaptation: In order to achieve accurate gesture recognition in the new domain, some samples are selected from each class in the new domain for domain adaptation to further calibrate the parameters of the training dataset and complete the training of the model. S2 includes the following steps: Step 2.1: Feature Extraction; Useful gesture features are extracted from the wireless signal using a CNN; The output features of this feature extractor are represented as follows: ; in Parameters used to represent the feature extractor, including weights and biases; Step 2.2: Calculate the loss of the gesture classifier and the multi-adversarial domain discriminator; the gesture classifier is connected to the feature extractor, and by maximizing the accuracy of gesture label recognition, it helps the model obtain the most useful features from the feature extractor; the probability distribution of gesture label prediction is as follows: ; in All parameters, including weights and biases, are used to represent the gesture classifier. The predicted label probability for each gesture sample is obtained through the gesture classifier, and then the cross-entropy function is used as the loss function, as shown below: ; in It represents the number of gesture samples; by minimizing the loss, the gesture classifier achieves the maximum recognition accuracy. The purpose of a multi-adversarial domain discriminator is to eliminate differences between domains and help learn transferable features. It is essentially an adversarial learning process. Feature extraction uses gesture-related domain-invariant features to confuse the multi-adversarial domain discriminator as much as possible. The goal of the multi-adversarial domain discriminator is to distinguish which domain a sample comes from as much as possible. A multi-adversarial domain discriminator is employed, with each discriminator associated with a gesture category. This better adapts to specific patterns in different domains, thus avoiding negative transitions. Predicted labels for the gesture feature domain are obtained from each discriminator. : ; in, This represents all parameters of the k-th multi-adversarial domain discriminator, including weights and biases; using probability-weighted data points. Train the multi-adversarial domain discriminators so that each discriminator focuses more on relevant data points; calculate the multi-adversarial domain discriminator's loss by summing the losses of the individual discriminators. Total loss: ; in The number of gesture samples, For real domain tags; Step 2.3: Minimize the gesture classifier loss and simultaneously minimize the multi-adversarial domain discriminator loss; therefore, the final loss function is: ; Where λ is the weighting parameter; by minimizing the loss function Ultimately, a domain-independent Doppler variation pattern is obtained.
2. The cross-domain small-sample gesture recognition method based on Wi-Fi signals as described in claim 1, characterized in that, In step S1, conjugate multiplication is used to eliminate phase shift. First, phase shift is eliminated by calculating the conjugate multiplication between antennas, and the coefficient of the ratio of amplitude to variance of the CSI signal is used as an indicator for antenna selection. In addition, high-pass and low-pass filters are used to eliminate static components and high-frequency noise. To further denoise and compress the CSI data dimension, PCA is applied to retain the significant components caused by target motion. Finally, short-time Fourier transform is performed on the main components extracted by PCA to obtain the Doppler spectrum.
3. The cross-domain few-sample gesture recognition method based on Wi-Fi signals as described in claim 1, characterized in that, In S2, the Doppler change pattern is extracted. The feature extractor focuses on extracting features of the Doppler change pattern. The gesture classifier maximizes the prediction accuracy, making the feature extractor tend to extract more features. The multi-domain adversarial discriminator is used to predict domain labels. By maximizing the cross-entropy loss, it reduces the prediction accuracy of the domain, making the feature extractor try its best to deceive the multi-domain adversarial discriminator. After a game of extreme small and large, the feature extractor finally extracts Doppler change features that are independent of the domain.
4. The cross-domain small-sample gesture recognition method based on Wi-Fi signals as described in claim 1, characterized in that, The prototype network initialization in S3 includes a meta-training phase and a meta-testing phase. In the meta-training phase, a large number of gesture samples from the source domain are used to form a training dataset, and the prototype network is initialized using the training dataset. In the subsequent meta-testing phase, a small number of samples from the target domain are used to quickly adapt the well-trained meta-learning model.
5. The cross-domain small-sample gesture recognition method based on Wi-Fi signals as described in claim 1, characterized in that, In S5, a domain generalization task generation scheme based on the Widar3.0 dataset is used during the training phase to generate a large number of domain variations and to retrain and update the parameters of the prototype network using the generated tasks.
6. The cross-domain few-sample gesture recognition method based on Wi-Fi signals as described in claim 1, characterized in that, In step S6, some samples from the new domain are used for domain adaptation to calibrate the parameters of the training dataset. Specifically, the updated prototype network parameters are first used as the initial values of the prototype network, and then the parameters are updated using a few samples from the new domain.
7. A cross-domain few-shot gesture recognition system based on Wi-Fi signals, implementing the cross-domain few-shot gesture recognition method based on Wi-Fi signals as described in any one of claims 1 to 6, characterized in that, The system includes: Gesture data signal acquisition module: Uses commercial transceivers based on Wi-Fi signals to collect signals. All transceivers are readily available mini desktop computers, and Linux CSI Tool is installed on the devices to record CSI measurement values. Signal preprocessing module: used for phase shift correction and elimination of static components and noise in the signal; Multi-domain adversarial module: Through the combined action of feature extractor, gesture classifier and multi-adversarial domain discriminator, Doppler change features independent of domain are extracted after a maximum-minimum game adversarial process; Meta-learning module: It uses meta-learning strategies to teach the prototype network how to learn, while using a small number of samples from the target domain to enable the model to have rapid domain adaptation capabilities and complete accurate gesture recognition.
8. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the cross-domain few-sample gesture recognition method based on Wi-Fi signals as described in any one of claims 1 to 6: Commercial transceiver equipment based on Wi-Fi signals was used to acquire signals and record CSI measurements on the equipment. Doppler spectrum was extracted using signal processing methods such as principal component analysis and short-time Fourier transform. Doppler variation patterns were extracted using a multi-domain adversarial network. In the meta-training phase, the meta-learning model is initialized using the training dataset of the source domain; a task generation scheme based on the training dataset is adopted to generate a large number of tasks to simulate different domain changes and use the generated multiple tasks to update the parameters of the prototype network, while using cross-entropy loss to evaluate the loss of each task; finally, some samples are selected from each class of the new domain for domain adaptation to further calibrate the parameters of the training dataset and complete the training of the model; the trained prototype network is used for gesture recognition.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the steps of the cross-domain few-sample gesture recognition method based on Wi-Fi signals as described in any one of claims 1 to 6.
10. An information data processing terminal, characterized in that, The information data processing terminal is used to implement the cross-domain small sample gesture recognition system based on Wi-Fi signals as described in claim 7.
Citation Information
Patent Citations
Domain self-adaptive Wi-Fi gesture recognition method based on discrete wavelet transform
CN112733609A
Switch cabinet partial discharge audio identification method based on deep meta learning
CN116110425A