WiFi open set gesture recognition method based on symbiotic metric learning and uncertainty perception

By using symbiotic measurement and uncertainty perception methods, the feature space topology of Wi-Fi gesture recognition technology is optimized, which improves the recognition accuracy and robustness in complex environments, solves the problems of loose topology and weak noise resistance, and achieves stable cross-domain recognition.

CN121708657APending Publication Date: 2026-03-20CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610079003.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-21
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing Wi-Fi gesture recognition technology suffers from problems such as loose feature space topology, weak noise resistance, and insufficient cross-domain generalization performance in complex dynamic environments, resulting in blurred rejection boundaries for unknown classes, poor noise resistance, and decreased cross-domain recognition performance.

Method used

A method based on coexistence metric and uncertainty perception is adopted. Dual-view samples are generated through random transformation, a multi-branch neural network is constructed, and a backbone network with shared weights and a feature embedding branch are combined. The feature space is optimized by majority voting mechanism and coexistence fusion loss. Doppler spectrum physical constraints and feature noise injection are introduced to optimize the decision boundary and improve recognition accuracy and robustness.

Benefits of technology

It significantly improves recognition accuracy and robustness in low signal-to-noise ratio environments, enhances the ability to reject unknown gestures, and maintains stable recognition performance in cross-domain scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121708657A_ABST
    Figure CN121708657A_ABST
Patent Text Reader

Abstract

The invention belongs to a pattern recognition technology, and particularly relates to a WiFi open set gesture recognition method based on symbiotic metric learning and uncertainty perception, which comprises the following steps: constructing a feature memory library according to original feature vectors of known gesture image samples, acquiring K neighbor view sample sets closest to the cosine distance of the original feature vector of the to-be-identified image from the feature memory library, determining candidate prediction tags of a to-be-detected sample by adopting a majority voting mechanism, and calculating a neighbor distance score value of the to-be-detected sample; and if the neighbor distance score value is greater than a set threshold value, taking the candidate tag as the tag of the to-be-detected sample, otherwise, judging that the to-be-detected sample is an unknown gesture. Aiming at the problems that Wi-Fi signal environment noise interference is strong and feature distribution is loose in an open set scene, the generalization ability of the model in a cross-domain scene is effectively improved through a double-view consistency constraint and symbiotic fusion mechanism, and the rejection rate of unknown gestures and the recognition precision of known gestures are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless sensing and pattern recognition technology, and specifically relates to a WiFi open set gesture recognition method based on symbiotic metric learning and uncertainty perception. Background Technology

[0002] With the rapid development of the Internet of Things (IoT) and ubiquitous computing technologies, contactless gesture recognition technology based on Channel State Information (CSI) of commercial Wi-Fi devices has become an important technological direction in the field of human-computer interaction due to its advantages such as non-intrusiveness, privacy protection, and all-weather monitoring. However, in practical application scenarios, systems inevitably encounter abnormal gestures or non-target actions that are not defined in the training set. Therefore, it is crucial to have open-set gesture recognition technology that can accurately recognize known types of gestures while effectively detecting and rejecting unknown types of gestures.

[0003] In existing technologies, implementation schemes for Wi-Fi open set gesture recognition typically include three stages: data preprocessing, feature extraction, and classification decision. Specifically, the original CSI data is first preprocessed to remove noise, such as by using CSI quotient calculation to eliminate phase noise or extracting the Doppler frequency shift spectrum; then, a convolutional neural network or recurrent neural network model is constructed to extract high-dimensional features; finally, fully connected layers or metric learning-based methods are used for feature clustering, and whether a gesture is unknown is determined by setting a classification confidence threshold or calculating the distance between the sample and known class centers. Despite extensive research in open set recognition, several challenges remain:

[0004] 1. Loose feature space topology leads to blurred rejection boundaries for unknown classes. Existing metric learning or classification methods often focus only on the separability of samples during optimization, neglecting the compactness of the overall feature space topology. In the absence of strong constraints, the feature distribution of known classes tends to exhibit loose strip-like, ring-like, or irregular manifold structures, rather than compact spherical clusters. This loose topology causes known classes to occupy an excessively large portion of the feature space, compressing the "unknown space." This makes it easy for a large number of unknown gesture samples to fall into the gaps or edges of the known class distribution, thus causing distance- or threshold-based discrimination to fail.

[0005] 2. Lack of a mechanism to perceive signal uncertainty, resulting in poor noise robustness. Wi-Fi signals are inherently susceptible to interference, affected by multipath effects and hardware noise floor, resulting in a large number of highly uncertain noise samples in the collected CSI data. Existing deep learning models typically treat all input data equally, lacking the ability to quantify the confidence level of samples. When the model forces a fit to these highly uncertain noise samples, it distorts the decision boundary, not only reducing the accuracy of recognizing known gestures but also exacerbating the risk of misjudging unknown noise signals.

[0006] 3. Lack of tolerance space in decision boundaries and insufficient cross-domain generalization performance. Existing technologies often pursue the ultimate fit of feature points to the source domain data during training, resulting in the learned decision boundary closely following the data distribution. When minor changes in the application scenario (such as personnel movement or changes in environmental layout) cause signal feature drift, test samples are highly likely to cross this closely spaced boundary. Due to the lack of proactive feature enhancement or boundary expansion mechanisms, existing models struggle to reserve sufficient safety margins in the feature space, leading to a significant decrease in open-set recognition performance in cross-domain scenarios. Summary of the Invention

[0007] To address the problems of loose feature space topology, weak noise resistance, and insufficient cross-domain generalization performance in existing Wi-Fi gesture recognition technologies under complex dynamic environments, this invention provides a Wi-Fi open-set gesture recognition method based on co-occurrence metric and uncertainty awareness, specifically including the following steps:

[0008] Preprocessed Wi-Fi Channel State Information (CSI) image samples are obtained, and random transformation techniques are used to generate first and second view samples with identical content but different spatial topology.

[0009] A multi-branch neural network including shared weight units and feature embedding branches is constructed. The first view sample and the second view sample are input into the backbone network with shared weights to extract feature maps. The feature embedding branch uses the embedding structure of cascaded global adaptive average pooling layer, fully connected layer and L2 normalization layer to process the feature maps and obtain the original feature vectors of the first view sample and the second view sample respectively.

[0010] The original feature vectors of known labeled image samples are obtained to construct a feature memory library. The set of K nearest neighbor view samples that are closest to the cosine distance of the original feature vector of the image to be identified is obtained from the feature memory library.

[0011] Based on the labels of K nearest neighbor view samples, a majority voting mechanism is used to determine the candidate predicted labels of the sample to be tested.

[0012] Based on the candidate labels, the minimum Euclidean distance between the test sample and its similar label samples and the minimum Euclidean distance between its dissimilar label samples are calculated, and the ratio of the two minimum distances is used as the nearest neighbor distance score.

[0013] If the nearest neighbor distance score is greater than the set threshold, the candidate label is used as the label of the sample to be tested; otherwise, the sample to be tested is judged as an unknown gesture.

[0014] This invention utilizes CSI quotient calculation preprocessing and dual-view... Figure 1The triple noise reduction mechanism of consistency constraints and feature noise injection, combined with a confidence weighting strategy based on prediction entropy, effectively filters out multipath effects, environmental noise, and outlier interference commonly found in Wi-Fi signals. This allows the model to focus solely on the essential motion features of gestures during feature extraction, significantly improving the system's robustness in low signal-to-noise ratio environments. Simultaneously, this invention innovatively utilizes a co-existing fusion loss comprising prediction guidance (NCA) and label-aware (AaD). Through real label correction and a low-rejection factor strategy, it forcibly reshapes the feature space from a loose strip-like or manifold distribution in existing technologies to a compact spherical island distribution. This optimized topology significantly increases the safety margin and separability between known and unknown classes. Combined with a minimum distance-based discrimination mechanism, this greatly enhances the ability to reject unknown gestures and the recognition accuracy of known gestures. Furthermore, by introducing Doppler spectrum physical constraints through a physical reconstruction branch, the network is forced to learn motion patterns independent of the environment. Combined with the decision boundary tolerance space provided by feature noise injection, the model maintains stable recognition performance even in cross-domain scenarios such as changes in user location or environmental layout. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the framework of the Wi-Fi open set gesture recognition method based on symbiosis measurement and uncertainty perception of the present invention;

[0016] Figure 2 This is a flowchart of the Wi-Fi open set gesture recognition method based on symbiosis measurement and uncertainty perception of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] This invention provides a Wi-Fi open-set gesture recognition method based on symbiosis metric and uncertainty awareness, specifically including the following steps:

[0019] Preprocessed Wi-Fi Channel State Information (CSI) image samples are obtained, and random transformation techniques are used to generate first and second view samples with identical content but different spatial topology.

[0020] A multi-branch neural network including shared weight units and feature embedding branches is constructed. The first view sample and the second view sample are input into the backbone network with shared weights to extract feature maps. The feature embedding branch uses the embedding structure of cascaded global adaptive average pooling layer, fully connected layer and L2 normalization layer to process the feature maps and obtain the original feature vectors of the first view sample and the second view sample respectively.

[0021] The original feature vectors of known labeled image samples are obtained to construct a feature memory library. The set of K nearest neighbor view samples that are closest to the cosine distance of the original feature vector of the image to be identified is obtained from the feature memory library.

[0022] Based on the labels of K nearest neighbor view samples, a majority voting mechanism is used to determine the candidate predicted labels of the sample to be tested.

[0023] Based on the candidate labels, the minimum Euclidean distance between the test sample and its similar label samples and the minimum Euclidean distance between its dissimilar label samples are calculated, and the ratio of the two minimum distances is used as the nearest neighbor distance score.

[0024] If the nearest neighbor distance score is greater than the set threshold, the candidate label is used as the label of the sample to be tested; otherwise, the sample to be tested is judged as an unknown gesture.

[0025] like Figures 1-2 In this embodiment, the present invention’s Wi-Fi open-set gesture recognition method based on symbiosis measurement and uncertainty perception is divided into four parts: data acquisition and enhancement module, three-branch neural network processing module, symbiotic fusion loss optimization module, and open set inference and result output module. Only the following description is given. (I) Data acquisition and enhancement module.

[0026] In this embodiment, the data acquisition and enhancement module includes a raw CSI signal acquisition unit, a CSI quotient preprocessing and phase map generation unit, and a dual-view data enhancement unit. The CSI quotient preprocessing and phase map generation unit is used to execute an antenna selection strategy based on amplitude variance ratio and CSI quotient calculation, eliminate phase noise, and generate a two-dimensional CSI phase image. The dual-view data enhancement unit is used to perform two independent random transformations on the phase image to generate a first view enhancement sample V1 and a second view enhancement sample V2.

[0027] Specifically, this embodiment first uses an antenna selection strategy to select data from the original CSI data stream. Selecting a reference antenna With the target antenna Perform complex quotient operation, that is, iterate through all antennas m at the receiver and calculate the ratio of the mean to the variance of the CSI amplitude for each antenna, which is defined as the stability score. :

[0028]

[0029] in, This represents the amplitude magnitude of the m-th antenna. This represents the mean over the time dimension. Indicates variance.

[0030] Further, select The largest antenna as the reference antenna (Corresponding to static path dominance), select The smallest antenna as the target antenna (Following dynamic gestures). Perform complex division to calculate the CSI quotient. And extract phase information to generate a phase image. Specifically, it is expressed as follows:

[0031]

[0032]

[0033] in, The phase information of the denoised CSI quotient sequence can be approximately mapped to the gesture-induced phase caused by human motion. ; The complex CSI sequence for determining the target antenna using an antenna selection strategy; and This represents the complex CSI sequence used to determine the reference antenna via an antenna selection strategy. This indicates the random phase deviation contained in the target antenna; This indicates the random phase deviation contained in the reference antenna; This represents the complex amplitude coefficient remaining after the signal undergoes quotient operation; This indicates taking the modulus of the complex number A; j is the imaginary unit; This indicates that a phase-taking operation is performed on the CSI quotient sequence to extract the gesture-induced phase value; This refers to the `unwrap` function in MATLAB.

[0034] To establish viewpoint independence and robustness of features, this embodiment introduces a symmetric training mechanism. Random data augmentation techniques are used to... Implement two independent transformation operators and The process generates a first view sample V1 and a second view sample V2. These two views share the same gesture semantic label y, but exhibit minor perturbations in spatial morphology, providing an input basis for subsequent consistency constraints. A benchmark Doppler frequency spectrum is then constructed, and the CSI quotient sequence is analyzed. Perform a short-time Fourier transform to generate a reference frequency spectrum in physical space, which, together with the first view sample and the second view sample, constitutes the bimodal input / supervision pair of the model.

[0035] (II) Three-branch neural network processing module. The three-branch neural network processing module is the core of feature extraction and prediction. It specifically includes shared ResNet backbone network units and three parallel functional branches: physical reconstruction branch, feature embedding branch, and gesture classification branch, wherein:

[0036] Shared ResNet backbone network units: used to extract high-dimensional spatial feature maps of the first view sample V1 and the second view sample V2;

[0037] Physical reconstruction branch: contains a deconvolution decoder used to restore the feature map to the Doppler frequency shift spectrum (DFS);

[0038] Feature embedding branch: includes fully connected mapping and normalization units, and innovatively introduces a feature noise injection unit to output robust feature vectors with fault-tolerant margins;

[0039] Gesture classification branch: Contains classifier units that output the predicted probability distribution of gesture categories.

[0040] Specifically, in this embodiment, the input view v flows through a shared backbone network and extracts spatial features, the network extracting a high-dimensional spatial feature map. Subsequently, feature maps Simultaneously, three functional branches are processed in parallel, specifically including:

[0041] The physical reconstruction branch consists of a decoder D composed of multiple deconvolutional layers (DeConv), batch normalization layers (BN), and ReLU activation function layers. Its structure and design logic lie in utilizing the reconstructed Doppler frequency spectrum. This forces the backbone network to capture the physical motion of gestures, rather than overfitting to a static background;

[0042] The feature embedding branch constructs a layer containing a globally adaptive average pooling (GAP) layer and a fully connected layer. and the embedding structure of the L2 normalized layer; first, the feature map Flatten the vector, then map it to a low-dimensional feature vector, and perform L2 normalization to obtain the original feature vector. Specifically, it is expressed as follows:

[0043]

[0044] in, This is the original feature vector; This is the trainable weight matrix for the fully connected layer; It is a globally adaptive average pooling layer; This indicates the calculation of the L2 norm.

[0045] During the training phase, the original feature vector is fed... Injecting random noise following a Gaussian distribution into the vector generates robust feature vectors. The specific calculation formula is as follows:

[0046]

[0047] in, For a random noise vector that follows a standard normal distribution, The noise intensity coefficient is preset (preferably 0.1 in this embodiment). The robust feature vector in this invention... For the calculation of subsequent co-occurrence fusion loss, by constructing in the feature space... The perturbation region centered on the model forces the model to learn a robust decision boundary with a fault-tolerant margin.

[0048] The classification branch uses a fully connected mapping to project features onto the class space and outputs a predicted probability distribution p.

[0049] (III) Symbiotic Fusion Loss Optimization Module.

[0050] The co-occurrence fusion loss optimization module is used to calculate the total loss function during the training phase to update the network parameters. Specifically, it includes a basic supervised loss calculation unit, a co-occurrence fusion loss calculation unit, and a total loss calculation and parameter update unit. The co-occurrence fusion loss calculation unit further includes a prediction-guided neighborhood component analysis (NCA) unit and a label-aware attraction-repulsion (AaD) unit, which calculate weights through the interaction of confidence and labels.

[0051] In the initial training phase of the model, this embodiment establishes discriminative capability through multi-dimensional supervision signals. Specifically, the classification prediction distribution... With real labels Calculate the classification cross-entropy loss between them ,Right now:

[0052]

[0053] in, Predicting distribution for classification With real labels Calculate the classification cross-entropy loss between them; C represents the total number of samples in the current training batch; C represents the total number of predefined known gesture categories. This indicates the true indicator variable that the i-th sample belongs to the c-th category; This represents the predicted probability that the i-th sample in the classification branch belongs to the c-th category.

[0054] Meanwhile, to ensure that the features extracted by the network conform to the physical perception logic, the mean squared error loss between the reconstructed DFS and the true DFS in the physical reconstruction branch is calculated:

[0055]

[0056] in, The mean square error loss represents the difference between the true Doppler frequency shift spectrum of the CSI image sample and the Doppler frequency shift spectrum reconstructed by the physical reconstruction branch; T represents the total length of the time dimension of the gesture signal in the Doppler frequency spectrum space. This represents the true Doppler spectral features, which are pre-extracted from the original channel state information using short-time Fourier transform and used as a supervision label; This represents the reconstructed Doppler spectrum features generated by deconvolution of the spatial feature map in the physical reconstruction branch of a multi-branch neural network.

[0057] Furthermore, this embodiment introduces a consistency loss for dual-view samples v1 and v2. The KL divergence metric is used to measure the similarity of the predicted distributions of two views, specifically expressed as:

[0058]

[0059] in, The consistency loss between the first view sample and the second view sample; To smooth the temperature coefficient, For the Softmax function; This indicates the calculation of the KL divergence; This represents the classification prediction vector output by the backbone network and classification branches of the first view sample; This represents the classification prediction vector output by the backbone network and classification branches of the second view sample.

[0060] This invention employs an interaction between "prediction-guided NCA loss" and "label-aware AaD loss" to dynamically correct the topological structure of the feature space, specifically including:

[0061] The prediction guidance NCA unit first uses the prediction distribution p of sample i. i The information entropy is calculated to quantify the uncertainty of the node in the current feature plane and mapped to a confidence weight. This weight increases as the ability to identify neighboring nodes improves, thus achieving a training evolution from global coarse attraction to high-confidence precise clustering. The confidence weight of the i-th training sample is... Represented as:

[0062] Subsequently, when calculating the neighborhood component analysis loss, this weight is used to reweight similar nearest neighbors to construct the NCA loss function. :

[0063]

[0064] Where N is the total number of samples in the current training batch; Let k be the set of the k nearest neighbors of the i-th training sample in the feature space. Represents a set The label of the j-th sample. The label represents the i-th training sample; Let be the confidence weight of the j-th nearest neighbor sample; Let cosine similarity be used to represent the original feature vector of the i-th training sample and its j-th nearest neighbor sample. The temperature parameter is used. This invention introduces the NCA loss function to ensure that the network automatically filters out noisy samples with high uncertainty (i.e., ambiguous predictions) when learning neighborhood relationships.

[0065] Based on this, the label-aware AaD unit corrects the clustering target using the real label y, implementing enhanced attraction and repulsion operations. The calculation formula is as follows:

[0066]

[0067] in, Let be the set of samples with the same label as the i-th training sample in the nearest neighbor set; Let be the set of samples whose labels are different from those of the i-th training sample in the nearest neighbor set; For set The number of elements in the middle; Let be the confidence level of the i-th training sample. Let be the confidence level of the j-th neighbor sample; As a repulsion factor, the present invention preferably sets a small γ value of 0.1 to avoid manifold breakage caused by strong repulsion, and promotes the formation of a compact spherical cluster structure in the feature space under the balance of attraction and weak repulsion. Let be the confidence weight of the i-th training sample; It is a linear rectifier function.

[0068] To prevent the co-occurrence loss from misleading the model in the early stages before convergence during training, this embodiment employs a linear wamp-up learning strategy. The total loss function is constructed as follows:

[0069]

[0070] in, and These are weighting coefficients that evolve with the number of training epochs, t. (In the warm-up epochs...) The coefficients are set to 0 initially, at which point the network only learns basic classification and physical reconstruction; after the warm-up rounds, the coefficients are adjusted accordingly. The proportion increases linearly. This gradual training logic ensures that the model can refine its topology based on stable features.

[0071] (iv) Open set reasoning and result output module.

[0072] The open set reasoning and result output module is used to make decisions on unknown gestures during the application phase. Specifically, it includes a feature memory construction unit, a nearest neighbor distance ratio (NNDR) calculation unit, a threshold decision logic unit, and a result output unit.

[0073] After the model training is completed, using all known class samples in the training set, the corresponding original feature vectors f are extracted through a shared backbone network and mapped one-to-one with their preset gesture labels y to build a global feature memory. Where M is the total number of samples in the training set. Let m be the gesture label corresponding to the m-th training sample. Let be the original feature vector of the m-th training sample.

[0074] For the hand gesture samples to be tested that have been collected in real time and preprocessed, their original feature vectors are extracted. Using cosine similarity as a metric The distance to each vector in the feature library G is used to retrieve the set of K nearest neighbors with the highest similarity. Based on the nearest neighbor set obtained from the retrieval The distribution of the corresponding labels is statistically analyzed. This embodiment uses a majority voting mechanism to determine the candidate predicted labels for the samples to be tested. The calculation formula is as follows:

[0075]

[0076] in, This is an indicator function. The candidate predicted label... This represents the preset gesture category that best matches the sample within the closed set space.

[0077] Furthermore, regarding the characteristics of the sample to be tested Calculate the minimum Euclidean distance between a sample and its nearest neighbors of the same class. and the minimum Euclidean distance to the nearest neighbor of the different species. :

[0078]

[0079]

[0080] The final nearest neighbor distance score is calculated as follows:

[0081]

[0082] like If a preset threshold is set, the gesture is identified as a known gesture and its corresponding preset gesture category is output. Otherwise, it is judged as an unknown gesture. Among them, It is a limiting constant used to prevent the denominator from being zero.

[0083] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A WiFi open-set gesture recognition method based on symbiotic metric learning and uncertainty awareness, characterized in that, Includes the following steps: Preprocessed Wi-Fi Channel State Information (CSI) image samples are obtained, and random transformation techniques are used to generate first and second view samples with identical content but different spatial topology. A multi-branch neural network including shared weight units and feature embedding branches is constructed. The first view sample and the second view sample are input into the backbone network with shared weights to extract feature maps. The feature embedding branch uses the embedding structure of cascaded global adaptive average pooling layer, fully connected layer and L2 normalization layer to process the feature maps and obtain the original feature vectors of the first view sample and the second view sample respectively. The original feature vectors of known labeled image samples are obtained to construct a feature memory library. The set of K nearest neighbor view samples that are closest to the cosine distance of the original feature vector of the image to be identified is obtained from the feature memory library. Based on the labels of K nearest neighbor view samples, a majority voting mechanism is used to determine the candidate predicted labels of the sample to be tested. Based on the candidate labels, the minimum Euclidean distance between the test sample and its similar label samples and the minimum Euclidean distance between its dissimilar label samples are calculated, and the ratio of the two minimum distances is used as the nearest neighbor distance score. If the nearest neighbor distance score is greater than the set threshold, the candidate label is used as the label of the sample to be tested; otherwise, the sample to be tested is judged as an unknown gesture.

2. The WiFi open set gesture recognition method based on symbiotic metric learning and uncertainty awareness according to claim 1, characterized in that, In the feature embedding branch, the input feature map is flattened and then processed, including: in, This is the original feature vector; This is the trainable weight matrix for the fully connected layer; It is a globally adaptive average pooling layer; This indicates the calculation of the L2 norm.

3. The WiFi open set gesture recognition method based on symbiotic metric learning and uncertainty awareness according to claim 1, characterized in that, The nearest neighbor distance score is expressed as follows: in, The score is the nearest neighbor distance score; The minimum Euclidean distance between the sample to be detected and the sample among the samples with the same candidate predicted label as the sample to be detected in the set of K nearest neighbors; The minimum Euclidean distance between the sample to be detected and the sample among the samples with different labels from the candidate predicted labels of the sample to be detected in the set of K nearest neighbors; To prevent the minimum value where the denominator is 0; This is the original feature vector of the sample to be detected. This is the original feature vector of the sample in the set of K nearest neighbors of the sample to be detected; express Corresponding tags; These are the candidate predicted labels for the samples to be detected.

4. The WiFi open set gesture recognition method based on symbiotic metric learning and uncertainty awareness according to any one of claims 1 to 3, characterized in that, The multi-branch neural network also includes a physical reconstruction branch and a classification branch. The physical reconstruction branch is used to reconstruct the Doppler frequency shift spectrum based on the feature maps extracted by the backbone network, and the classification branch is used to directly predict the gesture label based on the feature maps extracted by the backbone network. During training, the total loss function is calculated based on the results of the physical reconstruction branch, the classification branch, and the feature embedding branch to update the network parameters of the multi-branch neural network. The total loss function is expressed as: in, This is the total loss function; , These are the weight coefficients that evolve with the training round t. To monitor losses; This indicates the loss in neighborhood component analysis; This represents the loss of label perception.

5. The WiFi open set gesture recognition method based on symbiotic metric learning and uncertainty awareness according to claim 4, characterized in that, Set the number of warm-up training rounds In the warm-up training rounds Inside , The value is set to 0 after exceeding the number of warm-up training rounds. and according to The proportion increases linearly. This represents the maximum number of training iterations.

6. The WiFi open set gesture recognition method based on symbiotic metric learning and uncertainty awareness according to claim 4, characterized in that, The supervision loss consists of the sum of the classification loss of the classification branch, the reconstruction loss of the physical reconstruction branch, and the consistency loss between the first-view samples and the second-view samples, including: The classification loss of the classification branch is expressed as: in, Predicting distribution for classification With real labels Calculate the classification cross-entropy loss between them; C represents the total number of samples in the current training batch; C represents the total number of predefined known gesture categories. This indicates that the i-th sample belongs to the c-th category; This represents the predicted probability value of the i-th sample belonging to the c-th class in the classification branch output; The reconstruction loss of the physical reconstruction branch is expressed as: in, The mean square error loss represents the difference between the true Doppler frequency shift spectrum of the CSI image sample and the Doppler frequency shift spectrum reconstructed by the physical reconstruction branch; T represents the total length of the time dimension of the gesture signal in the Doppler frequency spectrum space. This represents the true Doppler spectral features, which are pre-extracted from the original channel state information using short-time Fourier transform and used as a supervision label; This represents the reconstructed Doppler spectrum features generated by deconvolution of the spatial feature map in the physical reconstruction branch of a multi-branch neural network. The consistency loss between the first-view sample and the second-view sample is expressed as: in, The consistency loss between the first view sample and the second view sample; To smooth the temperature coefficient, For the Softmax function; This indicates the calculation of the KL divergence; This represents the classification prediction vector output by the backbone network and classification branches of the first view sample; This represents the classification prediction vector output by the backbone network and classification branches of the second view sample.

7. The WiFi open set gesture recognition method based on symbiotic metric learning and uncertainty awareness according to claim 6, characterized in that, The physical reconstruction branch consists of a decoder composed of multiple deconvolutional layers, batch normalization layers, and ReLU activation function layers. During the training phase, random noise following a Gaussian distribution is injected into the original feature vector obtained from the feature embedding branch to generate robust feature vectors. , , This is the original feature vector of the image. For a random noise vector that follows a standard normal distribution, This is the preset noise intensity coefficient; The classification branch uses a fully connected mapping to project features onto the category space and outputs a predicted probability distribution.

8. The WiFi open set gesture recognition method based on symbiotic metric learning and uncertainty awareness according to claim 4, characterized in that, Neighborhood component analysis loss Represented as: Where N is the total number of samples in the current training batch; Let k be the set of the k nearest neighbors of the i-th training sample in the feature space. Represents a set The label of the j-th sample. The label represents the i-th training sample; Let be the confidence weight of the j-th nearest neighbor sample, denoted as: C represents the total number of predefined known gesture categories. Let be the scalar value of the probability that the j-th nearest neighbor sample belongs to the c-th known gesture category; Let cosine similarity be used to represent the original feature vector of the i-th training sample and its j-th nearest neighbor sample. This refers to the temperature parameter.

9. The WiFi open set gesture recognition method based on symbiotic metric learning and uncertainty perception according to claim 4, characterized in that, Label-perceived loss Represented as: in, Let be the set of samples with the same label as the i-th training sample in the nearest neighbor set; Let be the set of samples whose labels are different from those of the i-th training sample in the nearest neighbor set; For set The number of elements in the middle; Let be the confidence level of the i-th training sample. Let be the confidence level of the j-th neighbor sample; As a repulsion factor, Let be the confidence weight of the i-th training sample, denoted as C represents the total number of predefined known gesture categories. Let be the scalar value of the probability that the j-th training sample belongs to the c-th known gesture category; It is a linear rectifier function.

10. The WiFi open set gesture recognition method based on symbiotic metric learning and uncertainty awareness according to any one of claims 1-3 and 5-9, characterized in that, CSI image sample acquisition includes: Iterate through all antennas m at the receiver and calculate the ratio of the mean to the variance of the CSI amplitude for each antenna, defining it as the stability score. ; Select The largest antenna as the reference antenna Select The smallest antenna as the target antenna ; Perform complex division to calculate CSI entropy And extract phase information to generate a phase image. ,Right now: in, The phase information of the denoised CSI quotient sequence can be approximately mapped to the gesture-induced phase caused by human motion. ; To determine the complex CSI sequence of the target antenna using an antenna selection strategy; and This represents the complex CSI sequence used to determine the reference antenna via an antenna selection strategy. This indicates the random phase deviation contained in the target antenna; This indicates the random phase deviation contained in the reference antenna; This represents the complex amplitude coefficient remaining after the signal undergoes quotient operation; This indicates taking the modulus of the complex number A; j is the imaginary unit; This indicates that a phase-taking operation is performed on the CSI quotient sequence to extract the gesture-induced phase value; This refers to the `unwrap` function in MATLAB.