Cross-domain user identity identification method based on Wi-Fi channel state information
By extracting environment-independent features in Wi-Fi CSI and combining personnel classifiers and data reconstruction decoders, the problems of cross-domain user identity recognition and unknown user detection are solved, and the ability to accurately identify and detect unknown users in different environments is realized.
Patent Information
- Application Number
- CN202510185175.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-30
AI Technical Summary
The existing Wi-Fi CSI-based identity recognition method is difficult to achieve cross-domain user identity recognition, and there is a risk of misjudgment when detecting unknown users. How to reasonably set the threshold is a difficult problem.
By extracting environment-independent features in CSI, combining personnel classifiers and data reconstruction decoders, a cross-domain user identity recognition method is designed. The method includes data preparation, a first training phase and a second training phase, utilizing an adaptive multi-level gradient inversion layer and an autoencoder, mitigating the environmental dependence of the model, and detecting unknown users by reconstructing losses.
It realizes the identification of known user identities in different environments, improves the detection ability of unknown users, enhances the robustness of the system, and can accurately distinguish known users from unknown users.
Smart Images

Figure CN120074700A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to wireless sensing technology, and particularly to cross-domain user identity recognition technology based on Wi-Fi channel state information. Background Art
[0002] Traditional gait recognition methods usually rely on camera devices, wearable devices or sensors to obtain raw data. However, these methods have certain limitations, such as privacy issues, dependence on devices, and environmental conditions. In recent years, gait recognition technology based on Wi-Fi signals has gradually attracted wide attention. The radio waves of Wi-Fi signals can not only penetrate clothing, but also be reflected by the human body, and gait recognition can be achieved by analyzing Wi-Fi signals. Compared with the above methods, the gait recognition technology based on Wi-Fi signals shows many significant advantages: it can better protect privacy, is not affected by light conditions and line-of-sight occlusion, and relies on the existing Wi-Fi infrastructure, with low cost and strong practicability.
[0003] As a unique biometric feature, human gait has the characteristics of being difficult to disguise and imitate. Compared with other biometric identification technologies, user identity recognition based on Wi-Fi channel state information (CSI) through gait features does not require physical contact, can monitor at a long distance, and is suitable for identity recognition of a large number of people, showing broad application prospects in fields such as security monitoring and identity verification.
[0004] However, the identity recognition method based on Wi-Fi CSI still has some challenges. First, there are difficulties in achieving cross-scene identity recognition. Existing Wi-Fi signal-based identity recognition mainly relies on extracting identity features from CSI signals. However, CSI signals are extremely sensitive to environmental changes. When the environment changes, it may lead to a significant shift in data distribution. If the differences between the target domain and the source domain are ignored, the recognition performance of the model may decrease significantly. Specifically, the "domain" in this application can be understood as a set of data distributions or a specific environmental background, usually referring to the data generation environment or scene.
[0005] In addition, although existing models show high accuracy in the identity recognition of known users, there are still obvious deficiencies in detecting unknown users. A perfect human body recognition system should have the ability to recognize potential unknown users. Especially in actual application scenarios, the system needs to deal with more unknown users rather than just known users.
[0006] Currently, most methods improve the detection rate of unknown users by setting relatively strict thresholds. However, detection systems based on thresholds usually face a key problem: how to balance the threshold setting. If the threshold is set too low, the system may increase the risk of misclassifying known users as unknown users. On the contrary, if the threshold is set too high, although the false alarm rate will decrease, the system's detection ability for unknown users will also be correspondingly weakened. Therefore, in the practical application of identity recognition systems, how to reasonably set the threshold is an urgent problem to be solved. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to propose a technical solution for cross-domain user identity recognition by extracting environment-independent features in CSI and combining a personnel classifier and a data reconstruction decoder.
[0008] The technical solution adopted by the present invention to solve the above technical problem is a cross-domain user identity recognition method based on Wi-Fi CSI, including the following steps:
[0009] Data preparation stage: Obtain sample data of gait CSI data of legal personnel in each known environment; Build an identification model including an encoder, a decoder, and a classifier; The classifier includes a legal personnel classifier and a known environment classifier;
[0010] The first training stage: Input the sample data into the encoder. The encoder outputs user features after feature extraction and compression, and then input the user features into the legal personnel classifier and the known environment classifier respectively. The legal personnel classifier outputs the predicted user identity category and calculates the identity recognition loss, and the known environment classifier outputs the predicted environment category and calculates the environment recognition loss. Use the total loss obtained by subtracting the weighted environment recognition loss from the identity recognition loss to constrain the first training process, and take the user features output by the encoder that are independent of the environment as the goal of the first training stage;
[0011] The second training stage: Freeze the encoder obtained in the first training stage, and form an autoencoder with this encoder and the decoder; Input the sample data into the encoder in the autoencoder. The encoder outputs user features to the decoder, and the decoder outputs the reconstructed CSI data and calculates the reconstruction loss. Take accurately restoring the input sample data as the goal of the second training stage;
[0012] Detection stage:
[0013] Input the sample data into the trained autoencoder and calculate the reconstruction loss of each sample. Determine the detection threshold by referring to the reconstruction losses of all samples;
[0014] Input the user data to be tested into the trained autoencoder and calculate the reconstruction loss of the user data to be tested. Determine whether the reconstruction loss of the user data to be tested is greater than the detection threshold. If so, output the user identity recognition result as an unknown user; otherwise, input the user data to be tested into the encoder to obtain the user features, and then input the user features into the legitimate personnel classifier, and output the user identity category predicted by the legitimate personnel classifier as the user identity recognition result.
[0015] The beneficial effects of the present invention are as follows:
[0016] (1) A domain-invariant feature extraction method based on an adaptive multi-level gradient reversal layer is proposed to reduce the environmental dependence of the model and achieve the recognition of known user identities in different environments.
[0017] (2) By designing a two-stage training mechanism and combining the detection method of the autoencoder and the reconstruction loss, the efficient detection of unknown users is achieved, and the known users and unknown users can be accurately distinguished in different environments, thereby improving the robustness of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a cross-domain user identity recognition system diagram applied to the method of the present invention;
[0019] Figure 2 It is a schematic diagram of the principle of a cross-domain user identity recognition model based on Wi-Fi CSI;
[0020] Figure 3 It is a detailed schematic diagram of a cross-domain user identity recognition model based on Wi-Fi CSI;
[0021] Figure 4 It is a schematic diagram of a residual network;
[0022] Figure 5 It is a schematic diagram of a decoder module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] In order to illustrate the present invention in detail, the technical solutions of the present invention will be further described below in conjunction with the drawings in the embodiments of the present invention.
[0024] As Figure 1 shown, it is a complete process of a CSI-based user identity recognition system. The overall process can be divided into four main parts: data preparation, first-stage training, second-stage training, and detection stage.
[0025] S1. Data preparation stage:
[0026] 1-1 The CSI data preprocessing process includes the following steps:
[0027] 1-1-1 Discrete wavelet transform denoising
[0028] In scenarios of different environments, namely "domains", such as Scenario 1, Scenario 2, and Scenario 3, CSI data is collected. Since the collected CSI data contains a large amount of noise, these interference factors may have a significant impact on the final result of gait recognition. Therefore, it is crucial to denoise the signal. The present invention uses discrete wavelet transform (DWT) to denoise the extracted CSI. The data after denoising where N 1 represents the number of samples, T represents the time dimension, R represents the antenna dimension, and F represents the subcarrier dimension.
[0029] Specifically, the present invention uses wavelet decomposition based on symmetric wavelet functions to perform multi-scale decomposition on the signal, and combines the soft threshold method to suppress noise. The number of layers of wavelet decomposition is adaptively selected according to the characteristics of the signal to ensure effective noise removal while maintaining the local characteristics of the signal. Finally, the coefficients after threshold processing are reconstructed through inverse wavelet transform to generate the denoised CSI signal. DWT is implemented through a series of filters: First, the samples pass through a low-pass filter and a high-pass filter with impulse responses g and h:
[0030]
[0031] where n represents the sampling sequence number, x is the input signal, and y low is the output signal of the low-pass filter, and y high is the output signal of the high-pass filter, g is the impulse response function of the low-pass filter, h is the impulse response function of the high-pass filter, and k is a variable of the order of the filter.
[0032] Repeat this wavelet decomposition to further improve the frequency resolution and approximation coefficients, decompose with high-pass and low-pass filters, and then perform downsampling. In this way, a filter bank binary tree can be formed, and each level of the signal is decomposed into low frequency and high frequency. After the signal undergoes DWT, the original signal can be reconstructed through inverse transformation:
[0033]
[0034] where X(z), G(z), and H(z) are the Z-transforms of the input signal x, the impulse response function g of the low-pass filter, and the impulse response function h of the high-pass filter respectively, z is a complex variable, and G 1 (z) is the Z-transform of the response function of the reconstructed low-pass filter respectively, and H 1 (z) is the Z-transform of the response function of the reconstructed high-pass filter.
[0035] 1-1-2 Extract gait data
[0036] Although the CSI signal processed by denoising has been improved, it still cannot be directly used for gait recognition. Because the collected CSI segments contain data of the user's entire walking process. Therefore, in this embodiment, a sliding window segmentation algorithm is adopted to intercept gait data.
[0037] First, calculate the variance in the time dimension for the data of each antenna channel. Then, use a sliding window to calculate the local variance of the variance signal to capture the local change characteristics of the signal. Among them, the window size is 2000 data packets, and the window slides along the time series data stream, moving a fixed step length each time, generating a series of overlapping or non-overlapping sub-windows. Next, set a dynamic threshold based on the mean and standard deviation of the local variance, and find the regions where the local variance is higher than the threshold. We consider that this window contains gait data, record the indexes of the high variance regions, and then determine the start and end of the walking activity. To ensure data consistency, finally, the gait data is expanded or trimmed to a fixed 6000 data packets.
[0038] 1-2 Data augmentation steps:
[0039] To improve the robustness and generalization ability of the subsequent model, the present invention introduces data augmentation technology. By performing various forms of perturbation or transformation on the CSI signal during the training process, data augmentation can simulate different wireless environments, thereby generating more diverse training data. This not only helps the model learn more channel characteristics but also avoids the model overfitting to certain specific channel conditions. The CSI data after data augmentation is where N 2 represents the number of samples.
[0040] Specifically, the present invention adopts the following enhancement methods:
[0041] Gaussian white noise: By adding Gaussian white noise to the CSI signal, it simulates the noise interference in the actual wireless environment and enhances the model's robustness to noise.
[0042] Multipath fading simulation: The multipath effect is a common phenomenon in wireless signal propagation, and the signal reaches the receiving end through different paths. We simulate the multipath fading phenomenon by applying different delays and scalings to the signal, enabling the model to better adapt to the multipath environment.
[0043] Frequency selective fading: Frequency selective fading in the wireless channel will cause signals of different frequency components to be affected to different degrees. We simulate this phenomenon by applying variation coefficients to different frequency components to enhance the model's adaptability to frequency selective fading.
[0044] Antenna Sequence Transformation: In wireless communication, the arrangement and selection of antennas can affect the received signal. We randomly shuffle the antenna sequence to simulate the impact of different antenna arrangements on the signal, enhancing the robustness of the model in antenna selection.
[0045] Subcarrier Masking: In some cases, some subcarriers may be interfered with or become invalid. We randomly mask some subcarriers to simulate this situation, enhancing the model's tolerance to subcarrier failures.
[0046] Through the above data augmentation techniques, we generate more diverse CSI signal data, helping the recognition model maintain high performance in different wireless environments.
[0047] As Figure 2 shown, the recognition model uses an encoder, a decoder, and a classifier during training. The encoder is used to receive the input CSI data, extract the features of the user, and output after data compression. The classifier includes a legitimate personnel classifier and a known environment classifier. The legitimate personnel classifier is used to receive the feature data output by the encoder and output the predicted category of legitimate personnel. The known environment classifier is used to receive the feature data output by the encoder and output the predicted environment category. The decoder is used to receive the feature data output by the encoder and output the reconstructed CSI data.
[0048] During detection, the recognition model uses an autoencoder composed of a trained encoder and decoder to identify unknown users, and uses a legitimate personnel classification model composed of an encoder and a legitimate personnel classifier to identify the identity of legitimate users.
[0049] S2. The First Phase of the Training Phase of the Recognition Model:
[0050] This phase is used to train the recognition model for cross-domain recognition of known user identities.
[0051] As Figure 3 shown, during the training process of the first phase, the recognition model will use an encoder, a classifier, and an adaptive multi-level gradient reversal layer, aiming to achieve the identity recognition function by processing CSI data in different environments.
[0052] Input the CSI data processed in step S1 into the encoder. The encoder includes two linear layers, a residual network module, and a multi-head attention mechanism Transformer encoding module. Specifically, to reduce data complexity, first, perform dimensionality reduction processing through two linear layers to obtain the dimensionality-reduced data as the low-level feature representation. Among them, N represents the number of samples after dimensionality reduction, R represents the antenna dimension, f represents the subcarrier dimension after dimensionality reduction, t represents the time dimension after dimensionality reduction, and F represents the subcarrier dimension. Next, take F lDivided into R tensors according to the antenna dimension R. In the embodiment, 3 antennas are used, divided into 3 tensors, and respectively input into the residual network module.
[0053] As Figure 4 shown, the residual network module includes a plurality of mainly composed of a plurality of residual networks ResNet. Each residual network ResNet is composed of a plurality of convolutional layers Conv and activation function Mish. Each residual network extracts features from the input, and the i-th residual network block outputs features where k and g represent the intermediate feature dimensions. Then, the features extracted by the R residual modules are aggregated to form an intermediate-level feature representation F m . Subsequently, it is input into the Transformer encoding module.
[0054] The Transformer encoding module captures the spatio-temporal dependence of the CSI data through the multi-head attention mechanism and extracts a higher-level feature representation F h . The output F of the Transformer encoding module h is used as the output feature of the encoder where d represents the feature dimension.
[0055] The classifier module consists of two independent classifiers: one for classifying known environments and one for identifying known personnel identities.
[0056] Input the low-level feature representation F l , the intermediate-level feature representation F m and the high-level feature representation F h into the known environment classifier, so that the environment classifier can predict the environment category to which the data belongs based on feature information at different levels such as scenario 1, scenario 2,..., scenario m, where i ∈ {l, m, h}. represents the prediction probability after the i-th level is input into the environment classifier. To measure the predicted environment category and the true environment label The difference between them, the environment label is the domain label, and the loss of each level is calculated by the cross-entropy function as follows:
[0057]
[0058] To ensure that the generalization ability between different levels is consistent, a generalization consistency objective is introduced to measure the generalization gap between feature representations at different levels. The generalization ability G of each level i iMeasured by the regularization loss, the regularization term helps control the complexity of the model and prevent overfitting. If the feature representation at a certain level can achieve good performance with a small regularization loss, it indicates that this level has strong generalization ability. Thus, the generalization consistency objective function can be defined as:
[0059]
[0060] Combining the weighted aggregation of the environmental recognition loss at each level and the generalization consistency objective, the global objective, i.e., the total environmental recognition loss, is obtained. In this global objective, both the loss weighted sum of hierarchical features and the fairness between levels are considered. Let the loss weighting coefficient for each level be θ i , i ∈ {l, m, h}, θ i will be adaptively adjusted according to the generalization ability G i of each level. If the generalization ability G i of a certain level is poor and the variance is large, the weight θ i of this level will be automatically adjusted to increase its influence. If the generalization ability G i of a certain level is good and the variance is small, the weight θ i of this level will be decreased to prevent this level from contributing too much to the total loss, thus avoiding the model relying too much on the features of this level. The environmental recognition loss can be expressed as:
[0061]
[0062] where is a hyperparameter used to balance the weights between the loss and the generalization consistency objective. This global objective will guide the model in the optimization process to not only focus on minimizing the loss but also ensure the relative balance of the contributions of different levels, avoiding a certain level from overly dominating the model's learning.
[0063] The features F extracted by the encoder are input into the legitimate personnel classifier, which is used to predict the identity category of the user The user's identity recognition is completed by classifying the user's identity type. Legitimate personnel are those who have collected CSI data and are known. Similarly, the cross-entropy loss function is used to calculate the loss of identity recognition
[0064]
[0065] where y env and y id represent the true environmental label and personnel label of the data respectively, c is the category serial number, m represents the total number of environmental categories, and n represents the number of legitimate personnel categories.
[0066] An adaptive multi-level gradient reversal layer is introduced between the encoder and the known environment classifier to minimize the prediction accuracy of the domain labels. The trade-off between different source domains is controlled by adaptively adjusting the weights of the multi-level gradient reversal layer. When the weight is close to 0, the recognition model pays more attention to retaining the features of the source domain, and when the weight is close to 1, it pays more attention to reducing the differences between domains. This stage enables the model to extract identity features independent of the environment. That is, under different environmental conditions, it can still accurately extract features related to the personnel identity. The final total loss function is defined as:
[0067]
[0068] where λ is a hyperparameter.
[0069] The encoder that has completed the first-stage training extracts user identity features independent of the environment from the input cross-domain CSI data.
[0070] S3. The second training stage of the recognition model:
[0071] This stage is used to train the reconstruction ability of the recognition model for CSI data.
[0072] Before training, the parameters of the encoder trained in the first stage are frozen, and the encoder and decoder are combined to construct a complete autoencoder to complete the reconstruction of the data, so as to ensure that the model can accurately restore the features of the input data. In this stage, only the denoised CSI data X is used for training, and the encoder encoder is used to extract the features F of X:
[0073] F = encoder(X)
[0074] As Figure 5 shown, the decoder includes 3 linear layers Linear, multiple convolutional layers Conv and activation function Mish. The decoder is used to decode the extracted CSI features F to obtain the reconstructed data X rec :
[0075] X rec = decoder(F)
[0076] Finally, the reconstruction loss is calculated for training:
[0077]
[0078] where, ‖‖ 2 is the L2 norm.
[0079] The training purpose of the second stage is to generate an autoencoder that can calculate a more appropriate detection threshold σ.
[0080] S4. Detection stage
[0081] Before the detection starts, use the training data Input the autoencoder obtained from the second-stage training again to obtain the reconstruction loss of each sample And select the 99th percentile among them as the detection threshold σ.
[0082] For any test data Input the test data x into the autoencoder, and use the above method to obtain its reconstruction loss l rec , first compare it with the threshold. If l rec >σ, it is considered that the data is an unknown user; otherwise, it is considered a legitimate user, and the test data x is input into the encoder to obtain the feature encoder(x), and the feature encoder(x) is then sent to the legitimate personnel classifier for legitimate personnel classification.
[0083] The above are embodiments for explaining the technical idea of the present invention. The protection scope of the present invention cannot be limited thereby. Based on the professional ability and knowledge of ordinary technicians in the art, on the basis of the above description, other different forms of changes or modifications can be made without departing from the purpose of the present invention. Any modification, replacement, and improvement made on the basis of the technical solution proposed by the present invention shall fall within the protection scope of the claims of the present invention.
Claims
1. A cross-domain user identity identification method based on WiFi channel status information, characterized in that: The following steps are involved: Data preparation stage: obtain sample data of gait CSI data of legal persons in various known environments; build a recognition model including encoder, decoder and classifier; the classifier includes legal person classifier and known environment classifier; The first training stage: the sample data is input into the encoder, and the encoder outputs user features after feature extraction and compression. The user features are then input into the legal person classifier and the known environment classifier respectively. The legal person classifier outputs the predicted user identity category and calculates the identity recognition loss. The known environment classifier outputs the predicted environment category and calculates the environment recognition loss. The total loss obtained by subtracting the weighted environment recognition loss from the identity recognition loss is used to constrain the first training process. The goal of the first training stage is to output user features that are independent of the environment from the encoder. Second training stage: freeze the encoder obtained in the first training stage, and use the encoder and decoder to form an autoencoder; input sample data into the encoder in the autoencoder, the encoder outputs user features to the decoder, the decoder outputs reconstructed CSI data and calculates the reconstruction loss, with the goal of accurately restoring the input sample data as the second training stage; Detection phase: Input the sample data into the trained autoencoder and calculate the reconstruction loss of each sample. Determine the detection threshold with reference to the reconstruction loss of all samples. The user data to be tested is input into the trained automatic encoder and the reconstruction loss of the user data to be tested is calculated to determine whether the reconstruction loss of the user data to be tested is greater than the detection threshold. If so, the user identity recognition result is output as an unknown user; otherwise, the user data to be tested is input into the encoder to obtain user features, and then the user features are input into the legal person classifier, and the user identity category predicted by the legal person classifier is output as the user identity recognition result.
2. The method according to claim 1, characterized in that: In the data preparation stage, discrete wavelet transform DWT is used for denoising.
3. The method according to claim 1 or 2, characterized in that: In the data preparation stage, the gait CSI data of legal personnel in various known environments are collected and denoised to obtain the original sample data, and then the original sample data is enhanced to obtain the expanded sample data.
4. The method according to claim 3, characterized in that: In the data preparation stage, data enhancement methods include adding Gaussian white noise, multipath fading simulation, frequency selective fading, antenna sequence transformation and random subcarrier masking.
5. The method according to claim 4, characterized in that: In the first training phase, the augmented sample data is input into the encoder; In the second training stage, the original sample data is input into the encoder in the autoencoder, the encoder outputs the user features to the decoder, the decoder outputs the reconstructed CSI data and calculates the reconstruction loss, with the goal of accurately restoring the input original sample data; In the detection phase, the original sample data is input into the trained autoencoder and the reconstruction loss of each sample is calculated.
6. The method according to claim 1, characterized in that: The method of determining the detection threshold with reference to the reconstruction loss of all samples is: selecting the 99th percentile in the reconstruction loss of all samples as the detection threshold.
7. The method according to claim 1, characterized in that: In the first training stage, the cross entropy loss function is used to calculate both the identity recognition loss and the environment recognition loss.
8. The method according to claim 1, characterized in that: In the first training stage, an adaptive multi-level gradient reversal layer is introduced between the encoder and the known environment classifier to help the encoder extract features with strong cross-domain adaptability to minimize the prediction accuracy of domain labels.
9. The method according to claim 1 or 8, characterized in that: In the first training stage, the adaptive multi-level gradient reversal layer dynamically adjusts the contribution of different layer features to the total loss by learning weights, and introduces a generalization consistency objective to measure the generalization gap between different layer features.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to claim 1 are implemented.
Citation Information
Patent Citations
Method and device for realizing multi-user identity recognition based on WiFi signal detection gait
CN110621038A
Field situation future guiding technology based on mimicry adversarial learning mechanism
CN110807291A
Mixed reality open set human body posture recognition method based on deep learning
CN113705507A
Method for training deep neural network and apparatus
US20210012198A1