A Cross-Environment Human Activity Recognition Method Based on Dual Adversarial Networks

Through cross-adversarial training and domain adversarial training of dual adversarial networks, the performance degradation of WiFi sensing technology when applied across domains in different environments is solved, and a better cross-environmental human activity recognition effect is achieved.

CN116884090BActive Publication Date: 2025-07-22HEBEI UNIV OF ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310885458.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-19
Publication Date
2025-07-22
Estimated Expiration
2043-07-19

AI Technical Summary

Technical Problem

When existing WiFi sensing technology is applied across domains in different environments, the model performance is severely degraded and human activities cannot be effectively identified. This is mainly due to the limitation of domain-invariant feature extraction ability due to the difference in feature distribution of source domain and target domain.

Method used

A cross-environmental human activity recognition method based on dual adversarial network is adopted to extract features through convolutional neural networks, and virtual samples are generated by combining the generation adversarial network and self-attention. Cross-adversarial training and domain adversarial training are carried out to reduce the impact of environmental changes and improve the generalization ability of the model.

Benefits of technology

It effectively improves the cross-environmental human activity recognition performance of the model in different environments, reduces the impact of environmental characteristics, and improves the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116884090B_ABST
    Figure CN116884090B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross - environment human activity recognition method based on a dual adversarial network, belonging to the technical field of human activity recognition. It includes: S1, denoising processing; S2, when the activity distributions of the target domain are far apart, combining self - attention mechanism and generative adversarial network to generate virtual samples; S3, inputting the corresponding source - domain features into the classifier to calculate the cross - entropy loss; S4, inputting the features in another source domain into the classifier of the current source domain for cross - adversarial training; S5, combining domain - adversarial training in the cross - adversarial training of the classifier, and calculating the adversarial loss in the same batch of training; S6, using self - prediction learning technology to generate pseudo - labels; S7, calculating the self - prediction loss; S8, minimizing the final loss; S9, for the target domain whose accuracy rate does not reach the preset standard, fine - tuning the model with a small number of labeled samples. The present invention effectively improves the generalization ability of the model for different domains by combining cross - adversarial training between source domains and domain - adversarial training between the source domain and the target domain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human activity recognition, and in particular to a cross-environment human activity recognition method based on a dual adversarial network. Background Art

[0002] Human activity recognition is a core component in the field of artificial intelligence such as smart home, health monitoring, and virtual reality. Activity recognition attempts to utilize various sensing methods ranging from vision, wearable, to acoustics. However, they have inherent drawbacks such as privacy leakage and limited sensing range. Due to the characteristics of WiFi such as privacy protection, device-free, and ubiquitous, it has become a powerful sensing technology. With the popularization of wireless networks and the development of wireless communication technologies (Multiple-Input-Multiple-Output, MIMO) and Orthogonal Frequency-Division Multiplexing (OFDM), WiFi devices can obtain channel state information (CSI) from each transmitting antenna to each receiving antenna along multiple paths at each carrier frequency. WiFi sensing reveals activity information by capturing the change pattern of CSI in the propagation path. When a user is in a static environment, multiple paths can be divided into static paths and dynamic paths. Static paths are signals reflected from static objects (including walls, desks, chairs, etc.), and dynamic paths are signals reflected from active people. Therefore, in different environments, the amplitude and phase in the CSI of the same activity are different, resulting in a significant performance degradation when the model trained in the source domain is applied to the new domain. In practice, it is impossible for us to collect an infinite number of activity data in different environments, which greatly limits the popularization of WiFi sensing technology.

[0003] The Domain Adaptation Network (DAN) is a representative network model for learning the mapping relationship between different domains. The most widely used DAN in cross-domain WiFi sensing is the Domain Adversarial Neural Networks (DANN). It has three main components: a feature extractor, an activity recognizer, and a domain discriminator. The feature extractor learns the features of labeled samples in the training domain and unlabeled samples in the test domain; based on the learned features, the activity recognizer uses two fully connected networks and then a softmax layer to obtain the predictions of all input samples. During training, the feature extractor tries to learn domain-invariant features to deceive the domain discriminator while improving the performance of the activity recognizer. However, when applying the domain adaptation network to cross-domain activity sensing tasks, the ability of the model to extract domain-invariant features is limited by the differences between the source domain and the target domain. Only when the feature distribution distance between different domains is small can the model learn effective activity features. Summary of the Invention

[0004] The purpose of the present invention is to provide a cross-environment human activity recognition method based on a dual adversarial network, which uses labeled data from multiple source domains and unlabeled data from the target domain as inputs, and uses a Convolutional Neural Network (CNN) structure as a feature encoder to extract features; through cross-adversarial training between two groups of activity classifiers and domain adversarial training between the classifier and the domain discriminator, the influence brought by environmental changes is removed, and the cross-environment adaptive activity recognition performance of the human body is improved.

[0005] To achieve the above purpose, the present invention provides a cross-environment human activity recognition method based on a dual adversarial network, including the following steps:

[0006] S1. For all CSI data collected in the source domain and the target domain, interpolate the subcarriers into an arithmetic time series data according to the timestamp information, and use a filter to perform signal denoising to obtain clean subcarriers as the input of the neural network model;

[0007] S2. For the target domain where the same activity sample distribution is difficult to converge, use the technology combining the generative adversarial network and self-attention to generate virtual samples. Specifically, given the target domain samples and the source domain samples , minimize the following adversarial loss function:

[0008]

[0009]

[0010] where represents the discriminator of the source domain, The discriminator representing the target domain The generator representing the source domain The generator representing the target domain The source domain dataset The target domain dataset;

[0011] S3. Given a labeled source domain including a dataset of samples and the corresponding label set , each sample is associated with a label in it; given an unlabeled target domain dataset consisting of samples; divide the dataset into two source domain datasets and , and design corresponding classifiers and for them; calculate the cross-entropy loss by inputting the corresponding source domain features into the classifier: and ; where

[0012]

[0013] represents the number of samples in the corresponding source domain, represents the target domain encoder, represents the transpose of the label vector;

[0014] S4. Fix the trained and , input the features extracted from the other source domain into the classifier of the current source domain to calculate the cross-entropy loss, and perform cross-adversarial training;

[0015] ;

[0016] S5. Combine domain adversarial training in the classifier cross-adversarial training, send the features extracted from the source domain data and the features extracted from the target domain data to the source domain environmental discriminator , and perform domain adaptive adversarial training by learning the mapping features of the target domain. Calculate the adversarial loss in the same batch of training. The adversarial loss function is expressed as:

[0017]

[0018] ;

[0019] S6. Mine the discriminative information in the unlabeled target domain data using the self-prediction learning method. Utilize the similarity between the source domain samples and the batch of target domain samples to establish a matching criterion, select samples in the target domain to assign pseudo-labels. When the recognition results of two classifiers and are the same and the sum of the output results is greater than the set threshold, assign a pseudo-label to the target sample. This process is expressed as:

[0020]

[0021]

[0022] represents the recognition result of classifier , represents the recognition result of classifier , m represents the subscript of the maximum value among them, represents soft label prediction, represents the average similarity probability;

[0023] S7. Calculate the self-prediction loss. The expression of the self-prediction loss is:

[0024] ;

[0025] S8. Optimize the target domain encoder by minimizing the final loss ;

[0026] S9. For the case where the recognition accuracy in the target domain is low, use the fine-tuning technique in transfer learning to make the model trained in the source domain adapt to the activity data in the new scenario. The structure of Softmax depends on the number of action categories in the CSI dataset, and the number of actions in each dataset is not the same. Therefore, when migrating to a new recognition scenario, use the target domain data to fine-tune the given source model, reload the parameters of the pre-trained model, and retrain the weights of the fully connected layer and the Softmax layer in the classifier with a higher learning rate.

[0027] Preferably, in S1, the filter signal denoising is specifically as follows: Use a hampel filter to remove outliers far from the data stream trend; Remove high-frequency noise through a sixth-order Butterworth low-pass filter; Use Daubechies8 wavelets to suppress environmental noise within the frequency band.

[0028] Preferably, in S2, the self-attention mechanism is used to assign more weights to the activity features. The specific approach is as follows in the formula:

[0029]

[0030]

[0031] wherein represents a fully connected layer, represents the output of the first 4 convolutional neural network layers of the generator, represents the weight of the corresponding layer; the finally generated virtual samples are represented as follows:

[0032]

[0033] wherein is the feature map generated by the second convolutional layer.

[0034] In addition, when learning the mapping of the generator and the discriminator, it is required that and . Therefore, the following cycle consistency loss is adopted to make the network more stable. The cycle consistency loss is defined as:

[0035] .

[0036] Preferably, in the S4, the discriminant information of the unlabeled target domain samples is used to reduce the consistency loss:

[0037]

[0038] wherein represents the L1-norm, is the number of target domain samples in a batch, is the classifier and output dimension.

[0039] Preferably, in the S4, when calculating the cross-entropy loss function, a weighted mixture of one-hot labels and a uniform distribution is used, and the specific formula is

[0040]

[0041] wherein represents the domain label; is a smoothing parameter, set to 0.3.

[0042] Preferably, in the S4, a triplet loss function is defined, P activities are randomly selected, and K pieces of data are randomly selected for each activity. The triplet loss is expressed as;

[0043]

[0044] wherein represents the An anchor of a sample, Indicates the positive sample of the th sample, Indicates the negative sample of the th sample; Is the margin parameter of the triplet loss. If it is set too small, it will be difficult to distinguish between positive and negative samples.

[0045] Preferably, in the step S8, the final loss function is

[0046]

[0047] Wherein, , , , respectively represent the weights of the consistency loss, adversarial loss, triplet loss, and self-prediction loss functions.

[0048] Preferably, a cross-entropy loss function is defined in the step S9:

[0049]

[0050] Wherein represents a small amount of labeled target domain data.

[0051] The advantages and positive effects of the cross-environment human activity recognition method based on a dual adversarial network according to the present invention are as follows: By combining cross-adversarial training between source domains and domain adversarial training between the source domain and the target domain, environmental features are more effectively removed, and the generalization ability of the model to different domains is improved.

[0052] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a flowchart of an embodiment of a cross-environment human activity recognition method based on a dual adversarial network according to the present invention;

[0054] Figure 2 is a model framework diagram of an embodiment of a cross-environment human activity recognition method based on a dual adversarial network according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0055] The technical solutions of the present invention will be further described below through the drawings and embodiments.

[0056] Embodiment

[0057] As Figure 1 shown, a cross-environment human activity recognition method based on a dual adversarial network includes the following steps:

[0058] S1. For all the CSI data collected in the source domain and the target domain, interpolate the subcarriers into an arithmetic time series data according to the timestamp information, and downsample the packet length to 250.

[0059] After that, use the hampel filter to remove the outliers far from the data stream trend. Considering that the channel frequency change caused by the hardware itself is relatively fast and is in the high-frequency part of the spectrogram, remove the high-frequency noise through a sixth-order Butterworth low-pass filter. Use the discrete wavelet transform technology and adopt the Daubechies8 wavelet to suppress the environmental noise in the frequency band. The obtained clean subcarriers are used as the input of the neural network model.

[0060] S2. Given the labeled data in the source domain and the unlabeled data in the target domain , generate virtual samples through the generator composed of the convolutional layer and the transposed convolutional layer and , and use the discriminator composed of the convolutional layer to distinguish the authenticity of the samples. Specifically, minimize the following adversarial loss function:

[0061]

[0062]

[0063] To better generate the active data in the target domain, use the output of the fully connected layer as the attention mask to assign more weights to the active features. The specific method is as follows in the formula:

[0064]

[0065]

[0066] where represents the fully connected layer, represents the output of the first four convolutional neural network layers of the generator, represents the weights of the corresponding layer. The finally generated virtual samples are as follows:

[0067]

[0068] where is the feature map generated by the second convolutional layer.

[0069] In addition, when learning the mapping between the generator and the discriminator, it is required that and . Therefore, adopt the following cycle consistency loss to make the network more stable. The cycle consistency loss is defined as:

[0070] .

[0071] S3. Assume that CSI frames and labels (true labels of activities) are collected in an environment (referred to as the original environment, the source domain). The first step is to train the source representation mapping (source encoder) and an accurate source gesture classifier . The present invention uses a convolutional neural network (CNN) architecture, as shown in the appendix Figure 2 . It consists of a cascade of three 2D convolutional layers and downsampling layers. The output dimensions of the convolutional layers and the kernel size are , and respectively. A non-linear activation function (ReLu) is used to extract local features. The downsampling layer aims to reduce the dimensionality of the data while ensuring the invariance of the feature map through max pooling. The size of the feature tensor obtained through is . It is mapped to a binary label output through a classifier composed of fully connected layers . The first layer of the fully connected layer uses ReLu as the activation function, and the second layer uses Softmax as the activation function. To prevent the network from overfitting to the source domain data, we add a dropout layer after the second linear layer. After calculating the cross-entropy loss, Adam' is used as the optimizer. Backpropagation is performed layer by layer to optimize the model parameters. The objective can be summarized as the following optimization:

[0072] .

[0073] Given a labeled source domain including a dataset of samples and the corresponding label set , each sample is associated with a label in . Given an unlabeled target domain consisting of samples; the dataset is divided into two source domain datasets and , and corresponding classifiers and are designed for and ; the target domain encoder is initialized with the parameters of the source domain encoder and extracts features. The classifier first inputs the corresponding source domain features to calculate the cross-entropy loss:

[0074]

[0075] represents the number of samples in the corresponding source domain, represents the target domain encoder, represents the transpose of the label vector.

[0076] S4. First, fix the trained and . To improve the classifier's ability to identify active features, we send the features extracted from another source domain into the classifier of the current source domain to calculate the cross-entropy loss and perform cross-adversarial training;

[0077]

[0078] Through adversarial learning between the loss functions, the recognition abilities of the two classifiers will gradually improve, forcing the target domain encoder to extract domain-invariant features and ensuring and the consistency of the outputs before and after cross-using on the same sample. Therefore, the target domain encoder has the ability to extract domain-invariant features.

[0079] By using the unlabeled samples of the target domain to learn robust and discriminative feature representations. As mentioned above, and can correctly identify the activities corresponding to the source domain data. If only cross-trained using source domain data, and may lose the original ability to recognize new data due to environmental changes. To solve this problem, use the discriminative information of the unlabeled target domain samples to reduce the following consistency loss

[0080]

[0081] where, represents the L1-norm, is the number of target domain samples in a batch, is the classifier and output dimension.

[0082] To prevent the extracted features from overfitting to misclassified data, we use label smoothing to soften the labels of the samples. Specifically, when calculating the cross-entropy loss function, instead of directly using one-hot labels, we use a weighted mixture of one-hot labels and a uniform distribution. The specific formula is:

[0083]

[0084] Among them, represents the domain label; is a smoothing parameter, set to 0.3.

[0085] To further improve the discriminability of the learned features, a triplet loss function is defined, which is mainly used in deep learning to train samples with small differences. For each batch, P activities are randomly selected, and K pieces of data are randomly selected for each activity. The triplet loss is expressed as;

[0086]

[0087] Among them represents the anchor of the th sample, represents the positive sample of the th sample, represents the negative sample of the th sample; is the margin parameter of the triplet loss. If it is set too small, it will be difficult to distinguish between positive and negative samples, and its value is empirically set to 0.3.

[0088] S5. Design the source domain environment discriminator network Reduce the distribution difference between the features of the target domain space and the source domain space. It contains three fully connected layers to map the encoder features to a binary output. Among them, the first two layers use the ReLU activation function, and the last layer uses the LogSoftmax activation function. Here, a dropout layer is also added. Similar to the adversarial training mode in the generative adversarial network, but the labels are domain labels (source and target), rather than fake labels and real labels. We send the features extracted from the source domain data and the features extracted from the target domain data into the domain discriminator network structure, so that the discriminator cannot distinguish the domain labels of the source and target samples, thereby improving the generalization performance of the model when applied to a new domain. The adversarial loss can be expressed as;

[0089]

[0090] .

[0091] S6. Select some samples in the target domain and improve the recognition performance of the model by assigning pseudo-labels. When the recognition results of the two classifiers and are the same and the sum of the output results is greater than the set threshold, assign pseudo-labels to the target samples. This process is expressed as:

[0092]

[0093]

[0094] Represents the recognition result of the classifier Represents the recognition result of the classifier m Represents the subscript of the maximum value among them Represents Soft label prediction of, average similarity probability The threshold is set to 0.8

[0095] S7. According to the above selection criteria, select some samples from the unlabeled target domain to further optimize the model. Optimize the target domain encoder by minimizing the self-prediction loss , and the expression of the self-prediction loss is:

[0096] .

[0097] S8. Optimize the target domain encoder by minimizing the final loss .

[0098] The final loss function is

[0099]

[0100] Among them, , , , respectively represent the weights of the consistency loss, adversarial loss, triplet loss, and self-prediction loss functions

[0101] S9. In the target domain with low recognition accuracy, use the pre-trained model to reload the model parameters. During this process, fine-tune the entire network, and the weights of the fully connected layer and Softmax layer in the classifier are retrained with a higher learning rate. After a small number of iterations, a better model adapted to the new scenario is obtained

[0102] When fine-tuning the entire network, the cross-entropy loss function is:

[0103]

[0104] Among them represents a small amount of labeled target domain data

[0105] A cross-environment human activity recognition method based on a dual adversarial network according to the present invention does not require the labeled activity information in the target domain, effectively suppresses the problem that the model trained in the source domain cannot be applied to the target environment due to environmental changes, and improves the cross-environment recognition performance of the activity sensing model

[0106] ​​Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements do not enable the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A cross-environment human activity recognition method based on a dual adversarial network, characterized in that It includes the following steps: S1. For all CSI data collected in the source domain and the target domain, interpolate the subcarriers into equidistant time series data according to the timestamp information, and use a filter to denoise the signal to obtain clean subcarriers as the input of the neural network model; S2. Use the silhouette coefficient to judge whether the distribution of activities in the target domain converges. For the target domain where the same activity distribution has a large distance and is difficult to converge, use the technology combining the generative adversarial network and self-attention to generate virtual samples that conform to the activity feature distribution of the target domain, and input them into the model for training; specifically, given the target domain samples and the source domain samples, minimize the following adversarial loss function: ; ; Among them represents the discriminator of the source domain represents the discriminator of the target domain represents the generator of the source domain represents the generator of the target domain represents the source domain dataset represents the target domain dataset; S3. Given a labeled source domain including a dataset of samples and the corresponding label set , each sample is associated with a label in it; Given an unlabeled target domain dataset consisting of samples; Divide the dataset into two source domain datasets and , and design corresponding classifiers and for them and ; The classifier inputs the corresponding source domain features to calculate the cross-entropy loss: ; represents the number of samples in the corresponding source domain, represents the target domain encoder, represents the transpose of the label vector; S4. Fix the trained classifier and , input the features extracted from another source domain into the classifier of the current source domain to calculate the cross-entropy loss, and perform cross-adversarial training; ; S5. Combine domain adversarial training in the cross-adversarial training of classifiers, and send the features extracted from the source domain data and the features extracted from the target domain data to the environment discriminator of the source domain , perform domain adaptive adversarial training by learning the mapping features of the target domain, calculate the adversarial loss in the same batch of training, and the adversarial loss function is expressed as: ; ; S6. Use the self-prediction learning method to mine discriminative information in unlabeled target domain data. Utilize the similarity between source domain samples and a batch of target domain samples to establish a matching criterion, select samples in the target domain to assign pseudo-labels. When the recognition results of two classifiers and are the same and the sum of the output results is greater than a set threshold, assign a pseudo-label to the target sample. This process is expressed as: ; ; Indicates the recognition result of the classifier , Indicates the recognition result of the classifier , m Indicates the subscript of the maximum value among them Indicates the soft label prediction of the average similarity probability; S7. Calculate the self-prediction loss, and the expression of the self-prediction loss is: ; S8. Optimize the target domain encoder by minimizing the final loss ; S9. When the recognition accuracy in the target domain is low, use the fine-tuning technology in transfer learning, that is, use a small number of labeled samples in the target domain to make the model better adapt to the new scenario; Reload the parameters of the pre-trained model, and the weights of the fully connected layer and the Softmax layer in the classifier are retrained with a higher learning rate.

2. The cross-environment human activity recognition method based on a dual adversarial network according to claim 1, characterized in that In S1, the filter signal denoising is specifically as follows: use the hampel filter to remove the outliers far from the data stream trend; remove the high-frequency noise through a sixth-order Butterworth low-pass filter; use the Daubechies8 wavelet to suppress the environmental noise in the frequency band.

3. A cross-environment human activity recognition method based on a dual adversarial network according to claim 1, characterized in that: In S2, the virtual sample generation method is to use the self-attention mechanism to assign more weights to the activity features, and the specific method is as follows in the formula: ; ; Among them represents the fully connected layer, represents the output of the first 4 convolutional neural network layers of the generator, represents the weights of the corresponding layer; the finally generated virtual samples are shown as follows: ; Among them is the feature map generated by the second convolutional layer; When learning the mapping between the generator and the discriminator, it is required that and ; The following cycle consistency loss is adopted to make the network more stable; The cycle consistency loss is defined as: 。 4. A cross-environment human activity recognition method based on a dual adversarial network according to claim 1, characterized in that: In S4, the discriminative information of the unlabeled target domain samples is used to reduce the consistency loss: ; Among them, represents the L1-norm, is the number of target domain samples within a batch, is the classifier and is the output dimension.

5. A cross-environment human activity recognition method based on a dual adversarial network according to claim 1, characterized in that: In S4, when calculating the cross-entropy loss function, use a weighted mixture of one-hot labels and a uniform distribution, and the specific formula is ; Among them, represents a domain label; is a smoothing parameter.

6. The cross - environment human activity recognition method based on a dual adversarial network according to claim 1, characterized in that: In S4, define the triplet loss function, randomly select P activities, and randomly select K pieces of data for each activity. The triplet loss is expressed as; ; Among them represents the anchor of the th sample, represents the positive sample of the th sample, represents the negative sample of the th sample; is the margin parameter of the triplet loss. If it is set too small, it will be difficult to distinguish between positive and negative samples.

7. A cross - environment human activity recognition method based on a dual adversarial network according to claim 1, characterized in that: In S8, the final loss function is; ; Among them, , , , respectively represent the weights of the consistency loss, adversarial loss, triplet loss, and self-prediction loss functions.

8. A cross-environment human activity recognition method based on a dual adversarial network according to claim 1, characterized in that: Define the cross-entropy loss function in S9: ; Among them represents a small amount of labeled target domain data.