Electronic nose data correction method based on generative adversarial network

Through the electronic nose data correction method based on the generative adversarial network, the data distribution inconsistency caused by sensor drift is solved, and the classification accuracy and application range of the electronic nose system are improved.

CN114511031BActive Publication Date: 2025-07-01CHONGQING UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210138121.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-15
Publication Date
2025-07-01
Estimated Expiration
2042-02-15

AI Technical Summary

Technical Problem

Sensor drift leads to inconsistent distribution of electronic nose data, which leads to intra-class asymmetric between data sets, affecting the classification accuracy of machine learning models, and limiting the promotion and application of electronic nose systems.

Method used

The electronic nose data correction method based on the generative adversarial network is adopted, and the FEDA neural network is built to carry out domain adversarial training, and the L2 norm of the domain invariant features extracted by the feature extractor is calculated, and the distribution differences between the source domain and the target domain are reduced and the intra-class homogeneity is increased through adaptive feature norm loss and class conditional entropy loss.

Benefits of technology

The distribution differences between the source domain and the target domain are reduced, the intra-class homogeneity is increased, the domain adaptation problem of electronic nose data is solved, and the classification accuracy of the sensor drift data set is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114511031B_ABST
    Figure CN114511031B_ABST
Patent Text Reader

Abstract

The electronic nose data correction method based on the generative adversarial network of the present invention includes: 1) building a neural network named FEDA, 2) performing domain adversarial training: adding a gradient reversal layer to the feature extractor G f and the domain discriminator G d respectively. First, during the forward propagation of data, train the feature extractor G f to learn domain-invariant features, so that the domain discriminator G d cannot distinguish whether the features come from the source domain or the target domain. Then, by minimizing the domain classification loss L d train the domain discriminator G d so that the domain discriminator G d can distinguish source domain and target domain features; then, when the data propagates backward through the gradient reversal layer, reverse the gradient, so that the feature extractor G f cannot correctly judge the domain-invariant features, thereby completing the adversarial training. The electronic nose data correction method based on the generative adversarial network of the present invention reduces the distribution difference between the source domain and the target domain, increases the within-class homogeneity, solves the domain adaptation problem of electronic nose data, and can improve the classification accuracy rate of the sensor drift data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sensor data recognition, and in particular to a correction method for electronic nose data. Background Art

[0002] Sensor drift refers to the phenomenon that the output of a sensor changes over time while the input remains unchanged. One cause of sensor drift is non-subjective factors such as sensor aging, poisoning, or environmental fluctuations. The resulting sensor drift data set includes a long-term drift data set and a short-term drift data set. One cause of sensor drift is inter-board differences, that is, the deviations caused by the sensors and corresponding hardware during manufacturing. The resulting sensor drift data is an inter-board difference data set. In addition to time drift and inter-board drift, a more complex situation is that the sensor has both time drift and inter-board drift. The resulting sensor drift data set is a mixed drift data set.

[0003] The default assumption of machine learning is that the training set and test data are independent and identically distributed. The above two phenomena make it impossible for existing models to accurately classify data that produces drift (sensor drift and inter-board differences are collectively referred to as drift). Specifically in the field of electronic nose systems, sensor drift is an unavoidable problem for electronic nose systems. Electronic nose data has inconsistent data distribution due to time drift or inter-board differences, which in turn leads to intra-class heterogeneity between data sets, affecting the classification accuracy of machine learning models, thereby limiting the promotion and application of electronic nose systems. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide an electronic nose data correction method based on a generative adversarial network to solve the technical problem that sensor drift causes inconsistent electronic data distribution, which in turn leads to intra-class heterogeneity between data sets.

[0005] The electronic nose data correction method based on the generative adversarial network of the present invention comprises the following steps:

[0006] 1) Building a neural network named FEDA, which includes a feature extractor G for extracting domain-invariant features of the source domain and the target domain f , domain discriminator G used to distinguish data from source domain and target domain d , L2 norm module G for computing the L2 norm loss of domain invariant features l , label classifier G for data category classification y , class conditional probability entropy G used to calculate class entropy loss e and a gradient reversal layer for performing gradient reversal, the gradient reversal layer being connected to the feature extractor G f and domain discriminator G d between,

[0007] The feature extractor G f output is used as the class-conditional probability entropy G e , the domain discriminator G d , the L2 norm module G l , and the label classifier G y input; The data is divided into a source domain with rich labels and a target domain without labels. Define the source domain where n s represents the number of source domain samples, represents the i-th sample in the source domain, represents the label of the i-th sample in the source domain; where n t represents the number of source domain samples, represents the j-th sample in the target domain; The distribution of the source domain data is P(X s ,Y s ), and the distribution of the target domain data is Q(X t ,Y t ), P≠Q;

[0008] 2) Perform domain adversarial training: Add a gradient reversal layer to the feature extractor G f and the domain discriminator G d respectively. First, train the feature extractor G f during the forward propagation of the data to learn domain-invariant features, so that the domain discriminator G d cannot distinguish whether the features come from the source domain or the target domain. Then, minimize the domain classification loss L d to train the domain discriminator G d so that the domain discriminator G d can distinguish the source domain and target domain features; Then, when the data is backpropagated through the gradient reversal layer, reverse the gradient so that the feature extractor G f cannot correctly judge the domain-invariant features, thus completing the adversarial training;

[0009] During the domain adversarial training process, calculate the L2 norm of the features extracted by the feature extractor G f , and make the L2 norms of the source domain and the target domain balanced on a large scale through the adaptive feature norm loss L f ; And the class-conditional probability entropy G e adopts minimizing the target domain conditional entropy L h to reduce the inter-class overlap in the target domain and increase the intra-class homogeneity.

[0010] Furthermore, in step 2), train the feature extractor G d to learn domain-invariant features through the domain classification loss L f The domain classification loss Ld It is expressed as follows:

[0011]

[0012] Where X s , X t represent the source domain and the target domain, represents the i-th sample of the source domain, represents the j-th sample of the target domain.

[0013] Furthermore, in step 2), the gradient reversal layer reverses the gradient of L d The gradient of is to replace the gradient with where σ represents the weight parameter; the pseudo-function of the gradient reversal layer is defined as R σ (x), then the gradient reversal is expressed as the following two functions:

[0014]

[0015] where I represents the identity matrix.

[0016] Furthermore, the adaptive feature norm loss in step 2) is constructed by means of the maximum mean feature distribution difference. The construction steps include:

[0017] Define the maximum mean feature distribution difference between the source domain and the target domain as follows:

[0018]

[0019] where MMFDD[G f , X s , X t is the maximum mean feature distribution difference; x i , x j represent the data of the source domain and the target domain respectively.

[0020] Construct a distance Z to fit the feature norm gap L between the source domain and the target domain Z , so that the L2 norms of the source domain and the target domain converge to Z respectively, thereby minimizing MMFDD[G f , D s , D t ,

[0021]

[0022] Then construct the adaptive feature norm loss L f , and the formula is as follows,

[0023]

[0024] Where Δz represents the residual feature norm, and w g represents the weight parameter.

[0025] Furthermore, in step 2), minimizing the target domain conditional entropy L h has the following formula:

[0026]

[0027] where C represents the number of categories, represents the prediction probability of the k-th category in the target domain, and w e represents the weight parameter.

[0028] Furthermore, the above-mentioned electronic nose data correction method based on a generative adversarial network further includes optimizing backpropagation, and the optimization objective is:

[0029]

[0030] where θ f , θ y , θ d respectively represent the parameters of G f , G y , G d , and α, β, γ respectively represent weight parameters; respectively represent the optimal parameters of θ f , θ y , θ d ; L y is the cross-entropy loss used to guide the feature extractor G f and the label classifier G y to make correct class predictions. On the source domain, L y is expressed as:

[0031]

[0032] Advantages of the present invention:

[0033] The electronic nose data correction method based on a generative adversarial network of the present invention reduces the distribution difference between the source domain and the target domain, increases the within-class homogeneity, solves the domain adaptation problem of electronic nose data, and can improve the classification accuracy rate of the sensor drift data set. Description of the drawings

[0034] Figure 1 is the overall framework diagram of the FEDA neural network. Specific implementation manners

[0035] The present invention will be further described below with reference to the drawings and embodiments.

[0036] In this embodiment, the electronic nose data correction method based on the generative adversarial network includes the following steps:

[0037] 1) Build a neural network named FEDA. The FEDA includes a feature extractor G for extracting domain-invariant features of the source domain and the target domain f , a domain discriminator G for distinguishing whether the data comes from the source domain or the target domain d , an L2 norm module G for calculating the L2 norm loss of the domain-invariant features l , a label classifier G for classifying data categories y , a class conditional probability entropy G for calculating the class entropy loss e and a gradient reversal layer for performing gradient reversal. The gradient reversal layer is connected between the feature extractor G f and the domain discriminator G d .

[0038] The output of the feature extractor G f serves as the input of the class conditional probability entropy G e , the domain discriminator G d , the L2 norm module G l , and the label classifier G y . Divide the data into a source domain with rich labels and a target domain without labels, and define the source domain where n s represents the number of source domain samples, represents the i-th sample in the source domain, represents the label of the i-th sample in the source domain; where n t represents the number of source domain samples, represents the j-th sample in the target domain; the distribution of the source domain data is P(X s ,Y s ), and the distribution of the target domain data is Q(X t ,Y t ), P≠Q.

[0039] 2) Conduct domain adversarial training: Add a gradient reversal layer to the feature extractor G f and the domain discriminator G d respectively. First, train the feature extractor G f during the forward propagation of the data to learn the domain-invariant features, so that the domain discriminator G d cannot distinguish whether the features come from the source domain or the target domain. Then, train the domain discriminator G d by minimizing the domain classification loss L d so that the domain discriminator G d can distinguish the source domain and target domain features; then reverse the gradient when the data backpropagates through the gradient reversal layer, and let the feature extractor Gf It is impossible to correctly judge the domain-invariant features to complete adversarial training.

[0040] During the domain adversarial training process, calculate the L2 norm of the features extracted by the feature extractor G f and, through the adaptive feature norm loss L f make the L2 norms of the source domain and the target domain balanced in a large range; and the class-conditional probability entropy G e adopts minimizing the conditional entropy L of the target domain h to reduce the inter-class overlap of the target domain and increase the intra-class homogeneity.

[0041] In step 2), train the feature extractor G through the domain classification loss L d to learn domain-invariant features and solve the drift problem of the electronic nose in a domain adversarial manner. The domain classification loss L f is expressed as follows: d as follows:

[0042]

[0043] where X s , X t represent the source domain and the target domain, represents the i-th sample of the source domain, represents the j-th sample of the target domain.

[0044] In step 2), the gradient reversal layer reverses the gradient of L d by replacing the gradient with where σ represents the weight parameter; define the pseudo-function of the gradient reversal layer as R σ (x), then the gradient reversal is expressed as the following two functions:

[0045]

[0046] where I represents the identity matrix.

[0047] Parameters or features with smaller norms play a smaller information role during inference. To weaken the influence of noise samples, balance the feature norms of the source domain and the target domain over a large range of values, and increase the domain transfer effect, in this embodiment, in step 2), the L2 norm of the domain-invariant features extracted by the feature extractor G f is also calculated. The L2 norm is a non-negative evaluation index function. The adaptive feature norm loss described in step 2) is constructed by the method of maximum mean feature distribution difference. The construction steps include:

[0048] Define the maximum mean feature distribution difference between the source domain and the target domain as follows:

[0049]

[0050] where MMFDD[G f ,X s ,X t is the maximum average feature distribution difference; x i ,x j represent the data of the source domain and the target domain respectively.

[0051] Mathematically, the L2 norm can be understood as the distance from the center of the sphere (or the origin of the hypersphere) to the vector point. Therefore, a distance Z is constructed to fit the feature norm gap L between the source domain and the target domain Z , so that the L2 norms of the source domain and the target domain converge to Z respectively, thereby making MMFDD[G f ,D s ,D t minimum,

[0052]

[0053] When constructing formula (4), it is estimated from the overall data. Since the value of Z is determined in advance, the value of Z will determine whether the distance gap between the source domain and the target domain is optimized properly, without considering the local instance changes. In order to make Z instance-adaptive, an adaptive feature norm loss L is constructed f , and the formula is as follows,

[0054]

[0055] where Δz represents the residual feature norm, w g represents the weight parameter. Formula (5) can minimize MMFDD[G f ,D s ,D t while avoiding the size limitation of the mini-batch distribution and the change of the statistical estimator. Since Z changes with the data, this can better balance the feature norms of the source domain and the target domain on a large scale.

[0056] In step 2), the formula for minimizing the conditional entropy L of the target domain h is as follows,

[0057]

[0058] where C represents the number of classes, and both represent the prediction probabilities of the k-th class in the target domain, w eDenote the weight parameter. The smaller the entropy, the lower the degree of chaos of the model, and the greater the information accuracy. The low-entropy class prediction can effectively promote the low-density separation between classes. For the unlabeled target domain, reducing its chaos can effectively perform classification. Conditional entropy is a measure of class overlap. Therefore, minimizing the conditional entropy L of the target domain h can reduce the overlap between classes and improve the high compactness within classes, thereby reducing the non-homogeneity within classes.

[0059] As an improvement to the above embodiments, the electronic nose data correction method based on a generative adversarial network further includes optimizing backpropagation to obtain an optimal function solution, and the optimization objective is:

[0060]

[0061] where θ f , θ y , θ d respectively represent the parameters of G f , G y , G d , and α, β, γ respectively represent weight parameters; respectively represent the optimal parameters of θ f , θ y , θ d ; L y is the cross-entropy loss used to guide the feature extractor G f and the label classifier G y to make correct class predictions. On the source domain, L y is expressed as:

[0062]

[0063] Define the pseudo-objective function as follows:

[0064]

[0065] where R is the abbreviation of the gradient reversal layer, and α, β, γ are weight coefficients. Based on the above description, we can find the saddle point of formula (9):

[0066]

[0067] At the saddle point, the parameters θ d of the domain discriminator G d minimize the domain classification loss, and the parameters θ y of the label classifier G y minimize the label prediction loss, and the parameters θ f of the feature extractor G fMinimize the label prediction loss (to make the features discriminative) while maximizing the domain classification loss (to make the features domain-invariant). In this embodiment, the stochastic gradient descent method is used to update these parameters.

[0068] A saddle point can be found as the stationary point of the following stochastic update

[0069]

[0070] where η represents the learning rate, which is generally set dynamically. represents the corresponding loss function calculated at the i-th training example.

[0071] Next, experiments are carried out using four datasets of long-term drift, short-term drift, inter-board difference, and mixed drift of the electronic nose to verify the effectiveness of the method proposed in the embodiment.

[0072] Experiment on the long-term drift dataset

[0073] For the long-term drift experiment, an open long-term drift dataset of electronic nose sensors is used and experiments are carried out according to Experiment Setting 1, with the classification accuracy as the evaluation index.

[0074] Table 1 Long-term drift sensor dataset

[0075]

[0076] Experiment Setting 1: Batch1 is used as the training set (source domain) for training, and Batch K (K = 2, 3, 4,..., 10) is used as the test set (target domain) for testing.

[0077] Table 2 shows the classification accuracies of different algorithms under Experiment Setting 1. It can be clearly seen from the table that the classification accuracy of the method FEDA disclosed in the embodiment is 82.67%, which is 4.57 percentage points higher than the newly proposed subspace projection method MSCP-SSS. This shows that the method proposed in this embodiment can correct the distribution drift of sensor data well.

[0078] Table 2 Comparison of class accuracies of different methods on the long-term drift sensor dataset under Setting 1

[0079]

[0080] Experiment on the short-term drift dataset

[0081] For the short-term drift problem of the electronic nose, an electronic nose short-term drift dataset is used and experiments are carried out according to Experiment Setting 2, with the classification accuracy as the evaluation index.

[0082] Table 3 Short-term drift sensor dataset

[0083]

[0084] Experimental setup 2: Batch 1 is used as the source domain for model training, and Batch 2 is used as the target domain for model testing.

[0085] Table 4 Comparison of class accuracies of different methods on the short-term drift sensor dataset under Setup 2

[0086]

[0087] It can be clearly seen from Table 4 that the method proposed in the embodiment achieves the highest classification accuracy, which is 4.44 percentage points higher than the MCSP-SSS method and 3.12 percentage points higher than the recently proposed WDGRL method, indicating that the method proposed in this embodiment can effectively solve the short-term time drift problem of the electronic nose.

[0088] Experiment on the dataset of inter-board differences

[0089] For the inter-board difference problem of the electronic nose, the publicly available inter-board difference dataset (https: / / jordifonollosa.wordpress.com / downloads / public-datasets / ) and the lung cancer electronic nose dataset of the Chongqing Key Laboratory of Bio-Perception and Intelligent Information Processing (see the paper "MCSP-SSS: A Domain Adaptive Framework for High-Accuracy Sensor Data Classification") are used to conduct experimental verification according to Experimental Setup 3 and Experimental Setup 4, and the classification accuracy is used as the evaluation index.

[0090] Table 5 Public inter-board drift electronic nose dataset

[0091]

[0092] Experimental setup 3: Batch 1 is used as the source domain for model training, and Batch K (K = 2, 3, 4, 5) is used as the target domain for model testing.

[0093] Experimental setup 4: Prototype 1 is used as the source domain for model training, and Prototype 2 is used as the target domain for model testing.

[0094] The public inter-board difference electronic nose dataset is experimented according to Experimental Setup 3, and the experimental results are listed in Table 6. It can be clearly seen from the table that FEDA has achieved good classification results on the public inter-board, indicating that the method proposed in this embodiment can solve the inter-board difference problem of the electronic nose and improve the classification accuracy of the model.

[0095] Comparison of Class Accuracy of Different Methods on the Public Inter-board Difference Sensor Dataset under Setting 3 in Table 6

[0096]

[0097]

[0098] Experiments on the Hybrid Drift Dataset

[0099] For the hybrid drift problem of electronic nose sensors, that is, there are both inter-board differences and time drifts in the electronic nose. We use the publicly available hybrid drift dataset (see the paper "Anti-drift in E-nose: A subspace projection approach with drift reduction") and conduct experimental verification according to Experiment Setting 5, with classification accuracy as the evaluation index.

[0100] Table 7 Hybrid Drift Electronic Nose Dataset

[0101]

[0102] Experiment Setting 5: Master is used as the source domain for model training, and Salve is used as the target domain for model testing.

[0103] Comparison of Class Accuracy of Different Methods on the Hybrid Drift Sensor Dataset under Setting 5 in Table 8

[0104]

[0105] It can be seen from Table 8 that the method proposed in this embodiment has a good classification effect on the hybrid experimental dataset, which indicates that the method proposed in this embodiment can handle the hybrid drift problem of the electronic nose.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An electronic nose data correction method based on a generative adversarial network, characterized in that: Including the following steps: 1) Build a neural network named FEDA, where the FEDA includes a feature extractor G for extracting domain-invariant features of the source domain and the target domain f , a domain discriminator G for distinguishing whether the data comes from the source domain or the target domain d , an L2 norm module G for calculating the L2 norm loss of the domain-invariant features l , a label classifier G for classifying data categories y , a class conditional probability entropy G for calculating the class entropy loss e and a gradient reversal layer for performing gradient reversal, where the gradient reversal layer is connected between the feature extractor G f and the domain discriminator G d ; The feature extractor G f outputs the class-conditional probability entropy G e , the domain discriminator G d , the L2-norm module G l , and the label classifier G y as inputs; the data is divided into a source domain with rich labels and a target domain without labels. Define the source domain where n s represents the number of source domain samples, represents the i-th sample in the source domain, and represents the label of the i-th sample in the source domain; where n t represents the number of source domain samples, represents the j-th sample in the target domain; the distribution of the source domain data is P(X s , Y s ), and the distribution of the target domain data is Q(X t , Y t ), P ≠ Q; 2) Perform domain adversarial training: Add a gradient reversal layer to the feature extractor G f and the domain discriminator G d respectively. First, train the feature extractor G f during the forward propagation of data to learn domain-invariant features, so that the domain discriminator G d cannot distinguish whether the features come from the source domain or the target domain. Then, minimize the domain classification loss L d to train the domain discriminator G d so that the domain discriminator G d can distinguish source domain and target domain features; then reverse the gradient when the data backpropagates through the gradient reversal layer, making the feature extractor G f unable to correctly judge the domain-invariant features, thus completing the adversarial training; During the domain adversarial training process, calculate the L2 norm of the features extracted by the feature extractor G f and, through the adaptive feature norm loss L f balance the L2 norms of the source domain and the target domain on a large scale; and the class-conditional probability entropy G e adopts minimizing the conditional entropy L of the target domain h to reduce the inter-class overlap of the target domain and increase the intra-class homogeneity; The adaptive feature norm loss L f is constructed by means of the maximum mean feature distribution difference, and the construction steps include: Define the maximum mean feature distribution difference between the source domain and the target domain as follows: where MMFDD[G f ,X s ,X t is the maximum mean feature distribution difference; x i and x j respectively represent the data in the source domain and the target domain; construct a distance Z to fit the feature norm gap L between the source domain and the target domain Z , so that the L2 norms of the source domain and the target domain converge to Z respectively, thus making MMFDD[G f , X s , X t minimized Reconstruct the adaptive feature norm loss L f , and the formula is as follows Among them Δz represents the residual feature norm, and w g represents the weight parameter.

2. The electronic nose data correction method based on a generative adversarial network according to claim 1, characterized in that: In step 2), the feature extractor G is trained through the domain classification loss L d to learn domain-invariant features. The domain classification loss L f is expressed as follows: d is expressed as follows: where X s , X t represent the source domain and the target domain, represents the i-th sample of the source domain, represents the j-th sample of the target domain.

3. The method for calibrating electronic nose data based on a generative adversarial network according to claim 2, wherein: In step 2), the gradient reversal layer reverses the gradient of L d The gradient of is replaced with where σ represents the weight parameter; the pseudo-function of the gradient reversal layer is defined as R σ (x), then the gradient reversal is expressed as the following two functions: Where I represents the identity matrix and x represents the input of the function.

4. The method for calibrating electronic nose data based on a generative adversarial network according to claim 3, wherein: Minimize the target domain conditional entropy \(L\) in step 2) h The formula is as follows: where C represents the number of categories, represents the predicted probability of the k-th class in the target domain, and w e represents the weight parameter.

5. The method for correcting electronic nose data based on a generative adversarial network according to claim 4, wherein: It also includes optimizing backpropagation to obtain the optimal function solution, and its optimization objective is: where θ f , θ y , θ d represent the parameters of G f , G y , G d respectively, and α, β, γ represent the weight parameters; represent the optimal parameters of θ f , θ y , θ d respectively; L y is the cross - entropy loss used to guide the feature extractor G f and the label classifier G y to make correct class predictions. On the source domain, L y is expressed as: y i,k represents the i-th sample of the k-th class.

Citation Information

Patent Citations

  • IFTS spectrum processing method based on multi-step micro-reflector

    CN111208081A

  • Unsupervised depth field adaptation method based on distributed confrontation

    CN113011523A