Gas classification and identification method based on graph neural network of contrast learning and edge label framework
Patent Information
- Application Number
- CN202411219858.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-02
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2044-09-02
AI Technical Summary
[0006]现有技术缺点:传统的图神经网络(GNN)通常需要标记样本进行模型训练,且仅允许模型在有监督的设置下进行训练,导致模型对大量标记样本具有较高的依赖性,模型性能易受标记信息限制影响,且数据获取成本较高
[0111]The beneficial effects of this invention are as follows: This invention introduces the concept of contrastive learning into the CEGNNDrift model. This improvement not only significantly enhances model performance, but more importantly, it allows the model to be trained not only in supervised settings but also in unsupervised settings. This ability to train without supervision significantly reduces the model's dependence on a large number of labeled samples, thus providing a more efficient and practical solution in situations where data acquisition costs are high or labeled information is limited.
Smart Images

Figure CN119128722B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gas identification technology, and in particular to a graph neural network gas classification and identification method based on contrastive learning and edge labeling framework. Background Technology
[0002] Although the performance of chemical sensors is influenced by sensor properties, the gradual and unpredictable changes in their chemical sensory signal responses when exposed to the same analytes under identical conditions—known as sensor drift—have long been considered a major challenge in the field of chemical sensing. This phenomenon reduces the stability of electronic nose systems and renders models used to distinguish gases or odors obsolete. The causes can stem from a variety of effects, such as sensor aging from long-term use, poisoning of sensitive materials due to contamination, and thermal effects on the sensor during different operating processes. While these drifts are generally classified as first- and second-order components, it is difficult to empirically distinguish them, and it is not possible to develop methods suitable for correcting different drift sources because the origins are not yet determined. Therefore, regarding drift reduction, some researchers focus on high-performance sensor materials that can reversibly interact with gases, some minimize the irreversible effects of sensor poisoning through temperature regulation techniques, while others argue that drift countermeasures for both first- and second-order components should be implemented from a long-term perspective. Nevertheless, the drift mechanisms are so complex and unavoidable that they have plagued the sensor research community for many years and are likely to continue to do so.
[0003] Traditional methods mainly address the drift compensation problem of electronic noses from two aspects:
[0004] (1) Based on signal preprocessing techniques, the effects of drift can be reduced by improving data quality. For example, baseline processing methods improve the signal by subtracting a reference baseline value from the sensor response, and frequency domain filtering methods can filter out noise components in the signal by focusing on a specific frequency domain, especially those components that are irrelevant to the detection of the target gas.
[0005] (2) Drift component correction methods improve model stability and accuracy by correcting drift components in the data. These include Orthogonal Signal Correction (OSC): correcting drift by removing information from variables irrelevant to the target variable; Principal Component Analysis (PCA): extracting key features through data dimensionality reduction to reveal the most important hidden structures in the data; Generalized Least Squares Weighted (GLSW): optimizing the model's fit by assigning different weights to different data points; and Direct Standardization (DS): adjusting the model to accommodate potential drift by comparing the responses of standardized samples with those of actual samples. These methods have largely been replaced in recent years due to their lower sensitivity and accuracy.
[0006] Disadvantages of existing technologies: Traditional graph neural networks (GNNs) usually require labeled samples for model training and only allow the model to be trained under supervised settings. This results in the model being highly dependent on a large number of labeled samples, and the model performance is easily affected by the limitations of labeled information. In addition, the cost of data acquisition is high. Summary of the Invention
[0007] This invention provides a graph neural network-based gas classification and recognition method based on contrastive learning and edge labeling framework, which effectively improves the accuracy of gas classification results.
[0008] To achieve the above objectives, this invention provides a graph neural network gas classification and recognition method based on contrastive learning and edge labeling framework, the key of which includes a model training step A and a gas classification and recognition step B.
[0009] The model training step A includes the following steps:
[0010] Step A1: Construct a graph neural network training model, which includes a feature enhancement module, an encoder, a momentum encoder, a graph construction module, and a graph update module.
[0011] The feature enhancement module is connected to the encoder and the momentum encoder, the encoder and the momentum encoder are connected to the graph construction module, and the graph construction module is connected to the graph update module;
[0012] Step A2: The graph neural network training model is trained using a contrastive learning strategy under self-supervised learning. After training is completed, the momentum encoder, graph construction module and graph update module are retained. The momentum encoder is connected to the graph construction module and the graph construction module is connected to the graph update module, thereby obtaining a graph neural network based on contrastive learning and edge labeling framework.
[0013] Through the above design, a contrastive learning strategy under self-supervised learning is used during model training, and a momentum encoder is introduced to make the sample representations more consistent. When training can utilize labels from the training set, both self-supervised and supervised learning are performed simultaneously to improve model performance. When the data in the training set is unlabeled, self-supervised learning is used for training to improve model performance.
[0014] The graph neural network training model provided by this invention allows the model to be trained not only under supervised settings but also under unsupervised settings. This unsupervised training capability significantly reduces the model's dependence on a large number of labeled samples, thus providing a more efficient and practical solution in situations where data acquisition costs are high or labeled information is limited.
[0015] The gas classification and identification step B includes the following steps:
[0016] Step B1: The gas sensor array acquires gas characteristic data a1 in real time and transmits the gas characteristic data a1 to the preprocessing module;
[0017] Step B2: The preprocessing module performs preprocessing operations on the gas feature data a1 to obtain query set sample a2, and then passes it to the graph neural network;
[0018] Step B3: The input layer of the graph neural network inputs the query set sample a2 and N support set samples into the momentum encoder, where each support set sample contains R samples;
[0019] Step B4: The momentum encoder performs feature extraction operations on the query set sample a2 and the N-class support set samples respectively to obtain the query set feature vector a3 and the N-class support set feature vector, and then passes them to the graph construction module;
[0020] Step B5: The graph construction module constructs a fully connected graph a4 from the query set feature vector a3 and the N-class support set feature vectors, and initializes the node features and edge features in the fully connected graph a4 to obtain an initialized feature graph a5, which is then passed to the graph update module.
[0021] Step B6: The graph update module performs iterative operations on node feature update and edge feature update on the initial feature graph a5. After L iterations, it outputs the clustering feature graph a6 to the classification and recognition module.
[0022] Step B7: The classification and recognition module classifies and recognizes the cluster feature map a6 and outputs the gas classification result.
[0023] The preprocessing operation involves normalizing and standardizing the input data to make it conform to the input requirements of a graph neural network based on contrastive learning and edge labeling framework.
[0024] Through the above design, the similarity and differences between samples are explicitly simulated by iteratively updating edge labels. Compared with traditional GNN models, the CEGNNDrift graph neural network based on contrastive learning and edge labeling framework does not rely on implicit clustering similarity simulation, but directly uses edge labels to guide the learning process. This explicit clustering evolution enables the model to more accurately capture structural information between samples, thereby improving the ability to identify new gas categories.
[0025] Preferably, in step A2, a contrastive learning strategy under self-supervised learning is used to train the graph neural network training model, and the specific steps are as follows:
[0026] Step A21: The feature enhancement module acquires m unlabeled gas samples, and the first enhancement unit in the feature enhancement module performs feature enhancement operation on the unlabeled gas samples and outputs the first enhanced sample set to the encoder;
[0027] The second enhancement unit in the feature enhancement module performs feature enhancement operations on the unlabeled gas sample set and outputs the second enhanced sample set to the momentum encoder.
[0028] Step A22: The encoder performs feature extraction on the first enhanced sample set to obtain a first feature vector set, and passes it to the graph construction module;
[0029] The momentum encoder performs feature extraction on the second enhanced sample set to obtain a second feature vector set, which is then passed to the graph construction module.
[0030] Step A23: The graph construction module constructs the first feature vector set and the second feature vector set into a fully connected graph, initializes the node features and edge features in the fully connected graph, and then passes the initialized feature graph to the graph update module;
[0031] Step A24: The graph update module performs iterative operations on node feature update and edge feature update on the initialized feature map. After L iterations, it outputs the clustering feature map to the classification and recognition module.
[0032] Step A25: The classification and recognition module classifies and recognizes the clustering feature map, and calculates the model loss based on the classification and recognition results;
[0033] Step A26: Update the encoder parameters and graph update module parameters according to the model loss, and then update the momentum encoder parameters according to the encoder parameters;
[0034] Step A27: Repeat steps A21-A26 to obtain the optimal graph neural network training model. Retain the momentum encoder, graph construction module, and graph update module, as well as the optimal parameters of each module, to obtain a graph neural network based on contrastive learning and edge labeling framework.
[0035] The optimal parameters are those used by the model when the calculated model loss is less than the loss threshold during training. Under these parameter conditions, the model can achieve accurate classification of gas samples and exhibits excellent model performance.
[0036] Through the above design, unsupervised training of graph neural network training models is achieved, significantly reducing the model's dependence on a large number of labeled samples.
[0037] Preferably, in step A21, the first enhancement unit performs feature enhancement on the unlabeled gas sample using a feature zeroing enhancement method. The feature zeroing enhancement method is used for a given sample x. i , where each feature x i,j The probability of p1 is set to 0, as shown in the following expression:
[0038]
[0039] Where, x′ i,j Indicates sample x i After the enhanced j-th feature, x i,j Indicates sample x i The j-th feature;
[0040] The second enhancement unit uses an additive Gaussian perturbation enhancement method to perform feature enhancement on the unlabeled gas sample. The additive Gaussian perturbation enhancement method is used for a given sample x. i Each eigenvalue x i,j It will be multiplied by a number drawn from a Gaussian distribution with a mean of 1 and a variance of p², as shown in the following expression:
[0041] x′ i,j =x i,j ·(1+N(0, p2))
[0042] Where N(0, p2) represents a random number drawn from a normal distribution with mean 0 and variance p2.
[0043] Sample x i New samples were obtained through data augmentation methods. and Then, the new samples are processed by encoder E1 and momentum encoder E2 to obtain feature vectors. and
[0044] For sample x i Two different data augmentation methods were used to augment the data, resulting in two new samples. and Treat these two augmented samples as one class, and treat the samples augmented from these two samples and other samples as different classes. This will give you a series of positive and negative sample pairs.
[0045] In contrastive learning, the model is trained in this way to identify and distinguish subtle differences between different samples. Specifically, for each sample x... i Augmented samples generated using data augmentation methods A and B and These are considered positive sample pairs because they originate from the same original sample. Other samples, such as those generated after data augmentation, are considered positive samples. and (where j≠k), are considered negative sample pairs because they come from different original samples.
[0046] Preferably, in step A22, the encoder and the momentum encoder have the same structure and the same initial parameters;
[0047] In step A26, the encoder parameters are fully updated based on the model loss, and the momentum encoder parameters are partially updated based on the encoder parameters.
[0048] The core idea of a momentum encoder is to use a momentum-updated encoder. The parameters of the momentum encoder are updated using the network parameters with a small momentum coefficient. After the encoder parameters are updated, the momentum encoder parameters are updated again. The momentum encoder parameter update expression is as follows:
[0049]
[0050] in, This represents the momentum encoder parameters updated during the (t+1)th training iteration. This represents the momentum encoder parameters updated during the t-th training iteration. This represents the encoder parameters updated during the (t+1)th training iteration. Let represent the encoder parameters updated during the t-th training iteration, and u represent the momentum coefficient.
[0051] Preferably, in step B4, the momentum encoder performs feature extraction on the query set sample a2 or the support set sample, as follows:
[0052] Step B41: The spatial attention layer in the momentum encoder acquires the query set sample a2 or the support set sample, performs spatial feature extraction on it to obtain spatial feature data s1, and passes it to the first batch of normalization layers.
[0053] Step B42: The first batch of normalization layers performs batch normalization on the spatial feature data s1 to obtain the first batch of normalized data s2, and then passes it to the first residual block;
[0054] Step B43: The first residual block performs a residual convolution operation on the first batch of normalized data s2 to obtain the first residual data s3, and then passes it to the second batch of normalization layers;
[0055] Step B44: The second batch normalization layer performs batch normalization on the first residual data s3 to obtain the second batch normalized data s4, and then passes it to the second residual block;
[0056] Step B45: The second residual block performs a residual convolution operation on the second batch of normalized data s4 to obtain the query set feature vector c or the support set feature vector, and passes it to the encoding unit.
[0057] The momentum encoder is used to extract abstract features from the input data for use by the CEGNNDrift graph neural network based on contrastive learning and edge labeling framework, thereby improving gas classification efficiency.
[0058] The momentum encoder can not only extract deep features of samples, but also enhance the learning ability and stability of the model through spatial attention mechanism (SAM) and residual connections.
[0059] Preferably, in step B41, the spatial attention layer performs spatial feature extraction on the query set sample a2 or the support set sample, as follows:
[0060] Step B411: Obtain a one-dimensional feature map: Assume the input data of the spatial attention layer is X∈R C×W Where C and W represent the number of channels and channel width of the input data, respectively; by performing average pooling and max pooling on the input data X, the average pooling feature map and max pooling feature map are obtained, as expressed below:
[0061] F avg ∈R 1×W
[0062] F max ∈R 1×W
[0063] Among them, F avg F represents the average value of the average pooling feature map along the channel dimension. max This represents the maximum value of the max-pooled feature map in the channel dimension;
[0064] Step B412: Feature map concatenation: The average pooling feature map F avg and max pooling feature map F max The feature maps F are concatenated along the channel dimension. concat ∈R 2×W ;
[0065] Step B413: Obtain the spatial attention map: for the concatenated feature map F concat Perform a one-dimensional convolution operation and obtain the spatial attention map M through the sigmoid activation function. s ∈R2×W The expression is as follows:
[0066] M s =σ(Conv(F) concat ))
[0067] Where σ represents the Sigmoid function, and Conv represents one-dimensional convolution. The expression for one-dimensional convolution is as follows:
[0068]
[0069] Among them, y t It outputs the value of the feature vector at position t, x i w is the value of the input feature vector at position i. i,t-i b is the weight of the convolution kernel at position i. t It is a bias term.
[0070] Step B414: Attention Weighting: Combine the original input feature map X with the spatial attention map M s Multiply by the weighted output feature map X′:
[0071]
[0072] in, This represents element-level multiplication operations.
[0073] The Spatial Attention Mechanism (SAM) in the Spatial Attention Layer enhances the ability of the CEGNNDrift graph neural network, based on contrastive learning and edge labeling, to identify and respond to key spatial region features in the input data. This allows CEGNNDrift to dynamically focus on processing the most critical parts of the data for the current task, significantly improving processing efficiency and accuracy. Through the spatial attention mechanism, CEGNNDrift can not only better capture the essential features of the data but also optimize at the feature representation level, thereby enhancing the performance of subsequent tasks such as classification, detection, or segmentation.
[0074] Spatial Attention Mechanism (SAM) offers several significant advantages: First, it enhances model focus, enabling the model to concentrate resources on processing the most critical features in the data, thus reducing computational waste on unimportant regions. Second, SAM strengthens feature representation by highlighting important spatial features, contributing to the generation of richer and more discriminative feature representations, which is crucial for improving model performance on complex tasks. Furthermore, SAM is highly adaptable, allowing the model to adaptively adjust its focus and effectively handle data from different categories or under different conditions. SAM also offers end-to-end learning convenience, easily integrating into existing deep learning frameworks. Its generalization ability is also enhanced because the model can focus on key information in the data, thus demonstrating good performance even on unseen data. In terms of computational efficiency, SAM can improve model computation speed in certain situations by reducing processing of unimportant regions. Finally, SAM is conceptually intuitive and simple to implement, easily implemented and applied using existing deep learning libraries. These advantages make SAM a powerful tool for deep learning models processing complex spatial data.
[0075] By introducing a spatial attention mechanism (SAM) before the convolutional layer, the performance of the feature extraction module is significantly improved by enhancing the ability to focus on key information in the input feature space.
[0076] Preferably, the first residual block and the second residual block have the same structure, both of which are provided with a weight layer and an identity layer connected in parallel. The weight layer consists of 5 one-dimensional convolutional layers connected end to end, and the identity layer consists of 1 one-dimensional convolutional layer connected by residuals.
[0077] The expression for the first residual block or the second residual block is as follows:
[0078] x n+1 =F(x) n , {w n,m ))+Conv n (x n , {w′ n,m})
[0079] Where, x n The input feature is x in the first residual block. n For the first batch of normalized data s2, x in the second residual block n For the second batch of normalized data s4; x n+1 It is the output feature, x in the first residual block n+1 Given the first residual data s3, x in the second residual block n For the query set feature vector c or the support set feature vector; n represents the index of the residual layer; F(x n , {w n,m)) is the residual function of the weight layer, representing the input x. n The result after weighted layer transformation; Conv n (x n , {w′ n,m}) represents the input x n The result after identity layer transformation; {w n,m} is the weight set of the nth residual block, w′ n,m These are the parameters of the identity layer.
[0080] The convolutional layers in the residual block contain multiple convolutional kernels W, each responsible for extracting a specific type of local feature from the input data.
[0081] Within each residual block, the stacking of one-dimensional convolutional layers allows the model to learn local features at different scales. The first convolutional layer is responsible for capturing the most basic local features, while subsequent convolutional layers are able to capture more complex patterns. This stacking method not only increases the model's learning capacity but also maintains the network's stability through residual connections.
[0082] Preferably, in step B5, each node in the fully connected graph a4 represents a feature vector, and each edge represents the relationship type between two nodes; the node feature initialization expression is as follows:
[0083]
[0084] in, Let E(x) represent the feature of the i-th initial node. i ) represents the i-th input sample x i Extracted feature vectors;
[0085] The edge feature initialization expression is as follows:
[0086]
[0087] Among them, y i Let y represent the label of the i-th node. j Let y represent the label of the j-th node. ij The edge feature label between nodes i and j Represents the initial edge features between nodes i and j, each edge feature It is a two-dimensional vector, e when d=1. ij1 This represents the strength of the inter-class relationship between two connected nodes; when d = 2, e ij2 || represents the strength of the intra-class relationship between two connected nodes; || is the combination of intra-class and inter-class relationships; label represents the label, and NULL indicates that the sample does not contain a label.
[0088] Only edges between labeled samples in the support set can have their labels directly initialized to [1||0]. Edges between unlabeled samples in the support set or related to samples in the query set have their labels set to [0.5||0.5], indicating that the relationship between the edges connecting the vertices is unknown.
[0089] The edge feature initialization expression in the model training step is as follows:
[0090]
[0091] Preferably, in step B6, the graph update module is provided with an L-layer graph neural network, each layer of the graph neural network is provided with a node feature update unit, a similarity calculation unit and an edge feature update unit, and each layer of the graph neural network performs one iteration operation of node feature update and edge feature update.
[0092] The graph update module performs iterative operations on node feature updates and edge feature updates on the initialized feature graph a5, with the following steps:
[0093] Step B61: The node feature update unit in the graph update module performs node feature update operation on the initial feature graph a5 to obtain the node updated feature graph, and passes it to the similarity calculation unit and the edge feature update unit.
[0094] The node feature update expression is as follows:
[0095]
[0096] in, Represents normalized edge features. This represents summing all edges connected to node i, where k represents the index of all other nodes directly connected to node i. Represents a network with changing nodes. The node changes represent network parameters, where l represents the l-th layer of the graph neural network, and 1 ≤ l ≤ L;
[0097] In each layer of the graph neural network, node features Updating by aggregating features from neighboring nodes takes edge features into account. As weights, they are similar to attention mechanisms.
[0098] Step B62: The similarity calculation unit calculates the similarity between every two nodes in the node update feature map through the metric network, and passes the calculated similarity to the edge feature update unit;
[0099] The similarity calculation expression between any two nodes is as follows:
[0100]
[0101] in, This represents the similarity of the intra-class relationship between nodes i and j. This represents the similarity of the inter-class relationship between nodes i and j. Represents the metric network, These represent the parameters of the metric network;
[0102] Step B63: The edge feature update unit performs edge feature update operation on the node update feature map according to the similarity, obtains the edge update feature map, and passes it to the next layer graph feature network for iterative operation;
[0103] The edge feature update expression is as follows:
[0104]
[0105] in, The feature vector of the edge between nodes i and j; This represents the unnormalized edge feature vector calculated based on node features; express The L1 norm of a vector is the sum of the absolute values of all its elements.
[0106] Step B64: When the number of iterations equals L, the edge feature update unit in the Lth layer graph neural network outputs the clustering feature map a6 to the classification and recognition module.
[0107] The above design utilizes the graph structure information of the data to explicitly perform clustering by iteratively updating node and edge features, thereby improving the performance of few-shot learning tasks.
[0108] Preferably, in step B7, the classification and identification module extracts the similarity data between the query set sample nodes and each support set sample node in the clustering feature map a6, calculates the average similarity between the query set sample and each type of support set sample, and then takes the gas category corresponding to the support set sample with the largest average similarity as the gas classification result.
[0109] Each gas category corresponds to multiple support set samples. The similarity data between two support set samples of the same gas category and the query set sample are not the same. Therefore, by calculating the average similarity between the query set sample and each type of support set sample, the gas category corresponding to the support set sample with the largest average similarity is taken as the gas classification result of the query set sample.
[0110] The above design effectively improves the accuracy of gas classification results.
[0111] The beneficial effects of this invention are as follows: This invention introduces the concept of contrastive learning into the CEGNNDrift model. This improvement not only significantly enhances model performance, but more importantly, it allows the model to be trained not only in supervised settings but also in unsupervised settings. This ability to train without supervision significantly reduces the model's dependence on a large number of labeled samples, thus providing a more efficient and practical solution in situations where data acquisition costs are high or labeled information is limited. Attached Figure Description
[0112] Figure 1 This is a schematic diagram of the process of the present invention;
[0113] Figure 2 This is a schematic diagram of the neural network training model structure in Example 1;
[0114] Figure 3 This is a schematic diagram of the graph neural network structure based on contrastive learning and edge labeling framework in Example 1;
[0115] Figure 4 The diagram below shows the structure of the update module in Example 1;
[0116] Figure 5 This is a schematic diagram of the momentum encoder structure in Example 1;
[0117] Figure 6 This is a schematic diagram of the computational steps of the spatial attention mechanism in Example 1;
[0118] Figure 7 This is a schematic diagram of the Senet architecture in Example 1;
[0119] Figure 8 The accuracy of each algorithm in Example 2 when the categories are symmetric;
[0120] Figure 9 This is a graph showing the test accuracy of each algorithm in Example 2 when the categories are asymmetric. Detailed Implementation
[0121] The present invention will be further described in detail below with reference to the accompanying drawings and specific examples. The following embodiments or drawings are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0122] Example 1:
[0123] like Figure 1-3 As shown: A graph neural network gas classification and recognition method based on contrastive learning and edge labeling framework, including model training step A and gas classification and recognition step B;
[0124] The model training step A includes the following steps:
[0125] Step A1: Construct a graph neural network training model, which includes a feature enhancement module, an encoder, a momentum encoder, a graph construction module, and a graph update module.
[0126] The feature enhancement module is connected to the encoder and the momentum encoder, the encoder and the momentum encoder are connected to the graph construction module, and the graph construction module is connected to the graph update module;
[0127] Step A2: The graph neural network training model is trained using a contrastive learning strategy under self-supervised learning. After training is completed, the momentum encoder, graph construction module and graph update module are retained. The momentum encoder is connected to the graph construction module and the graph construction module is connected to the graph update module, thereby obtaining a graph neural network based on contrastive learning and edge labeling framework.
[0128] The gas classification and identification step B includes the following steps:
[0129] Step B1: The gas sensor array acquires gas characteristic data a1 in real time and transmits the gas characteristic data a1 to the preprocessing module;
[0130] Step B2: The preprocessing module performs preprocessing operations on the gas feature data a1 to obtain query set sample a2, and then passes it to the graph neural network;
[0131] Step B3: The input layer of the graph neural network inputs the query set sample a2 and N support set samples into the momentum encoder, where each support set sample contains R samples;
[0132] Step B4: The momentum encoder performs feature extraction operations on the query set sample a2 and the N-class support set samples respectively to obtain the query set feature vector a3 and the N-class support set feature vector, and then passes them to the graph construction module;
[0133] Step B5: The graph construction module constructs a fully connected graph a4 from the query set feature vector a3 and the N-class support set feature vectors, and initializes the node features and edge features in the fully connected graph a4 to obtain an initialized feature graph a5, which is then passed to the graph update module.
[0134] Step B6: The graph update module performs iterative operations on node feature update and edge feature update on the initial feature graph a5. After L iterations, it outputs the clustering feature graph a6 to the classification and recognition module.
[0135] Step B7: The classification and recognition module classifies and recognizes the cluster feature map a6 and outputs the gas classification result.
[0136] In step A2, the graph neural network training model is trained using a contrastive learning strategy under self-supervised learning. The specific steps are as follows:
[0137] Step A21: The feature enhancement module acquires m unlabeled gas samples, and the first enhancement unit in the feature enhancement module performs feature enhancement operation on the unlabeled gas samples and outputs the first enhanced sample set to the encoder;
[0138] The second enhancement unit in the feature enhancement module performs feature enhancement operations on the unlabeled gas sample set and outputs the second enhanced sample set to the momentum encoder.
[0139] Step A22: The encoder performs feature extraction on the first enhanced sample set to obtain a first feature vector set, and passes it to the graph construction module;
[0140] The momentum encoder performs feature extraction on the second enhanced sample set to obtain a second feature vector set, which is then passed to the graph construction module.
[0141] Step A23: The graph construction module constructs the first feature vector set and the second feature vector set into a fully connected graph, initializes the node features and edge features in the fully connected graph, and then passes the initialized feature graph to the graph update module;
[0142] Step A24: The graph update module performs iterative operations on node feature update and edge feature update on the initialized feature map. After L iterations, it outputs the clustering feature map to the classification and recognition module.
[0143] Step A25: The classification and recognition module classifies and recognizes the clustering feature map, and calculates the model loss based on the classification and recognition results;
[0144] Step A26: Update the encoder parameters and graph update module parameters according to the model loss, and then update the momentum encoder parameters according to the encoder parameters;
[0145] Step A27: Repeat steps A21-A26 to obtain the optimal graph neural network training model. Retain the momentum encoder, graph construction module, and graph update module, as well as the optimal parameters of each module, to obtain a graph neural network based on contrastive learning and edge labeling framework.
[0146] In step A21, the first enhancement unit performs feature enhancement on the unlabeled gas sample using a feature zeroing enhancement method. The feature zeroing enhancement method is used for a given sample x. i , where each feature x i,j The probability of p1 is set to 0, as shown in the following expression:
[0147]
[0148] Where, x′ i,j Indicates sample x i After the enhanced j-th feature, x i,j Indicates sample x i The j-th feature;
[0149] The second enhancement unit uses an additive Gaussian perturbation enhancement method to perform feature enhancement on the unlabeled gas sample. The additive Gaussian perturbation enhancement method is used for a given sample x. i Each eigenvalue x i,j It will be multiplied by a number drawn from a Gaussian distribution with a mean of 1 and a variance of p², as shown in the following expression:
[0150] x′ i,j =x i,j ·(1+N(0, p2))
[0151] Where N(0, p2) represents a random number drawn from a normal distribution with mean 0 and variance p2.
[0152] The graph neural network training model provided by this invention allows the model to be trained not only under supervised settings but also under unsupervised settings.
[0153] The supervised training process can be described as follows:
[0154] (1) During training, a batch of samples is collected from the source domain in each round based on the N-way K-shot paradigm for training. After class expansion, each sample x i Two augmented samples were obtained by performing data augmentation in different ways. and
[0155] (2) Enhance the sample and Input encoder E1 and momentum encoder E2 respectively to obtain the corresponding feature vectors. and
[0156] (3) Initialize vertices and edges. Set any edge between two different nodes as an unknown edge. Input the graph neural network, predict the unknown edges, and calculate the model loss. Note that when calculating the loss, samples augmented from similar samples should be considered as belonging to the same class, and vice versa.
[0157] (4) Update encoder E1 and graph neural network using model loss, and update momentum encoder E2 using the parameters of encoder E1.
[0158] Its unsupervised training process can be described as follows:
[0159] (1) During training, a batch of samples is collected from the source domain in each round for training. This batch of samples does not need to be expanded into classes. Each sample x i Two augmented samples were obtained by performing data augmentation in different ways. and
[0160] (2) Enhance the sample and Input encoder E1 and momentum encoder E2 respectively to obtain the corresponding feature vectors. and
[0161] (3) Initialize vertices and edges. Set any edge between two different nodes as an unknown edge. Input the graph neural network, predict the unknown edges, and calculate the model loss. Note that when calculating the loss, samples augmented from the same sample should be considered as the same class, and vice versa.
[0162] (4) Update encoder E1 and graph neural network using model loss, and update momentum encoder E2 using the parameters of encoder E1.
[0163] The model training process based on contrastive learning is as follows:
[0164] All samples x i New samples were obtained through data augmentation methods. and Then, the new samples are processed by encoder E1 and momentum encoder E2 to obtain feature vectors. and
[0165] When using supervised training, for edge labeling, due to the use of a contrastive learning strategy, the labels are set from samples of the same class x. i The data augmented samples have edges with a label of 1 between corresponding nodes, and edges with a label of 0 for all other edges.
[0166] When training using unsupervised learning, the same sample x is used. i The data augmented samples have edges with a label of 1 between corresponding nodes, and edges with a label of 0 for all other edges.
[0167] For edge initialization, regardless of whether supervised learning is used for training, the relationships between edges are not visible except for the labels between nodes themselves, which are set to [1||0], and are set to [0.5||0.5].
[0168] In step A22, the encoder and the momentum encoder have the same structure and the same initial parameters;
[0169] In step A26, the encoder parameters are fully updated based on the model loss, and the momentum encoder parameters are partially updated based on the encoder parameters.
[0170] like Figure 5 As shown: In step B4, the momentum encoder performs feature extraction on the query set sample a2 or the support set sample, and the steps are as follows:
[0171] Step B41: The spatial attention layer in the momentum encoder acquires the query set sample a2 or the support set sample, performs spatial feature extraction on it to obtain spatial feature data s1, and passes it to the first batch of normalization layers.
[0172] Step B42: The first batch of normalization layers performs batch normalization on the spatial feature data s1 to obtain the first batch of normalized data s2, and then passes it to the first residual block;
[0173] Step B43: The first residual block performs a residual convolution operation on the first batch of normalized data s2 to obtain the first residual data s3, and then passes it to the second batch of normalization layers;
[0174] Step B44: The second batch normalization layer performs batch normalization on the first residual data s3 to obtain the second batch normalized data s4, and then passes it to the second residual block;
[0175] Step B45: The second residual block performs a residual convolution operation on the second batch of normalized data s4 to obtain the query set feature vector c or the support set feature vector, and passes it to the encoding unit.
[0176] like Figure 6 As shown: In step B41, the spatial attention layer performs spatial feature extraction on the query set sample a2 or the support set sample, and the steps are as follows:
[0177] Step B411: Obtain a one-dimensional feature map: Assume the input data of the spatial attention layer is X∈R C×W Where C and W represent the number of channels and channel width of the input data, respectively; by performing average pooling and max pooling on the input data X, the average pooling feature map and max pooling feature map are obtained, as expressed below:
[0178] F avg ∈R 1×W
[0179] F max ∈R 1×W
[0180] Among them, F avg F represents the average value of the average pooling feature map along the channel dimension. max This represents the maximum value of the max-pooled feature map in the channel dimension;
[0181] Step B412: Feature map concatenation: The average pooling feature map F avg and max pooling feature map F max The feature maps F are concatenated along the channel dimension. concat ∈R 2×W ;
[0182] Step B413: Obtain the spatial attention map: for the concatenated feature map F concat Perform a one-dimensional convolution operation and obtain the spatial attention map M through the sigmoid activation function. s ∈R 2×W The expression is as follows:
[0183] M s =σ(Conv(F) concat ))
[0184] Where σ represents the Sigmoid function, and Conv represents one-dimensional convolution. The expression for one-dimensional convolution is as follows:
[0185]
[0186] Among them, y t It outputs the value of the feature vector at position t, x i w is the value of the input feature vector at position i. i,t-i b is the weight of the convolution kernel at position i. t It is a bias term;
[0187] Step B414: Attention Weighting: Combine the original input feature map X with the spatial attention map M s Multiply by the weighted output feature map X′:
[0188]
[0189] in, This represents element-level multiplication operations.
[0190] When performing unsupervised learning, since category expansion strategies are not applicable, channel attention mechanisms (CAM) can be used on the unscrambled dimensions. Specifically, SENet (Squeeze-and-Excitation Networks) can be added to model the relationships between sample channels.
[0191] SEnet is a method for implementing channel attention mechanism in deep convolutional neural networks (CNNs). Its core idea is to improve the representation ability of the network by explicitly modeling the interdependencies between convolutional feature channels. Its architecture is shown in Figure 7.
[0192] Its basic idea is as follows:
[0193] (1) Compression: This step aggregates the values of all elements in each channel into a single value using global average pooling. This allows us to obtain the overall statistics for each channel without losing information within each channel. The formula is as follows:
[0194]
[0195] Where x represents the initial sample, z represents the compressed sample x, dim_size represents the length of each dimension of the sample, i represents the i-th sample, and F sq This indicates a compression operation.
[0196] (2) Extraction: The compressed vector is processed through two fully connected layers to generate channel weights. To improve computational efficiency, a reduction factor ratio is set, reducing the number of neurons in the first layer to half of the original value. ReLU is used as the non-linear function. The second layer retains the same number of neurons as the input and uses the sigmoid function to constrain the weights between 0 and 1. These fully connected layers are parameterized by parameters w′1 and w′2.
[0197] s = F ex (z)=g(f(x,{w′1},{w′2}))
[0198] Where s represents the extracted sample from compressed sample z, g() and f() represent two fully connected layer operations, w′1 and w′2 represent the fully connected layer parameters, and F ex This indicates an extraction operation.
[0199] (3) Reweighting: The importance score for each channel is derived from the Extract stage, and then the channels are reweighted using this score. This process involves multiplying each channel by its corresponding importance score to produce finely calibrated and adjusted attention channels.
[0200]
[0201] in, F represents the final sample matrix, whose shape is the same as the initial sample x; sc This indicates a weighting operation.
[0202] SEnet explicitly models the interdependencies between channels by introducing a channel attention mechanism, and improves the network's representational power and performance through the utilization of global information and adaptive recalibration. Although this model is commonly used for tasks such as image recognition, we transfer it to our model to improve performance due to the similarity between gas samples and image samples.
[0203] like Figure 5 As shown: The first residual block and the second residual block have the same structure, both of which are provided with parallel connected weight layers and identity layers. The weight layer consists of 5 one-dimensional convolutional layers connected end to end, and the identity layer consists of 1 one-dimensional convolutional layer connected by residuals.
[0204] The expression for the first residual block or the second residual block is as follows:
[0205] x n+1 =F(x) n , {w n,m})+Conv n (x n , {w′ n,m})
[0206] Where, x n The input feature is x. n+1 It is the output feature, where n represents the index of the residual layer; F(x) n , {w n,m}) is the residual function of the weight layer, representing the input x. n The result after weighted layer transformation; Conv n (x n , {w′ n,m}) represents the input x n The result after identity layer transformation; {w n,m} is the weight set of the nth residual block, w′ n,m These are the parameters of the identity layer.
[0207] like Figure 2 , Figure 4 As shown: In step B5, each node in the fully connected graph a4 represents a feature vector, and each edge represents the relationship type between two nodes; the node feature initialization expression is as follows:
[0208]
[0209] in, Let E(x) represent the feature of the i-th initial node. i ) represents the i-th input sample x i Extracted feature vectors;
[0210] The edge feature initialization expression is as follows:
[0211]
[0212] Among them, y i Let y represent the label of the i-th node. j Let y represent the label of the j-th node. ij The edge feature label between nodes i and j Represents the initial edge features between nodes i and j, each edge feature It is a two-dimensional vector, e when d=1. ij1 This represents the strength of the inter-class relationship between two connected nodes; when d = 2, e ij2 || represents the strength of the intra-class relationship between two connected nodes; || is the combination of intra-class and inter-class relationships; label represents the label, and NULL indicates that the sample does not contain a label.
[0213] In step B6, the graph update module is equipped with an L-layer graph neural network. Each layer of the graph neural network is equipped with a node feature update unit, a similarity calculation unit, and an edge feature update unit. Each layer of the graph neural network performs one iteration of node feature update and edge feature update.
[0214] The graph update module performs iterative operations on node feature updates and edge feature updates on the initialized feature graph a5, with the following steps:
[0215] Step B61: The node feature update unit in the graph update module performs node feature update operation on the initial feature graph a5 to obtain the node updated feature graph, and passes it to the similarity calculation unit and the edge feature update unit.
[0216] The node feature update expression is as follows:
[0217]
[0218] in, Represents normalized edge features. This represents summing all edges connected to node i, where k represents the index of all other nodes directly connected to node i. Represents a network with changing nodes. The node changes represent network parameters, where l represents the l-th layer of the graph neural network, and 1 ≤ l ≤ L;
[0219] Step B62: The similarity calculation unit calculates the similarity between every two nodes in the node update feature map through the metric network, and passes the calculated similarity to the edge feature update unit;
[0220] The similarity calculation expression between any two nodes is as follows:
[0221]
[0222] in, This represents the similarity of the intra-class relationship between nodes i and j. This represents the similarity of the inter-class relationship between nodes i and j. Represents the metric network, These represent the parameters of the metric network;
[0223] Step B63: The edge feature update unit performs edge feature update operation on the node update feature map according to the similarity, obtains the edge update feature map, and passes it to the next layer graph feature network for iterative operation;
[0224] The edge feature update expression is as follows:
[0225]
[0226] in, The feature vector of the edge between nodes i and j; This represents the unnormalized edge feature vector calculated based on node features; express The L1 norm of a vector is the sum of the absolute values of all its elements.
[0227] Step B64: When the number of iterations equals L, the edge feature update unit in the Lth layer graph neural network outputs the clustering feature map a6 to the classification and recognition module.
[0228] In step B7, the classification and recognition module extracts the similarity data between the query set sample nodes and each support set sample node in the clustering feature map a6, calculates the average similarity between the query set sample and each support set sample, and then takes the gas category corresponding to the support set sample with the largest average similarity as the gas classification result.
[0229] Example 2:
[0230] Based on Example 1, Example 2 validates the graph neural network gas classification and recognition method based on the node labeling framework using the "Gas Sensor Array Drift Dataset" from the UCI Machine Learning Library. This dataset comprises 13,910 measurements collected over 36 months, using an advanced electronic nose system equipped with 16 robust metal-oxide-semiconductor gas sensors. These sensors are divided into four mini-sensor arrays, each containing four different models of Figaro sensors: TGS2600, TGS2602, TGS2610, and TGS2620. The primary objective is to detect and classify six volatile organic compounds: ammonia (NH3), acetaldehyde (C2H4O), acetone (C3H6O), ethylene (C2H4), ethanol (C2H5OH), and toluene (C7H8). All data are divided into 10 batches according to the collection order, and the collection period and number of measurements for each batch of data are shown in Table 1.
[0231] Table 1. Collection period and number of measurements for each batch of data.
[0232]
[0233] To construct multiple classification tasks, Example 2 employed two experimental configurations to verify the algorithm performance of CEGNNDrift:
[0234] Experimental Setup 1: The first batch was used as the source domain dataset for training, and subsequent batches were used as the target domain dataset for performance evaluation, thus forming a series of 9 independent classification tasks. This setup was primarily designed to investigate the model's performance under short-term and long-term drift.
[0235] Experimental Setup 2: A sequential training strategy was adopted. For each batch T numbered from 2 to 10, the previous batch, T-1, was used as the source domain dataset for training, and the next batch, T, was used as the target domain dataset for testing. This setup also generated 9 corresponding classification tasks. This setup was mainly to study the model's performance during short-term drift.
[0236] Using these two experimental setups, we can more comprehensively test the model's performance in terms of long-term drift (long time interval between source and target domain dataset collection) and short-term drift (short time interval between source and target domain dataset collection).
[0237] The input momentum encoder E has a matrix shape of [16,8]. The momentum encoder transforms this matrix into a matrix of shape [64,2], and then flattens it into a vector of shape
[128] . This vector contains the deep features extracted by the momentum encoder, providing rich information for subsequent graph neural network processing.
[0238] CEGNNDrift uses the binary cross-entropy loss BCELoss as its loss function, and its calculation formula is as follows:
[0239] BCELoss is particularly suitable for binary classification, that is, when the output has only two categories, such as 0 and 1, which usually represent "not" and "yes". The model output is a probability value, representing the probability of an event occurring. BCELoss calculates the cross-entropy between the model output probability and the true label (0 or 1).
[0240] The formula for the cross-entropy loss function is defined as follows:
[0241]
[0242] Where N is the total number of samples. i p is the true label of the i-th sample, either 0 or 1; i is the probability that the model predicts the i-th sample to be of class 1, and log is the natural logarithm.
[0243] CEGNNDrift uses Adam as its optimizer, a gradient descent optimization algorithm for deep learning that combines the advantages of momentum and RMSprop optimization algorithms. Given parameters θ, the gradient of t is g. t The Adam optimizer update rules are as follows:
[0244] (1) Calculate the first-order moment momentum estimate:
[0245] m t =β1m t-1 +(1-β1)g t
[0246] Where, m t β1 is the first moment at time t; β2 is the decay rate of the first moment, usually set to 0.9.
[0247] (2) Calculate the estimate of the variance of the second moment gradient:
[0248]
[0249] Among them, v t It is the first t The second moment at time t; β2 is the decay rate of the second moment, usually set to 0.999.
[0250] (3) Calculate the adjustment values for the first and second moments:
[0251]
[0252] (4) Update parameters:
[0253]
[0254] Where, θ t+1 α is the learning rate; ε is a very small constant used to prevent the denominator from being zero, usually set to 1e-8.
[0255] The Adam optimizer is designed to adaptively adjust the learning rate of each parameter, thereby improving training efficiency and model performance. Subsequent graph neural network training models also use this optimizer for parameter updates.
[0256] Some important parameters of graph neural networks are shown in Table 1.
[0257] This model uses accuracy as the evaluation metric, which is the most common evaluation metric in classification tasks, representing the proportion of samples that the model correctly predicts.
[0258] N-way K-shot is a common learning paradigm in few-shot learning, used to partition datasets for training or testing tasks, where N is the number of classes in the task and K is the number of samples in each class. Since this model is trained unsupervised, the training samples are randomly selected rather than partitioned using the N-way K-shot paradigm, while the testing samples are obtained using this paradigm.
[0259] Table 1
[0260]
[0261] Next, these seven models are used as comparison models, compared with GNNDrift, EGNNDrift, CEGNNDrift(S), and CEGNNDrift(U). GNNDrift, EGNNDrift, and CEGNNDrift(S) represent graph neural networks trained using supervised learning with a node-labeling framework, a graph neural network trained using an edge-labeling framework, and a graph neural network trained using contrastive learning and an edge-labeling framework, respectively, when each class has one labeled sample for reference (K=1). CEGNNDrift(U) represents a graph neural network trained using unsupervised learning when each class has five labeled samples for reference (K=5). Both EGNNDrift and CEGNNDrift use a transmission learning mechanism, while EGNNDrift does not use unlabeled samples to assist training.
[0262] The seven models are: Support Vector Machine (SVM), Principal Component Analysis (PCA), Multi-Constraint Subspace Projection (MCSP), Linear Discriminant Analysis (LDA), Domain Regularized Component Analysis (DRCA), Domain-Based Subspace Learning (DMDMR) which considers maximizing label feature dependency and minimizing feature redundancy, and Domain Regularized Component Analysis (DDRCA).
[0263] For Experiment 1 and Experiment 2, 11 networks were tested respectively to demonstrate the superiority of CEGNNDrift in long-range and short-range drift problems. In this case, the source and target domains have the same class, i.e., class symmetry. The test results were plotted as a heatmap as shown below. Figure 8 As shown, “ab” indicates that batch a is the source domain and batch b is the target domain, Avg represents the average accuracy of the model on each task, and U and S represent the CEGNNDrift model trained in unsupervised and supervised learning methods, respectively.
[0264] like Figure 8 As shown, the following conclusions can be drawn:
[0265] After evaluating the performance of short-term and long-term drift, it was found that all models showed good performance in handling short-term drift, but performance generally declined in the case of long-term drift. This decline was particularly pronounced in the 1-10 experimental settings. The model with the highest average accuracy (Avg) was CEGNNDrift(S), reaching 84.96%, demonstrating the best performance in almost all experiments. This was followed by EGNNDrift and EGNNDrift(U), reaching 80.54% and 80.33%, respectively. Next were DDRCA, MCSP, and GNNDrift, reaching 73.89%, 70.55%, and 68.03%, respectively. SVM, PCA, LDA, DRCA, and DMDMR performed poorly, with average accuracies ranging from 40% to 70%. It can be seen that our three proposed models, especially CEGNNDrift, exhibited high stability and accuracy in this experimental setting, regardless of whether the conditions were supervised or unsupervised.
[0266] Furthermore, it can be observed that different experimental settings have varying effects on the models. For example, most models perform well in experimental settings 1-4, 1-5, 1-7, and 1-8, but perform poorly in experimental settings 1-6, 1-9, and 1-10. This is likely due to the characteristics of these experimental settings affecting the models.
[0267] In most experimental settings, the CEGNNDrift model achieved the highest average accuracy, and because it was trained without labeled samples, its applicability is wider compared to other models. The reasons for its high accuracy can be summarized as follows:
[0268] ① Contrastive Learning Strategy: By introducing contrastive learning, the CEGNNDrift model can be trained in an unsupervised setting. Contrastive learning provides training signals by distinguishing between positive sample pairs (similar views) and negative sample pairs (dissimilar views), which helps the model learn to distinguish subtle differences between different samples.
[0269] ② Proxy task individual discrimination: Individual discrimination, as a proxy task, prompts the model to treat each training data as a unique category, which enhances the model's ability to identify sample features and ensures that the feature vectors of samples with consistent categories are closer.
[0270] ③ Data Augmentation Strategies: CEGNNDrift employs two data augmentation strategies—feature zeroing augmentation and additive Gaussian perturbation augmentation. These strategies increase the diversity of samples, forcing the model to learn more robust and generalized feature representations.
[0271] ④ Adjustment of the graph neural network: Based on the data augmentation strategy, CEGNNDrift adjusted the graph neural network by setting edge labels to distinguish samples augmented by the same feature vector from other samples, further improving the model's ability to identify samples of the same type.
[0272] ⑤ Iterative Update Mechanism: During the training of the CEGNNDrift model, the feature vectors and edge labels are iteratively updated. This mechanism allows the model to gradually extract signals useful for sample classification in each iteration. The slow update of the momentum encoder enables the model to update the encoder parameters stably and consistently, thereby generating more stable contrastive learning targets to improve the representation learning effect.
[0273] ⑥ Generalization ability: When using unsupervised learning, CEGNNDrift reduces its dependence on a large number of labeled samples, thus providing a more efficient and practical solution when data acquisition costs are high or label information is limited.
[0274] (1) The impact of class asymmetry on model performance:
[0275] The class asymmetry problem refers to the situation where the sample classes in the training set and the test set are inconsistent. In this case, the source domain dataset contains five gas samples: ammonia, acetaldehyde, acetone, ethylene, and ethanol, while the target domain dataset contains six gases: ammonia, acetaldehyde, acetone, ethylene, ethanol, and toluene.
[0276] In addition to the seven comparative models mentioned above, a sensor drift compensation model specifically designed for class asymmetry—UDA-CA—is also used. The UDA-CA model is an unsupervised domain adaptation method that aims to reduce the distributional differences between the source domain (labeled dataset) and the target domain (unlabeled dataset) through class-aware projection. This method is particularly suitable for situations where there are significant distributional differences between the source and target domains, while maintaining class separability within the source domain. Its core idea is to map the data from the source and target domains to a common feature space using a projection matrix W, while minimizing the distributional differences between the source and target domains and preserving class cohesion within the source domain.
[0277] When K=1, the accuracy of the three models EGNNDrift, CEGNNDrift(S), and CEGNNDrift(U) in the asymmetric case was verified, and eight comparison models were used for comparison. The test results were plotted as a heatmap, as shown below. Figure 9 As shown.
[0278] Experimental results consistently demonstrate that both EGNNDrift and CEGNNDrift models exhibit high accuracy in cases of class asymmetry, showcasing their strong adaptability and robustness to class inconsistencies between the training and test sets. This finding confirms the potential of these models in practical applications, particularly for classification tasks involving unknown classes.
[0279] Specifically, under supervised training, CEGNNDrift(S) outperforms EGNNDrift on most tasks. The EGNNDrift model has already proven its effectiveness in handling class asymmetry problems, while the CEGNNDrift model further enhances its performance. Unsupervised training of CEGNNDrift(U) also demonstrates excellent performance.
[0280] (2) Investigating the effect of K value on model efficiency:
[0281] In few-shot learning, the value of K directly affects the difficulty of the model's learning task. A larger K value means that there are more samples available for training in each class, which usually helps the model learn the features of each class better, thus improving classification accuracy. Conversely, when the K value is small, meaning there are only a few samples available for training in each class, the features learned by the model may be insufficient to distinguish between classes, leading to a decrease in performance.
[0282] Since introducing transfer learning mechanisms can have an effect similar to increasing the K value, as both lead to the model processing more samples simultaneously, we chose not to use transfer learning mechanisms in experiments investigating the impact of the K value on model performance to avoid the combined effects of both interfering with the experimental results.
[0283] For cases where K is 1, 2, 3, 4, and 5, the performance of the three CEGNNDrift models on Experiment 1 is tested. For the CEGNNDrift model, labeled data is not required for training, but labeled samples in the target domain are needed as a benchmark to determine the probability that an unknown class sample belongs to that class.
[0284] The performance of CEGNNDrift under supervised learning is shown in Table 2.
[0285] Table 2
[0286]
[0287] The performance of CEGNNDrift in unsupervised learning is shown in Table 3.
[0288] Table 3
[0289]
[0290] For CEGNNDrft(S) and CEGNNDrft(U), both models exhibit minimum efficiency when K=1. As K increases, model efficiency improves, but the rate of increase slows down as N increases. With K=1, meaning only one sample per class is used for training, all models are inefficient. This finding is not surprising, as in few-shot learning, models struggle to learn generalizable feature representations due to insufficient data to capture the diversity and complexity of the data. However, as the value of K increases, i.e., the number of training samples increases, the efficiency of all models improves. This trend suggests that more training data provides the model with richer information, enabling it to better understand and simulate the intrinsic patterns of sensor drift. This improvement is particularly significant in the early stages, demonstrating that even a small increase in samples can significantly improve model performance in few-shot learning.
[0291] As the value of K increases further, especially when K is large, the improvement in model efficiency begins to plateau. This may mean that the model has approached its performance ceiling or has been affected by overfitting to some extent. Furthermore, this may also indicate that there is an optimal number of training samples in few-shot learning, beyond which adding more samples contributes little to improving model performance.
[0292] (3) The impact of the transfer learning mechanism on CEGNNDrift
[0293] To thoroughly evaluate the role of the transduction learning mechanism in the CEGNNDrift model and its specific impact on model performance, this embodiment designs a series of comparative experiments. In these experiments, the accuracy of the model's performance with and without the transduction mechanism is directly compared. Since semi-supervised learning and transduction learning mechanisms have similar effects—both increasing model performance through unlabeled samples—supervised learning is used to train CEGNNDrift instead of semi-supervised learning.
[0294] This approach aims to reveal whether the transfer learning mechanism can significantly improve the classification efficiency of a model in few-shot learning tasks. Experimental results will effectively explain how the transfer learning mechanism affects the model's ability to capture similarities and differences between samples under different training and testing conditions, and the specific contribution of this mechanism to the accuracy and robustness of the final classification decision.
[0295] Furthermore, these comparative experiments will also verify the adaptability and effectiveness of the CEGNNDrift model for sensor drift compensation tasks under different configurations, thus providing experimental basis for model selection and optimization in practical applications.
[0296] For CEGNNDrift(S), K=1 was set, and for CEGNNDrift(U), K=5 was set. The experimental results are shown in Table 4.
[0297] Table 4
[0298]
[0299] The introduction of the transfer learning mechanism significantly improved the classification efficiency of the three models. Experimental results clearly demonstrate the importance of transfer learning in few-shot learning tasks, especially when there is only one support sample in each class. This mechanism enhances the model's understanding of the global structure of the data by considering the relationships between all test samples in the graph.
[0300] (4) The impact of category expansion strategies
[0301] To improve the class diversity of the source domain data, a dimensionality augmentation strategy was used for CEGNNDrift(S). This strategy increases the number of sample classes by randomly shuffling sensor channels, allowing the model to access a wider range of data types and thus improving its generalization ability. To verify the effectiveness of this strategy, the performance of the CEGNNDrift(S) model without this strategy was compared with that of the model using this strategy. The case of only one support sample in each class (i.e., K=1) was selected, and experiment setting 1 was used to conduct the experiment. EGNNDrift does not use unlabeled samples to assist training and testing. The experimental results are shown in Table 5. It should be noted that CEGNNDrift(S) uses a passivation learning strategy.
[0302] The category expansion strategy played a crucial role in improving the performance of both models. Experimental results show that the model using this strategy achieved a significant improvement in accuracy, indicating that by increasing the number of categories, the model can better generalize to unseen categories, thus exhibiting better adaptability and robustness in practical applications.
[0303] Table 5
[0304]
[0305] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A graph neural network gas classification and recognition method based on contrastive learning and edge labeling framework, characterized in that, This includes model training step A and gas classification and recognition step B; The model training step A includes the following steps: Step A1: Construct a graph neural network training model, which includes a feature enhancement module, an encoder, a momentum encoder, a graph construction module, and a graph update module. The feature enhancement module is connected to the encoder and the momentum encoder, the encoder and the momentum encoder are connected to the graph construction module, and the graph construction module is connected to the graph update module; The encoder performs feature extraction on the first enhanced sample set to obtain a first feature vector set, which is then passed to the graph construction module. The momentum encoder performs feature extraction on the second enhanced sample set to obtain a second feature vector set, which is then passed to the graph construction module. The graph construction module constructs the first feature vector set and the second feature vector set into a fully connected graph, initializes the node features and edge features in the fully connected graph, and then passes the initialized feature graph to the graph update module. The encoder parameters are fully updated based on the model loss, and the momentum encoder parameters are partially updated based on the encoder parameters. Step A2: The graph neural network training model is trained using a contrastive learning strategy under self-supervised learning. After training is completed, the momentum encoder, graph construction module and graph update module are retained. The momentum encoder is connected to the graph construction module and the graph construction module is connected to the graph update module, thereby obtaining a graph neural network based on contrastive learning and edge labeling framework. The gas classification and identification step B includes the following steps: Step B1: The gas sensor array acquires gas characteristic data a1 in real time and transmits the gas characteristic data a1 to the preprocessing module; Step B2: The preprocessing module performs preprocessing operations on the gas feature data a1 to obtain query set sample a2, and then passes it to the graph neural network; Step B3: The input layer of the graph neural network inputs the query set sample a2 and N support set samples into the momentum encoder, where each support set sample contains R samples; Step B4: The momentum encoder performs feature extraction operations on the query set sample a2 and the N-class support set samples respectively to obtain the query set feature vector a3 and the N-class support set feature vector, and then passes them to the graph construction module; Step B5: The graph construction module constructs a fully connected graph a4 from the query set feature vector a3 and the N-class support set feature vectors, and initializes the node features and edge features in the fully connected graph a4 to obtain an initialized feature graph a5, which is then passed to the graph update module. Step B6: The graph update module performs iterative operations on node feature update and edge feature update on the initial feature graph a5. After L iterations, it outputs the clustering feature graph a6 to the classification and recognition module. Step B7: The classification and recognition module classifies and recognizes the cluster feature map a6 and outputs the gas classification result.
2. The graph neural network gas classification and recognition method based on contrastive learning and edge labeling framework according to claim 1, characterized in that: In step A2, the graph neural network training model is trained using a contrastive learning strategy under self-supervised learning. The specific steps are as follows: Step A21: The feature enhancement module acquires m unlabeled gas samples, and the first enhancement unit in the feature enhancement module performs feature enhancement operation on the unlabeled gas samples and outputs the first enhanced sample set to the encoder; The second enhancement unit in the feature enhancement module performs feature enhancement operations on the unlabeled gas sample set and outputs the second enhanced sample set to the momentum encoder. Step A22: The encoder performs feature extraction on the first enhanced sample set to obtain a first feature vector set, and passes it to the graph construction module; The momentum encoder performs feature extraction on the second enhanced sample set to obtain a second feature vector set, which is then passed to the graph construction module. Step A23: The graph construction module constructs the first feature vector set and the second feature vector set into a fully connected graph, initializes the node features and edge features in the fully connected graph, and then passes the initialized feature graph to the graph update module; Step A24: The graph update module performs iterative operations on node feature update and edge feature update on the initialized feature map. After L iterations, it outputs the clustering feature map to the classification and recognition module. Step A25: The classification and recognition module classifies and recognizes the clustering feature map, and calculates the model loss based on the classification and recognition results; Step A26: Update the encoder parameters and graph update module parameters according to the model loss, and then update the momentum encoder parameters according to the encoder parameters; Step A27: Repeat steps A21-A26 to obtain the optimal graph neural network training model. Retain the momentum encoder, graph construction module, and graph update module, as well as the optimal parameters of each module, to obtain a graph neural network based on contrastive learning and edge labeling framework.
3. The graph neural network gas classification and recognition method based on contrastive learning and edge labeling framework according to claim 2, characterized in that: In step A21, the first enhancement unit performs feature enhancement on the unlabeled gas sample using a feature zeroing enhancement method. The feature zeroing enhancement method is used for a given sample... Each feature The probability of p1 is set to 0, as shown in the following expression: ; in, Indicates sample After enhancement, the j-th feature Indicates sample The j-th feature; The second enhancement unit employs an additive Gaussian perturbation enhancement method to perform feature enhancement operations on the unlabeled gas sample. This additive Gaussian perturbation enhancement method is used for a given sample... Each eigenvalue It will be multiplied by a factor with a mean of 1 and a variance of 1. The numbers drawn from the Gaussian distribution are expressed as follows: ; in, This indicates that the mean is 0 and the variance is... A random number drawn from a normal distribution.
4. The graph neural network gas classification and recognition method based on contrastive learning and edge labeling framework according to claim 2, characterized in that: In step A22, the encoder and the momentum encoder have the same structure and the same initial parameters; In step A26, the encoder parameters are fully updated based on the model loss, and the momentum encoder parameters are partially updated based on the encoder parameters. The momentum encoder parameter update expression is as follows: ; in, This represents the momentum encoder parameters updated during the (t+1)th training iteration. This represents the momentum encoder parameters updated during the t-th training iteration. This represents the encoder parameters updated during the (t+1)th training iteration. Let represent the encoder parameters updated during the t-th training iteration, and u represent the momentum coefficient.
5. The graph neural network gas classification and recognition method based on contrastive learning and edge labeling framework according to claim 1, characterized in that: In step B4, the momentum encoder performs feature extraction on the query set sample a2 or the support set sample, as follows: Step B41: The spatial attention layer in the momentum encoder acquires the query set sample a2 or the support set sample, performs spatial feature extraction on it to obtain spatial feature data s1, and passes it to the first batch of normalization layers. Step B42: The first batch of normalization layers performs batch normalization on the spatial feature data s1 to obtain the first batch of normalized data s2, and then passes it to the first residual block; Step B43: The first residual block performs a residual convolution operation on the first batch of normalized data s2 to obtain the first residual data s3, and then passes it to the second batch of normalization layers; Step B44: The second batch normalization layer performs batch normalization on the first residual data s3 to obtain the second batch normalized data s4, and then passes it to the second residual block; Step B45: The second residual block performs a residual convolution operation on the second batch of normalized data s4 to obtain the query set feature vector c or the support set feature vector, and passes it to the encoding unit.
6. The graph neural network gas classification and recognition method based on contrastive learning and edge labeling framework according to claim 5, characterized in that: In step B41, the spatial attention layer performs spatial feature extraction on the query set sample a2 or the support set sample, as follows: Step B411: Obtain a one-dimensional feature map: Assuming the input data of the spatial attention layer Where C and W represent the number of channels and channel width of the input data, respectively; by performing average pooling and max pooling on the input data X, the average pooling feature map and max pooling feature map are obtained, as expressed below: ; ; in, This represents the average value of the average pooling feature map along the channel dimension. This represents the maximum value of the max-pooled feature map in the channel dimension; Step B412: Feature map concatenation: The average pooling feature map is concatenated... and max pooling feature map The feature maps are concatenated along the channel dimension to obtain the concatenated feature maps. ; Step B413: Obtain the spatial attention map: for the concatenated feature map Perform a one-dimensional convolution operation and obtain a spatial attention map using the sigmoid activation function. The expression is as follows: ; in, This represents the Sigmoid function. This represents one-dimensional convolution. The expression for one-dimensional convolution is as follows: ; in, It outputs the value of the feature vector at position t. It is the value of the input feature vector at position i. These are the weights of the convolution kernel at position i. It is a bias term; Step B414: Attention Weighting: Compare the original input feature map X with the spatial attention map Multiply to obtain the weighted output feature map. : ; in, This represents element-level multiplication operations.
7. The graph neural network gas classification and recognition method based on contrastive learning and edge labeling framework according to claim 5, characterized in that: The first residual block and the second residual block have the same structure, both of which are provided with parallel connected weight layers and identity layers. The weight layer consists of 5 one-dimensional convolutional layers connected end to end, and the identity layer consists of 1 one-dimensional convolutional layer connected by residuals. The expression for the first residual block or the second residual block is as follows: ; in, These are input features. It is the output feature. The serial number representing the residual layer; It is the weighted layer residual function, representing the input. The result after weighted layer transformation; Indicates input The result after identity layer transformation; It is the first The set of weights for each residual block. These are the parameters of the identity layer.
8. The graph neural network gas classification and recognition method based on contrastive learning and edge labeling framework according to claim 1, characterized in that: In step B5, each node in the fully connected graph a4 represents a feature vector, and each edge represents the relationship type between two nodes; the node feature initialization expression is as follows: ; in, This represents the feature of the i-th initial node. Represents the i-th input sample Extracted feature vectors; The edge feature initialization expression is as follows: ; ; in, This represents the label of the i-th node. This represents the label of the j-th node. The edge feature label between nodes i and j Represents the initial edge features between nodes i and j, each edge feature It is a two-dimensional vector, when d=1 This represents the strength of the inter-class relationship between two connected nodes, when d=2. Indicates the strength of the intra-class relationship between two connected nodes; It is a collection of intra-class and inter-class relationships; NULL indicates that the sample does not contain a label.
9. The graph neural network gas classification and recognition method based on contrastive learning and edge labeling framework according to claim 1, characterized in that: In step B6, the graph update module is equipped with an L-layer graph neural network. Each layer of the graph neural network is equipped with a node feature update unit, a similarity calculation unit, and an edge feature update unit. Each layer of the graph neural network performs one iteration of node feature update and edge feature update. The graph update module performs iterative operations on node feature updates and edge feature updates on the initialized feature graph a5, with the following steps: Step B61: The node feature update unit in the graph update module performs node feature update operation on the initial feature graph a5 to obtain the node updated feature graph, and passes it to the similarity calculation unit and the edge feature update unit. The node feature update expression is as follows: ; ; in, Represents normalized edge features. This represents summing all edges connected to node i, where k represents the index of all other nodes directly connected to node i. Represents a network with changing nodes. This indicates changes in network parameters at the nodes. Indicates the first Layered graph neural networks, ; Step B62: The similarity calculation unit calculates the similarity between every two nodes in the node update feature map through the metric network, and passes the calculated similarity to the edge feature update unit; The similarity calculation expression between any two nodes is as follows: ; ; in, This represents the similarity of the intra-class relationship between nodes i and j. This represents the similarity of the inter-class relationship between nodes i and j. Represents the metric network, These represent the parameters of the metric network; Step B63: The edge feature update unit performs edge feature update operation on the node update feature map according to the similarity, obtains the edge update feature map, and passes it to the next layer graph feature network for iterative operation; The edge feature update expression is as follows: ; in, The feature vector of the edge between nodes i and j; This represents the unnormalized edge feature vector calculated based on node features; express The L1 norm of a vector is the sum of the absolute values of all its elements. Step B64: When the number of iterations equals L, the edge feature update unit in the Lth layer graph neural network outputs the clustering feature map a6 to the classification and recognition module.
10. The graph neural network gas classification and recognition method based on contrastive learning and edge labeling framework according to claim 1, characterized in that: In step B7, the classification and recognition module extracts the similarity data between the query set sample nodes and each support set sample node in the clustering feature map a6, calculates the average similarity between the query set sample and each support set sample, and then takes the gas category corresponding to the support set sample with the largest average similarity as the gas classification result.
Citation Information
Patent Citations
Gas contour identification method
CN110889418A
Data path circuit design using reinforcement learning
CN116070557A