Imbalanced node classification method based on graph contrast learning

The Encoder-Decoder architecture generates high-quality pseudo-labels and combines adaptive sampling and augmentation technology to solve the problem of node classification on category imbalance graph data in graph comparison learning, and improves the model's minority class recognition capabilities and overall performance in an unsupervised environment.

CN120336951APending Publication Date: 2025-07-18NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510376650.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing graph comparison learning method cannot effectively retain the information of a few class nodes when processing category imbalanced graph data, resulting in the degradation of the classification performance of the model on unbalanced data, especially in an unsupervised environment.

Method used

The Encoder-Decoder architecture is used for pre-training to generate high-quality pseudo-labels, and through adaptive sampling strategies and augmentation technology that prioritizes the preservation of minority nodes, combined with the pseudo-label information, data augmentation strategies are designed to improve the model's ability to identify minority classes.

Benefits of technology

Under the self-supervision conditions, the classification effect of a few types of nodes is significantly improved, the overall performance of the model on the unbalanced data set is improved, and the generalization ability of the model is enhanced through adaptive sampling and augmented technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336951A_ABST
    Figure CN120336951A_ABST
Patent Text Reader

Abstract

The invention discloses an unbalanced node classification method based on graph contrast learning, and belongs to the technical field of artificial intelligence. According to the method, a graph comparison learning framework of adaptive balance data is provided, minority classes can be automatically identified, the minority class performance is improved, and then the overall performance of the model is improved. Firstly, an Encoder-Decoder architecture is used for pre-training, and compared with a traditional pseudo tag generation method, an unbalance rate self-adaptive sampling strategy is designed, the unbalance rate of data is calculated according to pseudo tags, and the sampling strategy is selected in a self-adaptive mode. For a data set with a low unbalance rate, a simple downsampling method is adopted, and the proportion of minority class information is increased; for a data set with a relatively high unbalance rate, a mixed sampling strategy is adopted, and over-sampling and down-sampling are combined, so that the information loss of majority of nodes is reduced while the information proportion of minority of nodes is increased. In addition, the pre-training model used in the invention can provide more accurate label information, thereby improving the distinguishing ability of the model in subsequent GCL training. Then, a new data augmentation technology is designed, in the node masking process, pseudo label information is utilized, information of minority class nodes is reserved preferentially, and meanwhile majority class nodes are masked; the method is helpful for the model to better capture minority class features in an unbalanced data set. And finally, a linear classifier is used for classification. According to the method, the unbalanced node classification performance under the self-supervision condition can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of graph data processing and the field of imbalance learning. Specifically, it relates to a graph contrast learning method for improving the classification performance of unbalanced data nodes. Background Art

[0002] As a complex data structure, graph data is widely used in fields such as social networks, recommendation systems, bioinformatics, knowledge graphs, molecular property prediction, etc. In these applications, graph data usually consists of nodes and edges, where nodes represent entities (such as users, molecules, papers, etc.), and edges represent the relationships between entities (such as social relationships, chemical bonds, citation relationships, etc.). With the wide application of graph data, graph neural networks (GNNs) and graph representation learning (GRL) techniques have developed rapidly, especially achieving remarkable results in tasks such as node classification, graph classification, and link prediction. However, graph data in the real world often has the problem of class distribution imbalance, that is, the number of nodes in some classes far exceeds that in other classes. For example, in a social network, most users may belong to the ordinary user category, while a few users may belong to high-value users or abnormal users; in molecular property prediction, the number of molecular samples in some classes may be much less than that in other classes. This class imbalance problem will cause traditional graph neural networks and contrast learning methods to tend to learn the features of majority-class nodes and ignore the features of minority-class nodes when processing graph data, thus affecting the overall classification performance of the model.

[0003] Graph contrast learning (GCL), as a self-supervised learning technique, has demonstrated excellent performance in node representation learning and graph classification tasks of graph data in recent years. The core idea of graph contrast learning is to generate different views of graph data and maximize the representation consistency between these views, thereby learning discriminative node representations. Specifically, graph contrast learning usually includes three main modules: an augmentation function, an encoder, and a contrast loss function. The augmentation function is used to generate different views of graph data, the encoder (usually a graph neural network) is used to capture the feature information and structural information of graph data, and the contrast loss function is used to maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs.

[0004] Although graph contrast learning has achieved success in many tasks, it still faces significant challenges when dealing with class-imbalanced graph data. Existing graph contrast learning methods usually adopt a random augmentation strategy, that is, randomly masking nodes or edges during the augmentation process. However, this random augmentation strategy may cause the information of minority-class nodes to be lost or weakened during the augmentation process, thus further exacerbating the class imbalance problem. In addition, since graph contrast learning cannot obtain the label information of data during the training process, this important characteristic of class imbalance is ignored by most existing graph contrast learning methods.

[0005] Traditional methods for handling class imbalance are divided into data-level methods and algorithm-level methods. Data-level methods balance the class distribution by modifying the distribution of training data. Common strategies include downsampling the majority class, oversampling the minority class, and hybrid sampling methods that combine downsampling and oversampling. For example, the Synthetic Minority Over-sampling Technique (SMOTE) generates new samples by interpolating between minority class samples, thus increasing the number of minority class samples. Algorithm-level methods modify the learning algorithm itself to better handle class imbalance problems. Common strategies include cost-sensitive learning, ensemble learning, and loss function engineering. For example, the cost-sensitive learning method assigns higher weights to minority class samples, making the model pay more attention to minority class samples during training. Although imbalance learning has been widely applied in traditional machine learning tasks, its application in graph data still faces challenges. Graph data has complex structural information, and traditional class imbalance learning methods cannot be directly applied to graph data. In recent years, some studies have attempted to combine class imbalance learning with graph structural information, proposing methods such as data interpolation methods on graphs (such as GraphSMOTE), adversarial generation methods (such as GraphENS), etc. However, these methods still have certain limitations when dealing with class-imbalanced graph data and require the provision of label information, which cannot be provided by graph contrastive learning.

[0006] In summary, studying how to break through the limitations of labels and effectively retain the information of minority class nodes, thereby improving the overall performance of imbalanced data, is the core challenge in this field. In an unsupervised environment, label information cannot be obtained, and the random augmentation strategy of graph contrastive learning may lead to the loss of minority class information, making the difficulty of imbalanced learning on graphs in an unsupervised environment much higher than that in a supervised graph learning environment. Therefore, developing a method with high-quality pseudo-label generation and an augmentation strategy that preferentially retains minority class nodes has become an urgent need in the field of unsupervised graph imbalanced learning. Summary of the Invention

[0007] Object of the Invention: Aiming at the limitations of existing Graph Contrastive Learning (GCL) methods in dealing with class-imbalanced graph data, the present invention proposes a new imbalanced node classification method based on graph contrastive learning. By introducing pseudo-label generation, an adaptive sampling strategy, and an augmentation technique that preferentially retains minority class nodes, this method can automatically balance node representations and improve the classification effect of minority class nodes under self-supervised conditions.

[0008] Technical solution: The present invention provides an unbalanced node classification method based on graph contrast learning, which is characterized in that the method uses a pre-trained Encoder-Decoder architecture model to generate high-quality pseudo-labels, designs an adaptive sampling strategy using the pseudo-labels to balance node categories, and preferentially retains the information of minority-class nodes in the augmentation stage to improve the model's recognition ability for minority-class samples. The method includes the following steps:

[0009] S1. Obtain graph data, preprocess the data to obtain preprocessed data.

[0010] S2. Pre-train a variational autoencoder VGAE with an Encoder-Decoder structure, and obtain the output Z of the trained variational autoencoder VGAE 1VGAE .

[0011] S3. The present invention uses a simple and practical K-Means method to cluster Z VGAE to obtain high-quality pseudo-labels. The clustering update formula is: where K is the preset number of clusters, S k is the set of data points assigned to the k-th cluster center c k , c k is the k-th cluster center, and z i is the i-th data point. We assign the label of each cluster center c k as the pseudo-labels of all nodes in the cluster to obtain a high-quality pseudo-label set for subsequent graph contrast learning.

[0012] S4. Calculate the data imbalance rate N k is the number of nodes of the k-th class calculated according to the pseudo-labels . The present invention designs two sampling strategies according to the imbalance rate. Briefly speaking, when the imbalance rate is low, only downsampling is used to balance the data set and improve the accuracy of the minority class; when the imbalance rate is high, in order not to overly lose the information of the majority class, both downsampling and oversampling are used to better balance the data set.

[0013] S5. Calculate according to the category of each node Use the hyperparameter ɑ to weigh between the two sampling strategies. Combine the node centrality to calculate the probability Use the hyperparameter λ to weigh the two. After normalizing it, the obtained is used as the sampling probability of the node:

[0014]

[0015] ​Among them, N k is the number of nodes of the k-th class calculated according to the pseudo-labels , where K is the number of classes, is the degree of the normalized node, represents the sampling probability of the k-th class, and represents the probability that each node of the k-th class is sampled.

[0016] S6. Downsample the nodes according to .

[0017] S7. Oversample the nodes according to (whether this step is executed is determined by S4).

[0018] S8. Generate two augmented graphs for the generated new graph data.

[0019] S9. Training of the graph contrast model. Input the two augmented graphs into the model to obtain outputs z1, z1. The present invention optimizes the model by minimizing the redundancy between different augmented views while maximizing the mutual information between the views. The loss function is specifically:

[0020]

[0021] where N is the batch size, and ‖·‖ F is the Frobenius norm.

[0022] S10. Input the obtained features into a linear classifier to evaluate its performance.

[0023] Furthermore, the specific process of step S1 includes:

[0024] S1.1. The data used in the present invention is provided by the third-party package torch_geometric.datasets and runs in the PyTorch environment. If using graphs not in this third-party package or graph data from other sources, the data format should be processed into a format that can be processed by the PyTorch framework.

[0025] S1.2. Normalize the adjacency matrix. The formula is:

[0026]

[0027] where A is the adjacency matrix, D is the degree matrix, and I is the identity matrix. This processing makes the values of the adjacency matrix more stable and avoids problems of gradient explosion or gradient disappearance in subsequent graph convolution operations.

[0028] S1.3. To ensure the comparability of the loss function on graphs of different scales, use the normalization factor norm to adjust the loss function. The formula is:

[0029]

[0030] Among them, N is the number of nodes, and A is the adjacency matrix.

[0031] The specific process of step S2 includes:

[0032] S2.1. Use the encoder to map the input feature X and the normalized adjacency matrix A * to the hidden space, and output the hidden feature H v , the mean μ v , the variance and the latent variable Z v . Specifically:

[0033]

[0034] Z v = μ v + ∈ ⊙ σ v

[0035] Among them, is the weight matrix of different convolutional layers of the model; represents the random noise of the standard normal distribution.

[0036] S2.2. Use the output of the encoder to train the discriminator. The present invention uses the standard normal distribution to generate random noise as the real sample, regards the output of the encoder as the generated sample, and designs the following loss function to distinguish the real sample and the generated sample:

[0037]

[0038] Among them, D real,i is the output of the discriminator for the i-th real sample, and D fake,j is the output of the discriminator for the j-th generated sample. 1 represents a vector of all 1s, with the same shape as D real,i ; 0 represents a vector of all 0s, with the same shape as D fake,j . N and M are the numbers of real samples and generated samples respectively. BCE is the binary cross-entropy loss function, and y and represent the real label and the predicted probability respectively. The present invention makes the prediction result closer to the real label by minimizing BCE. Then, update the parameters of the discriminator according to the loss function:

[0039]

[0040] Among them, η D is the learning rate of the discriminator, is the loss function with respect to the parameter θD Gradient.

[0041] S2.3. Train the variational auto - encoder VGAE. The present invention combines the discriminator loss with the KL - divergence loss to design a reconstruction loss We use the reconstruction loss to train the variational auto - encoder to enable it to have the ability to reconstruct the input data and generate new data. The formula is as follows:

[0042]

[0043] where μ i and σ i are the mean and the logarithm of the standard deviation output by the encoder respectively. Then, the parameters of the auto - encoder are updated according to the loss function:

[0044]

[0045] where η V is the learning rate of the variational auto - encoder, is the loss function with respect to the parameter θ V Gradient.

[0046] The specific process of step S6 includes:

[0047] S6.1. Calculate the mean

[0048] S6.2. Calculate the number of nodes to be retained after down - sampling for all categories with the number of nodes greater than the mean:

[0049]

[0050] where β less is the down - sampling ratio hyper - parameter.

[0051] S6.3. Down - sample the nodes. The number of nodes retained for each category is determined by Less k and the probability of a node being selected in each category is determined by to generate a mask MASK less .

[0052] The specific process of step S7 includes:

[0053] S7.1. Calculate the mean

[0054] S7.2. Calculate the number of nodes to be retained after down - sampling for all categories with the number of nodes greater than the mean:

[0055]

[0056] Among them, β more is the downsampling ratio hyperparameter.

[0057] S7.3. The number of nodes that need to be oversampled for each class node is determined by More k , and the probability that a node is selected in each class is determined by . Generate a mask MASK for the nodes that need to be oversampled more .

[0058] S7.4. Use the SMOTE method to oversample each node marked by MASK more .

[0059] v new =δv target +(1 - δ)v neigh

[0060] Among them, v target is the node that needs to be oversampled, v neigh is one of the neighbors of node v calculated using K-nearest neighbors target , and δ is a random number, which increases the diversity of the training data and avoids overfitting.

[0061] S7.5. The new node v new generated in S7.4 is an isolated node. The present invention uses the nearest neighbor method to generate a new edge for it. First, find the set V new including its k neigh nearest neighbor nodes for the generated new node v neigh :

[0062]

[0063] Among them, V is the set of all nodes, V' is a subset of V, and the number of nodes in V' is k neigh . Then, we select n neigh nodes from these k e nodes to generate an edge between them and the new node v new generated in S7.4, and add a self-loop to v new .

[0064] The specific process of step S8 includes:

[0065] S8.1. Calculate the number of nodes in each class as N' k , and the majority class label set is

[0066] S8.2. Masking Probability Function Dynamically Adjusted According to Class Distribution:

[0067]

[0068] where P f is the base probability of node masking in the augmentation function, and N′ is the total number of nodes.

[0069] S8.3. Selecting Masked Nodes in Node Masking According to Probability: Mask f1 = (U < P′f). Where U ∈ [0, 1] N is a random vector sampled from a uniform distribution. Excluding the minority class nodes in Mask f1 to obtain the final masked nodes: Beneficial Effects. Compared with the prior art, the significant effects and substantial features of the present invention mainly lie in:

[0070] (1) The Encoder-Decoder architecture is adopted for pre-training. Compared with the traditional method of generating pseudo-labels, the pre-training model used in the present invention can provide more accurate label information, thereby improving the discrimination ability of the model in subsequent GCL training.

[0071] (2) An unbalanced rate adaptive sampling strategy is designed. The unbalanced rate of the data is calculated based on the pseudo-labels, and the sampling strategy is adaptively selected to balance the performance between the majority class and the minority class to better improve the overall performance. For datasets with a relatively low unbalanced rate, the present invention adopts a simple downsampling method to increase the proportion of minority class information; for datasets with a relatively high unbalanced rate, a hybrid sampling strategy is adopted, combining oversampling and downsampling, to increase the proportion of minority class node information while reducing the loss of majority class node information.

[0072] (3) A new data augmentation technology is developed. It utilizes pseudo-label information during the node masking process, preferentially retains the information of minority class nodes, and masks the majority class nodes at the same time. This method helps the model better capture the features of the minority class in unbalanced datasets.

[0073] (4) Extensive experiments have been conducted on multiple datasets, and the results show the effectiveness and superiority of the method in dealing with unbalanced node classification problems. The research on ablation experiments further demonstrates the effectiveness of the method. Description of the Drawings

[0074] Figure 1 is the workflow diagram of the method of the present invention;

[0075] Figure 2 is the training flowchart of the Encoder-Decoder model in the present invention;

[0076] Figure 3 It is the overall framework diagram of the system to which the method of the present invention is applied. Detailed implementation manners

[0077] To make the objectives, advantages and technical solutions of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be described completely and clearly below with reference to the accompanying drawings;

[0078] The research on unbalanced node classification under unsupervised conditions is of great value for understanding complex network structures and improving the robustness of classification algorithms. Graph contrast learning technology, based on graph neural networks (GNNs), can reveal the potential relationships between nodes through unsupervised learning and provide key bases for node classification. However, most graph contrast learning methods mainly rely on similarity metrics of node features, constructing classification models by calculating the distances or similarities between nodes. Although such methods have been widely applied, their neglect of the implicit feature of class distribution imbalance is very likely to lead to a decline in the classification performance of minority-class nodes. In addition, single-scale feature extraction cannot effectively capture the structural information of nodes at different levels, restricting the generalization ability of classification models. Both types of methods are affected by the low expressiveness of predefined similarity metrics and lack adaptability to unbalanced data. Finally, the key point of this research is to solve the problem of uneven class distribution in the dataset, thereby improving the classification performance of the model for minority-class samples. In addition, the research needs to be carried out in an unlabeled environment, making the difficulty even greater.

[0079] What the present invention provides is an unbalanced node classification method based on graph contrast learning, combined with Figure 1 the process shown in Figure 3 and the method framework diagram shown in, the embodiments adopt technologies such as pseudo-label generation, adaptive sampling strategies, and self-balanced augmentation to achieve unbalanced node classification.

[0080] S1. Obtain graph data, preprocess the data, and obtain the preprocessed data.

[0081] The data used in the present invention is provided by the third-party package torch_geometric.datasets and runs in the PyTorch environment. If graph data not available in this third-party package or from other sources is used, the data format should be processed into a format that can be processed by the PyTorch framework. Then, normalize the adjacency matrix:

[0082]

[0083] Among them, A is the adjacency matrix, D is the degree matrix, and I is the identity matrix. This processing makes the values of the adjacency matrix more stable and avoids the problems of gradient explosion or gradient disappearance in subsequent graph convolution operations. To ensure the comparability of the loss function on graphs of different scales, a normalization factor norm is used to adjust the loss function:

[0084]

[0085] Among them, N is the number of nodes and A is the adjacency matrix.

[0086] S2. Combine Figure 2 the training flowchart of the Encoder-Decoder model to pre-train a variational autoencoder VGAE with an Encoder-Decoder structure and obtain the output Z of the trained variational autoencoder VGAE VGAE .

[0087] First, use the encoder to map the input feature X and the normalized adjacency matrix A * to the hidden space and output the hidden feature H v , the mean μ v , the variance and the latent variable Z v . Specifically:

[0088]

[0089] Among them, is the weight matrix of different convolutional layers of the model; represents the random noise of the standard normal distribution. Subsequently, use the output of the encoder to train the discriminator. The present invention uses the standard normal distribution to generate random noise as the real sample, regards the output of the encoder as the generated sample, and designs the following loss function to distinguish the real sample and the generated sample:

[0090]

[0091] Among them, D real,i is the output of the discriminator for the i-th real sample, and D fake,j is the output of the discriminator for the j-th generated sample. 1 represents a vector of all 1s, with the same shape as D real,i ; 0 represents a vector of all 0s, with the same shape as D fake,j . N and M are the numbers of real samples and generated samples respectively. BCE is the binary cross-entropy loss function, and y and represent the real label and the predicted probability respectively. The present invention makes the prediction result closer to the real label by minimizing BCE. Then, update the parameters of the discriminator according to the loss function:

[0092]

[0093] Among them, η D is the learning rate of the discriminator, is the loss function with respect to the parameter θ D gradient. Subsequently, the present invention combines the discriminator loss with the KL divergence loss to design a reconstruction loss We use the reconstruction loss to train the variational autoencoder VGAE to enable it to have the ability to reconstruct input data and generate new data. The formula is as follows:

[0094]

[0095] Among them, μ i and σ i are the mean and the logarithm of the standard deviation output by the encoder respectively. Update the parameters of the autoencoder according to the loss function:

[0096]

[0097] Among them, η V is the learning rate of the variational autoencoder, is the loss function with respect to the parameter θ V gradient.

[0098] S3. The present invention uses the simple and practical K-Means method to cluster Z VGAE to obtain high-quality pseudo-labels. The clustering update formula is: Among them, K is the preset number of clusters, S k is the set of data points assigned to the k-th cluster center c k , c k is the k-th cluster center, z i is the i-th data point. We assign the label of each cluster center c k to the pseudo-labels of all nodes in this cluster to obtain a high-quality pseudo-label set for subsequent graph contrast learning.

[0099] S4. Calculate the data imbalance rate N k is based on the pseudo-labels The calculated number of nodes of the k-th class. The present invention designs two sampling strategies according to the imbalance rate. Briefly speaking, when the imbalance rate is low, only downsampling is used to balance the dataset and improve the accuracy of the minority class; when the imbalance rate is high, in order not to excessively lose the information of the majority class, both downsampling and oversampling are used to better balance the dataset.

[0100] S5. Calculate according to the class of each node Use the hyperparameter ɑ to weigh between the two sampling strategies. Calculate the probability in combination with the node centrality Use the hyperparameter λ to weigh the two. After normalizing it, the obtained is used as the sampling probability of the node:

[0101]

[0102] where N k is the number of nodes of the k-th class calculated according to the pseudo-label , K is the number of classes, is the degree of the normalized node, represents the sampling probability of the k-th class, represents the probability that each node of the k-th class is sampled.

[0103] S6. Downsample the nodes according to . First, calculate the mean value Then, calculate the number of nodes that need to be retained after downsampling for all classes with the number of nodes greater than the mean value:

[0104]

[0105] where β less is the downsampling ratio hyperparameter. After that, downsample the nodes. The number of nodes retained for each class is determined by Less k , and the probability that a node is selected in each class is determined by . Generate a mask MASK less for the retained nodes.

[0106] S7. Upsample the nodes according to (it is determined whether to execute this step by S4). First, calculate the mean value Then, calculate the number of nodes that need to be retained after downsampling for all classes with the number of nodes greater than the mean value:

[0107]

[0108] where β moreis the downsampling ratio hyperparameter. Then determine the nodes that need to be oversampled. The number of nodes that need to be oversampled for each class node is determined by More k decide, and the probability of a node being selected in each class is determined by decide, generate a mask MASK more for the nodes that need to be oversampled. For each node marked by MASK more , use the SMOTE method for oversampling.

[0109] v new =δv target +(1 - δ)v neigh

[0110] where v target is the node that needs to be oversampled, and v neigh is one of the nodes in the neighborhood of node v calculated using K - nearest neighbors target , and δ is a random number, which increases the diversity of the training data and avoids overfitting. Since the newly generated node v new is an isolated node, the present invention uses the nearest neighbor method to generate new edges for it. First, find the set V new including its k neigh nearest neighbor nodes for the newly generated node v neigh :

[0111]

[0112] where V is the set of all nodes, V’ is a subset of V, and the number of nodes in V’ is k neigh . Then, we select n neigh nodes from these k e nodes to generate the edges between them and the newly generated node v in S7.4 new , and add a self - loop to v new .

[0113] S8. Generate two augmented graphs for the newly generated graph data. First, calculate the number of nodes in each class as N’ k , and the set of majority class labels is Then, according to the masking probability function dynamically adjusted according to the class distribution:

[0114]

[0115] where P f is the base probability of node masking in the augmentation function, and N′ is the total number of nodes. Finally, select the masked nodes in node masking according to the probability: Mask f1 =(U < P' f ). Where U ∈ [0, 1] Nis a random vector sampled from a uniform distribution. Exclude Mask f1 The minority class nodes inside obtain the final masked nodes:

[0116] S9. Training of the graph contrast model. Two augmented graphs are input into the model to obtain outputs z1 and z1. The present invention optimizes the model by minimizing the redundancy between different augmented views while maximizing the mutual information between the views. The loss function is specifically:

[0117]

[0118] where N is the batch size, ‖·‖ F is the Frobenius norm.

[0119] S10. Input the obtained features into a linear classifier to evaluate its performance.

[0120] Through the embeddings generated by Encoder-Decoder, we obtain high-quality pseudo-labels, which can better guide the training of the model and also play a positive role in balancing the data distribution. The unbalanced rate adaptive sampling strategy designed for the model calculates the unbalanced rate of the data based on the pseudo-labels and adaptively selects the sampling strategy to balance the performance between the majority class and the minority class to better improve the overall performance. The graph contrast module uses a new data augmentation technique that utilizes pseudo-label information during the node masking process, preferentially retaining the information of minority class nodes while masking majority class nodes. This method helps the model better capture the features of the minority class in an unbalanced dataset. The present invention comprehensively uses the above technologies to enhance the performance of the unbalanced node classification task in a self-supervised environment.

Claims

1. An imbalanced node classification method based on graph contrast learning, characterized in that this method uses pre-training on the Encoder-Decoder architecture to generate high-quality pseudo-labels, designs an adaptive sampling strategy using the pseudo-labels to balance node classes, and preferentially retains minority-class node information in the augmentation stage to improve the model's recognition ability for minority-class samples. It includes the following steps: S1. Obtain graph data, preprocess this data to obtain preprocessed data. S2. Pre-train a variational autoencoder VGAE with an Encoder-Decoder structure, and obtain the output of the trained variational autoencoder VGAE as Z VGAE . S3. The simple and practical K-Means method of the present invention is used to cluster Z VGAE to obtain high-quality pseudo-labels. The clustering update formula is: where K is the preset number of clusters, S k is the set of data points assigned to the k-th cluster center c k , c k is the k-th cluster center, and z i is the i-th data point. We assign the label of each cluster center c k as the pseudo-labels of all nodes in the cluster, obtaining a high-quality pseudo-label set for subsequent graph contrastive learning. S4. Calculate the data imbalance rate N k is the number of nodes of the k-th class calculated according to the pseudo label The present invention designs two sampling strategies according to the imbalance rate. Briefly speaking, when the imbalance rate is low, only downsampling is used to balance the data set and improve the accuracy of the minority class; when the imbalance rate is high, in order not to overly lose the information of the majority class, both downsampling and oversampling are used to better balance the data set. S5. Calculate according to the category of each node Use the hyperparameter ɑ to balance between the two sampling strategies. Calculate the probability by combining node centrality Use the hyperparameter λ to balance the two. For After normalizing it, the obtained is used as the sampling probability of the node: Among them, N k is the number of nodes of the k-th class calculated according to the pseudo-label , K is the number of classes, is the degree of the normalized node, represents the sampling probability of the k-th class, represents the probability that each node of the k-th class is sampled. S6. According to downsample the nodes. S7. According to oversample the nodes (whether to execute this step is determined by S4) S8. Generate two augmented graphs for the generated new graph data. S9. Training of the graph contrast model. Input the two augmented graphs into the model to obtain outputs z1, z1. The present invention optimizes the model by minimizing the redundancy between different augmented views while maximizing the mutual information between views. The loss function is specifically: where N is the batch size, and ||·|| F is the Frobenius norm. S10. Input the obtained features into a linear classifier to evaluate its performance.

2. The unbalanced node classification method based on graph contrastive learning according to claim 1, wherein The specific process of step S1 includes: S1.

1. The data used in the present invention is provided by the third-party package torch_geometric.datasets and runs in the PyTorch environment. If using graphs not in this third-party package or graph data from other sources, the data format should be processed into a format that can be processed by the PyTorch framework. S1.

2. Normalize the adjacency matrix, and the formula is: where \(A\) is the adjacency matrix, \(D\) is the degree matrix, and \(I\) is the identity matrix. This processing makes the values of the adjacency matrix more stable and avoids problems such as gradient explosion or gradient disappearance in subsequent graph convolution operations. S1.

3. To ensure the comparability of the loss function on graphs of different scales, use the normalization factor norm to adjust the loss function, and the formula is: where \(N\) is the number of nodes and \(A\) is the adjacency matrix.

3. The unbalanced node classification method based on graph contrastive learning according to claim 1, characterized in that The specific process of step S2 includes: S2.

1. Use an encoder to map the input feature X and the normalized adjacency matrix A * to the hidden space, and output the hidden feature H v , the mean μ v , the variance and the latent variable Z v . Specifically: Z v = μ v + ε ⊙ o v Among them, is the weight matrix of different convolutional layers of the model; represents random noise of the standard normal distribution. S2.

2. Train the discriminator using the output of the encoder. The present invention uses the standard normal distribution to generate random noise as real samples, regards the output of the encoder as generated samples, and designs the following loss function to distinguish real samples from generated samples: Among them, D real,i is the output of the discriminator for the i-th real sample, and D fake,j is the output of the discriminator for the j-th generated sample. 1 represents a vector of all 1s, with the same shape as D real,i ; 0 represents a vector of all 0s, with the same shape as D fake,j . N and M are the numbers of real samples and generated samples respectively. BCE is the binary cross-entropy loss function, where y and represent the real label and the predicted probability respectively. In the present invention, the prediction result is made closer to the real label by minimizing BCE. Then, the parameters of the discriminator are updated according to the loss function: Among them, η D is the learning rate of the discriminator, is the loss function with respect to the parameter θ D gradient. S2.

3. Train the variational autoencoder VGAE. The present invention combines the discriminator loss with the KL divergence loss to design a reconstruction loss We use the reconstruction loss to train the variational autoencoder to have the ability to reconstruct the input data and generate new data. The formula is as follows: Among them, μ i and σ i are respectively the logarithms of the mean and standard deviation of the encoder output. Then, the parameters of the autoencoder are updated according to the loss function: Among them, η V is the learning rate of the variational autoencoder, is the loss function with respect to the parameter θ V gradient.

4. The unbalanced node classification method based on graph contrastive learning according to claim 1, characterized in that, The specific process of step S6 includes: First, calculate the mean value Then, calculate the number of nodes to be retained after downsampling for all categories with the number of nodes greater than the mean value: Among them, β less is the downsampling ratio hyperparameter. After that, downsampling is performed on the nodes, and the number of nodes retained for each category node is determined by Less k , and the probability of a node being selected in each category is determined by . A mask MASK less is generated for the retained nodes.

5. The unbalanced node classification method based on graph contrastive learning according to claim 1, wherein, The specific process of step S7 includes: First, calculate the mean value Then, calculate the number of nodes to be retained after downsampling for all categories with the number of nodes greater than the mean value: Among them, β more is the downsampling ratio hyperparameter. Then determine the nodes that need to be oversampled. The number of nodes that need to be oversampled for each category node is determined by More k . The probability of a node being selected in each category is determined by . Generate a mask MASK more for the nodes that need to be oversampled. Use the SMOTE method to oversample each node identified by MASK more . v new = δv target + (1 - δ)v neigh Among them, v target is the node that needs to be oversampled, and v neigh is one of the neighbors of node v calculated using K-nearest neighbors. δ is a random number that increases the diversity of the training data and avoids overfitting. Since the generated new node v target is an isolated node, the present invention uses the nearest neighbor method to generate a new edge for it. First, for the generated new node v new find the set V new that includes its k neigh nearest neighbor nodes: neigh : Among them, V is the set of all nodes, V' is a subset of V, and the number of nodes in V' is k neigh . After that, we select n neigh nodes from these k e nodes to generate the edges between them and the new node v generated in S7.4 new , and add a self-loop to v new .

6. The unbalanced node classification method based on graph contrastive learning according to claim 1, wherein The specific process of step S8 includes: First, calculate the number of nodes in each category as N’k, and the majority class label set is Then, the masking probability function dynamically adjusted according to the category distribution: where, P f is the base probability of node masking in the augmentation function, and N′ is the total number of nodes. Finally, the masked nodes in the node masking are selected according to the probability: Mask f1 = (U < P′ f ). Where, U ∈ [0, 1] N is a random vector sampled from a uniform distribution. Excluding the minority class nodes in Mask f1 gives the final masked nodes:

7. The unbalanced node classification method based on graph contrastive learning according to claim 1, characterized in that Through the embeddings generated by the Encoder-Decoder, we obtain high-quality pseudo-labels, which can better guide the training of the model and also play a positive role in balancing the data distribution. The designed imbalance rate adaptive sampling strategy of the model calculates the imbalance rate of the data according to the pseudo-labels and adaptively selects the sampling strategy, weighing between the performance of the majority class and the minority class to better improve the overall performance. The graph contrast module uses a new data augmentation technique, utilizes the pseudo-label information during the node masking process, preferentially retains the information of minority-class nodes, and at the same time masks the majority-class nodes. This method helps the model better capture the features of the minority class in an imbalanced dataset. The present invention comprehensively uses the above technologies to enhance the performance of the imbalanced node classification task in a self-supervised environment.

Citation Information

Cited By

  • Natural resource element identification method based on spatial correlation memory

    CN121095778A