Class imbalance node classification method based on graph contrast learning enhancement strategy
Through the graph comparison learning framework, the key nodes are identified and mixed nodes are generated. Combined with the construction of the weighted adjacency matrix, the problem of category imbalance in the graph node classification is solved, and the feature representation quality and model recognition ability of a few types of nodes are improved.
Patent Information
- Application Number
- CN202510258331.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-13
AI Technical Summary
In graph node classification, the problem of category imbalance causes traditional methods to ignore the characteristics and patterns of a few class nodes, affecting the fairness and accuracy of the model.
The graph comparison learning framework recognizes key nodes, generates new hybrid nodes to enhance the representation of a few classes, and builds a weighted adjacency matrix to maintain the structural information of the graph, thereby improving the classification performance of the graph neural network.
It effectively solves the problem of introducing noise when generating new nodes, improves the feature representation quality of a few types of nodes, enhances the model's ability to identify a few types, and maintains the structural information of the graph.
Smart Images

Figure CN120145231A_ABST
Abstract
Description
Technical Field:
[0001] The present invention relates to the field of computer science, especially machine learning and graph data processing technologies. Specifically, the present invention relates to a graph contrastive learning enhancement strategy for solving the class imbalance node classification problem. This method identifies key nodes through a graph contrastive learning framework, generates new hybrid nodes to enhance the representation of the minority class, and constructs a weighted adjacency matrix to preserve the graph structure information, thereby improving the classification performance of graph neural networks on class-imbalanced datasets. The technical solution of the present invention is applicable to various fields that need to process and analyze graph-structured data, including but not limited to social network analysis, bioinformatics, traffic system optimization, etc., aiming to provide an effective, accurate, and fair graph node classification method for all classes. Background Art:
[0002] In computer science, graph data structures are widely used to represent relationships between entities, where nodes represent entities and edges represent connections between entities. Graph node classification is a core task in graph data analysis, which involves predicting the class to which each node in the graph belongs. This task has important applications in multiple fields, such as social network analysis, bioinformatics, recommendation systems, etc. However, in practical applications, graph data often has the problem of class imbalance, that is, the number of nodes in some classes is much larger than that in other classes, and this phenomenon is similar to the long-tail distribution.
[0003] The class imbalance problem will cause traditional machine learning algorithms to tend to focus on majority-class nodes during training, thus ignoring the features and patterns of minority-class nodes, and ultimately affecting the fairness and accuracy of the model. To solve this problem, researchers have proposed various methods, mainly divided into two categories: generation methods and loss modification methods. Generation methods enhance the representativeness of the original dataset by creating additional minority-class nodes to achieve a more balanced class distribution; loss modification methods adjust the loss function to increase the model's attention and recognition of minority-class nodes.
[0004] The core of generative methods is to increase the proportion of minority classes in the data set by synthesizing new minority class samples, thereby providing more balanced data input. The key challenge of generative methods is how to effectively generate representative minority class samples, maintaining the distribution characteristics of the original data while avoiding the introduction of unnecessary noise. Common generative methods include simple oversampling and more complex model methods, such as methods based on generative adversarial networks. For example, some existing methods migrate the classic SMOTE method to the field of graph nodes, but their consideration of the graph topology is still insufficient, making it difficult to generate structurally reliable and representative nodes. In addition, some generative methods may select nodes from other categories during the mixing process. Although the number of samples is increased, noise from irrelevant categories is often introduced, which destroys the integrity of the internal information of the minority class, thereby compressing the feature space of the minority class and affecting the model performance.
[0005] The loss correction method assigns higher weights to the errors of minority class nodes, prompting the model to pay more attention to the learning of minority class samples. Studies have found that this method is particularly suitable for graph neural networks because it can capture complex relationships and dependencies between nodes. In practice, loss correction methods are usually implemented through a variety of strategies, including assigning higher category weights to minority class nodes, introducing focal loss to reduce the model's attention to easy-to-classify nodes, and designing cost-sensitive loss functions to directly reflect the impact of different types of errors. In the problem of graph node classification, loss correction methods can balance the sample weights between categories to a certain extent, but often require more parameter tuning to adapt to graph structure data.
[0006] In recent years, graph contrastive learning has gradually become an effective unsupervised learning framework and has been widely used in tasks such as graph node classification and link prediction. The basic idea of graph contrastive learning is to construct positive and negative sample pairs so that the model can distinguish between similar and dissimilar nodes, thereby learning higher-quality node representations. Since graph contrastive learning has achieved remarkable results in the field of computer vision, researchers have gradually introduced its methods into the graph field to improve the representation ability of the model.
[0007] Graph contrastive learning provides a new approach to the class imbalance problem in graph node classification: it can not only be used to learn graph structure, but also to identify and enhance the representativeness of minority class nodes. By using contrastive learning to construct positive sample pairs related to minority class nodes, the model can more accurately capture the structure and feature information of minority classes and effectively integrate it into node representation.
[0008] Although many studies have explored the application of generative methods, loss correction methods, and graph contrast learning in graph node classification, the class imbalance problem still faces many challenges. Generative methods may introduce noise when supplementing minority class samples, destroying the feature space of the minority class; loss correction methods require complex parameter adjustment and are sometimes inadequate for complex structures in graph data; graph contrast learning provides a new idea for enhancing node representations, but how to effectively remove the interference of the majority class still needs further exploration. Summary of the Invention:
[0009] The core content of the present invention is to propose a method for class-imbalanced node classification based on a graph contrast learning enhancement strategy to solve the class imbalance problem in graph node classification. This strategy is achieved through the following steps:
[0010] S1, Process the source graph using a graph contrast learning framework to generate multiple different views. Through random augmentation operations, perturb the source graph to obtain multiple views with different features. Input these views into a shared encoder, calculate the contrast loss, and calculate the gradient of the adjacency matrix through backpropagation. Based on the gradient information, identify the more important and representative nodes in the graph to form a candidate node set;
[0011] S2, Randomly select source nodes of the minority class from the unprocessed source graph; randomly select more representative target nodes of the minority class from the candidate node set generated in step (S1). By fusing the features of the selected source nodes and target nodes under the control of a mixing ratio, generate new mixed node feature information to enhance the representation ability of the minority class nodes, thereby improving the class imbalance problem;
[0012] S3, Process the original adjacency matrix using the personalized PageRank algorithm to generate a weighted adjacency matrix. Based on the degree centrality analysis of nodes, select appropriate neighbor nodes for the newly generated mixed nodes. In this way, ensure that the new nodes can be integrated into the graph structure and maintain the overall topological characteristics of the graph.;
[0013] S4, Add the generated new nodes to the original graph and connect them according to the neighbor nodes selected in step (S3) to construct a more balanced enhanced graph. Input the enhanced graph into a graph neural network model for classification tasks, thereby improving the classification performance of the model on the class-imbalanced dataset and the recognition ability of the minority class.
[0014] The innovation points of the present invention are:
[0015] 1. Application of graph contrast learning: By identifying key nodes through a graph contrast learning framework, it effectively solves the problem of introducing noise in the existing methods when generating new nodes and improves the feature representation quality of the minority class nodes.
[0016] 2. Hybrid Node Generation Strategy: A hybrid strategy based on in-class node sampling is adopted, which avoids the interference of the majority class on the feature space of the minority class and enhances the model's recognition ability for the minority class.
[0017] 3. Construction of Weighted Adjacency Matrix: By using the personalized PageRank algorithm and degree centrality analysis, suitable neighbors are selected for the new nodes, maintaining the structural information of the graph and further improving the quality of the generated nodes and the classification performance. Description of the Drawings:
[0018] Figure 1 For the sampling and mixing strategy module - GraphCAS. Specific Embodiments:
[0019] Next, in combination with the drawings in the embodiments of the present invention, the technical solutions in the embodiments of the invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0020] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0021] As Figure 1 shown, the present invention discloses a method for class-imbalanced node classification based on a graph contrast learning enhancement strategy, including the following steps:
[0022] In S1:
[0023] (1) Randomly enhance the original graph. The specific operations include randomly deleting or inserting edges, masking node features, extracting subgraphs, etc., to generate two views with different features.
[0024] (2) Input these two views into a shared encoder to calculate the contrast loss, and its calculation formula can be expressed as:
[0025]
[0026] Among them, and respectively represent the embedding representations of node i in the two views, β represents the similarity measurement function, and τ is the temperature parameter.
[0027] (3) Calculate the gradient of the adjacency matrix through backpropagation to identify representative edges and nodes in the graph, and its calculation formula can be expressed as:
[0028]
[0029] Among them, represents the contrast loss, A 1 and X 1 respectively represent the adjacency matrix and the node feature matrix of the view Figure 1 , and f represents the encoder function.
[0030] In S2:
[0031] (1) Randomly select source nodes from the minority class of the unprocessed source graph;
[0032] (2) Randomly select more representative target nodes from the minority class in the candidate node set obtained from S1;
[0033] (3) Mix the features of the sampled source nodes and target nodes under the control of the mixing ratio λ to generate the features of the new nodes
[0034] Its calculation formula can be expressed as:
[0035] X new = λX source +(1 - λ)X target , λ ∈ [0, 1];
[0036] Among them, X new represents the features of the new nodes, X source and X target respectively represent the features of the source nodes and target nodes, and λ is a parameter controlling the mixing ratio, and the representation ability of the minority class is enhanced through this operation.
[0037] In S3:
[0038] (1) Use the personalized PageRank algorithm to convert the original adjacency matrix into a weighted adjacency matrix, and its calculation formula can be expressed as:
[0039]
[0040] Among them, T represents the transition matrix, θ r represents the attenuation factor.
[0041] (2) Based on the degree centrality analysis of the nodes, assign weights to the neighbors of the newly synthesized nodes.
[0042] In S4:
[0043] (1) Add the new nodes generated in step S2 to the original graph and connect them according to the neighbor nodes selected in step S3 to construct a more balanced enhanced graph. In this way, ensure that the new nodes can form effective connections with other nodes in the graph while maintaining the overall structural characteristics of the graph;
[0044] As described above, only the preferred specific embodiments of the present invention are given, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be within the protection scope of the present invention.
Claims
1. A method for classifying class imbalanced nodes based on graph contrast learning enhancement strategy, characterized in that: The steps include: S1, uses the graph contrastive learning framework to process the source graph and generate multiple different views. Through random enhancement operations, the source graph is perturbed to obtain multiple views with different features. These views are input into the shared encoder, the contrastive loss is calculated, and the gradient of the adjacency matrix is calculated through back propagation. Based on the gradient information, the more important and representative nodes in the graph are identified to form a set of candidate nodes; S2, randomly selects source nodes of the minority class from the unprocessed source graph; Randomly select a more representative minority class target node from the candidate node set generated in step (S1). By fusing the features of the selected source node and the target node under the control of the mixing ratio, new mixed node feature information is generated to enhance the representation ability of the minority class node, thereby improving the class imbalance problem; S3, use the personalized PageRank algorithm to process the original adjacency matrix and generate a weighted adjacency matrix. Based on the degree centrality analysis of the nodes, select appropriate neighbor nodes for the newly generated hybrid nodes. In this way, ensure that the new nodes can be integrated into the structure of the graph and maintain the overall topological characteristics of the graph. ; S4, adding the generated new nodes to the original graph and connecting them according to the neighbor nodes selected in step (S3) to construct a more balanced enhanced graph. The enhanced graph is input into the graph neural network model for classification tasks, thereby improving the classification performance of the model on the category imbalanced dataset and the recognition ability of the minority class.
2. According to claim 1, a method for classifying class imbalanced nodes based on graph contrast learning enhancement strategy is characterized in that: In S1, the graph contrast learning framework includes the following operations: (1) Perform random enhancement on the original graph, including randomly deleting or inserting edges, masking node features, extracting subgraphs, etc., to generate two views with different features; (2) These two views are input into the shared encoder to calculate the contrast loss, which can be expressed as: in, and They represent the embedding representation of node i in the two views, β represents the similarity measure function, and τ is the temperature parameter. (3) The gradient of the adjacency matrix is calculated by back propagation to identify representative edges and nodes in the graph. The calculation formula can be expressed as: in, represents the contrastive loss, A1 and X1 represent the adjacency matrix and node feature matrix of view 1 respectively, and f represents the encoder function.
3. According to claim 1, a method for classifying class imbalanced nodes based on graph contrast learning enhancement strategy is characterized in that: In S2, the following operations are performed on the graph: (1) Randomly select source nodes from the minority class of the unprocessed source graph; (2) Randomly select a more representative target node from the minority class in the candidate node set obtained from S1; (3) The features of the source node and the target node obtained by sampling are mixed under the control of the mixing ratio λ to generate the features of the new node. The calculation formula can be expressed as: X new =λX source +(1-λ)X target ,λ∈[0,1]; Among them, X new Represents the characteristics of the new node, X source and X target They represent the characteristics of the source node and the target node respectively, and λ is a parameter that controls the mixing ratio. This operation is used to enhance the representation ability of the minority class.
4. According to claim 1, a method for classifying class imbalanced nodes based on graph contrast learning enhancement strategy is characterized in that: In S3, constructing a weighted adjacency matrix includes the following operations: (1) Use the personalized PageRank algorithm to convert the original adjacency matrix into a weighted adjacency matrix, and its calculation formula can be expressed as: Where T represents the transfer matrix, θ r Represents the attenuation factor. (2) Based on the degree centrality analysis of the nodes, weights are assigned to the neighbors of the newly synthesized nodes.
5. The method for classifying class imbalanced nodes based on graph contrast learning enhancement strategy according to claim 1, characterized in that: In step S4, The following operations are included: (1) Add the new node generated in step S2 to the original graph and connect it according to the neighbor nodes selected in step S3 to build a more balanced enhanced graph. In this way, it is ensured that the new node can form an effective connection with other nodes in the graph while maintaining the overall structural characteristics of the graph.