Unbalanced graph node classification method and system based on boundary node condition GAN

By introducing boundary node conditional GANs into GNNs and using class information and topological information for conditional adversarial training, the problem of poor classification effect of GNNs on class imbalance graph data is solved, and better node embedding and classification results are achieved.

CN120045762AActive Publication Date: 2025-05-27BEIJING TECH & BUSINESS UNIV

Patent Information

Application Number
CN202510104523.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

When existing GNNs models deal with graph data with unbalanced classes, it is difficult to effectively learn and classify samples of majority and minority classes, resulting in poor overall classification results.

Method used

A method of imbalanced graph node classification based on boundary node condition GAN is proposed. By using class information and context topology information as conditional input condition GAN for conditional confrontation training, the topology information in graph structure data is fully utilized to promote the learning of classifiers.

Benefits of technology

Through adversarial network training, the distance between categories is expanded, so that the graph neural network model can better learn node embedding, overcoming the problem of poor overall classification results caused by insufficient representation of subclass embedding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045762A_ABST
    Figure CN120045762A_ABST
Patent Text Reader

Abstract

The invention discloses an unbalanced graph node classification method and system based on a boundary node condition GAN. The method comprises the following steps: constructing a graph structure data set to be classified; inputting the to-be-classified graph structure data set into a graph convolutional network for node classification for processing, and outputting a classification result; wherein the graph convolutional network for node classification is trained through a training data set and is obtained according to convolutional layer parameters obtained through training, and the training set is a graph structure data set. According to the method, the adversarial network training process is applied to class decision boundary nodes, the inter-class distance is enlarged, the graph neural network model better learns node embedding, and the problem that the overall classification result is poor due to insufficient subclass embedding representation is solved under the condition that noise nodes are not introduced into an original graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an unbalanced graph node classification method and system based on boundary node condition GAN. Background Art

[0002] Since graph-structured data is very common in the real world, almost all complex systems can be naturally represented as graph structures, such as social networks, financial networks, and protein interaction networks. In recent years, Graph Neural Networks (GNNs) have received increasing attention and research. Different from traditional deep learning models that process regular grid data, GNNs focus on capturing and learning the complex relationships and topological structures between nodes in a graph. Through an iterative process of information passing and aggregation, GNNs enable each node representation to gradually focus on the information related to it, in order to obtain an effective information representation and handle various graph learning tasks, such as node classification.

[0003] However, most current GNNs models are implicitly trained based on the complete data state, assuming that the graph data can provide sufficient information. Due to the diverse states of real-world data distributions, nodes are inherently class-imbalanced, and this implicit assumption does not hold in many practical cases. For example, the number of high-risk enterprises in the financial knowledge network is much less than the number of normal enterprises. When the number of samples in the majority class in the training set is much larger than that in the minority class, researchers believe that machine learning models are prone to insufficient representation of the minority class, resulting in poor overall classification results. Therefore, this imbalance inevitably hinders the training of expressive graph neural network models on these graph data with relatively limited information. How to balance both the majority class and the minority class samples in the node classification task of graph-structured data is the key to the research and implementation of class-imbalanced graph data node classification methods. Summary of the Invention

[0004] To solve the above technical problems existing in the prior art, the present invention proposes an unbalanced graph node classification method and system based on boundary node condition GAN, which takes class information and context topological information as conditions to input into the conditional GAN for conditional adversarial training, making full use of the rich topological information in graph-structured data that is ignored by traditional class-imbalanced learning methods to promote the learning of the classifier.

[0005] On the one hand, to achieve the above object, the present invention provides an unbalanced graph node classification method based on boundary node condition GAN, including:

[0006] Construct a graph-structured data set to be classified;

[0007] Input the graph structure data set to be classified into a graph convolutional network for node classification for processing, and output the classification result; wherein, the graph convolutional network for node classification is trained by a training data set and obtained according to the convolutional layer parameters obtained by training, and the training set is a graph structure data set.

[0008] Preferably, obtaining the graph convolutional network for node classification includes:

[0009] Input the training data set into a neural network GCN structure including two layers of graph convolution to obtain the embedding vectors of the original nodes in the data set;

[0010] Calculate the uncertainty scores of each node according to the embedding vectors of each node, and perform balance calibration to obtain the node misclassification risk rate, and obtain the decision boundary node set based on the node misclassification risk rate;

[0011] Input the boundary nodes into the conditional generator, and use a multi-layer perceptron to encode the input nodes according to the input conditional labels and structures, and output synthetic nodes;

[0012] Input the original node embeddings and synthetic nodes into the conditional discriminator, and use a multi-layer perceptron to perform conditional encoding on the input synthetic nodes and original nodes respectively according to the input conditional labels and structures, and output two discriminator feedback values reflecting the true and false degrees of the synthetic nodes and original nodes;

[0013] Set the model optimizer, calculate the classification loss of the GCN; calculate the target loss function of the generator, and update the parameters of the conditional generator according to the parameter gradient calculated backward from the target loss function of the generator; calculate the target loss function of the conditional discriminator, and update the conditional discriminator according to the parameter gradient calculated backward from the target loss function of the conditional discriminator, and combine the parameter gradient calculated backward from the classification loss to update the parameters of the GCN convolutional layer;

[0014] Obtain the trained graph convolutional network for node classification according to the convolutional layer parameters obtained by training, that is, the graph convolutional network for node classification.

[0015] Preferably, obtaining the embedding vectors of the original nodes includes:

[0016] Input the graph structure data set into the first graph convolutional layer of the GCN to obtain the first graph convolutional vector;

[0017] Input the first graph convolutional vector into the activation function to obtain the activation vector;

[0018] Input the activation vector into the second graph convolutional layer to obtain the embedding vectors of the original nodes.

[0019] Preferably, obtaining the decision boundary node set includes:

[0020] According to the embedding vectors of each original node, calculate the node uncertainty score through the Kullback-Leibler divergence, and perform balanced calibration on the node uncertainty score to obtain the final misclassification risk rate of the node;

[0021] Through the final misclassification risk rate of the node, obtain the top K% nodes to get the decision boundary node set;

[0022] Among them, the calculation of the node uncertainty score is:

[0023]

[0024] In the formula, o v is the embedding vector of node v, o v (j) = P(y v = C j |G); C is the category label distribution set, C j is the node set of the j-th class, |C| is the total number of categories; U v is the uncertainty score of node v, D KL represents the function calculated by KL, represents the value of the category in the single-point distribution in, is the predicted label of node v;

[0025] The final misclassification risk rate of the node is obtained as:

[0026]

[0027] In the formula, r v is the final misclassification risk rate, is the number of nodes of class in the training set; R imb is the imbalance ratio of the training set.

[0028] Preferably, outputting the synthetic node includes:

[0029] Map the conditional label to a conditional vector one through one-hot, and integrate the context information of the node as a conditional structure and map it to a conditional vector two;

[0030] Obtain noise z through Gaussian perturbation, concatenate the noise z with the conditional vector one and the conditional vector two, and perform learning through a standard multi-layer perceptron to convert it into a fake sample similar to the real one, and output the synthetic node;

[0031] Among them, the conditional structure information TI of the node v obtains the formal expression as:

[0032]

[0033] In the formula, N(v) represents the neighborhood nodes of node v, and O u represents the original node embedding of node u.

[0034] Preferably, obtaining the discriminator feedback value includes:

[0035] Mapping the synthetic sample and the original node embedding output by the conditional generator and the conditional label corresponding to the original node to vectors h g and h r ;

[0036] Performing weighted summation using convolution and outputting the discriminator feedback value that respectively reflects the true and false degree of the synthetic sample g output by the conditional generator and the real node.

[0037] Preferably, the target loss function of the generator is:

[0038]

[0039] In the formula, L G is the target loss function of the conditional generator, V R is the set of decision boundary nodes, Z i is the noise vector of the i-th node, y i is the conditional label of the i-th node, D(G(z i |(y i , TI i )|y i )) is the true discrimination probability of the discriminator for the generated node;

[0040] The target loss function of the conditional discriminator is:

[0041]

[0042] In the formula, L D is the target loss function of the conditional discriminator, D(x i |y i ) is the discrimination probability of the discriminator for the real node being the category y i , 1 - D(G(z i |(y i , TI i )|y i )) is the non-true discrimination probability of the discriminator for the generated node;

[0043] The loss of the GCN is:

[0044] L = αL gcn +(1 - α)L D;

[0045] Wherein, L is the GCN loss, L gcn is the classification loss of GCN, and α is the weight value.

[0046] On the other hand, to achieve the above object, the present invention also provides an unbalanced graph node classification system based on boundary node conditional GAN, including:

[0047] Data acquisition module: used to construct a graph structure data set to be classified;

[0048] Model training module: used to input the graph structure data set to be classified into a graph convolutional network for node classification for processing, and output a classification result; wherein, the graph convolutional network for node classification is trained by a training data set and obtained according to the convolutional layer parameters obtained by training, and the training set is a graph structure data set.

[0049] Preferably, the model training module includes:

[0050] Boundary node evaluation unit: used to calculate the uncertainty score of each node according to the embedding vector of each node, perform balance calibration, obtain the node misclassification risk rate, and obtain a decision boundary node set based on the node misclassification risk rate;

[0051] Conditional generator unit: used to input boundary nodes into the conditional generator, and use a multi-layer perceptron to encode the input nodes according to the input conditional labels and structures, and output synthetic nodes;

[0052] Conditional discriminator unit: used to input the original node embedding and synthetic nodes into the conditional discriminator, and use a multi-layer perceptron to perform conditional encoding on the input synthetic nodes and real nodes respectively according to the input conditional labels and structures, input them into the discriminator, and obtain a feedback value;

[0053] Model update unit: used to update the conditional generative adversarial network model based on the feedback value until the loss function converges;

[0054] Model prediction unit: used to perform classification prediction on the graph structure data set to be classified according to the trained classification model.

[0055] The present invention also provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor, and when the program or instruction is executed by the processor, the steps of the unbalanced graph node classification method based on boundary node conditional GAN are implemented.

[0056] Compared with the prior art, the present invention has the following advantages and technical effects:

[0057] Compared with traditional class-imbalanced learning methods, the method proposed in the present invention is based on a graph neural network. Class information and context topology information are used as conditions to input into a conditional GAN for conditional adversarial training, making full use of the rich topology information in graph-structured data that traditional class-imbalanced learning methods ignore, and promoting the learning of the classifier. The present invention applies an adversarial network training process to class decision boundary nodes, expands the distance between classes, enables the graph neural network model to better learn node embeddings, and overcomes the problem of poor overall classification results caused by insufficient embedding representation of minority classes without introducing noise nodes into the original graph. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0059] Figure 1 is a flowchart of an imbalanced graph node classification method based on boundary node conditional GAN according to an embodiment of the present invention;

[0060] Figure 2 is a framework diagram of a conditional generator according to an embodiment of the present invention;

[0061] Figure 3 is a framework diagram of a conditional discriminator according to an embodiment of the present invention;

[0062] Figure 4 is a framework diagram of a model update according to an embodiment of the present invention;

[0063] Figure 5 is a schematic structural diagram of an imbalanced graph node classification system based on boundary node conditional GAN according to an embodiment of the present invention;

[0064] Figure 6 is a result block diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0066] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0067] The present invention proposes an imbalanced graph node classification method based on boundary node conditional GAN, as Figure 1 , including:

[0068] Construct a graph structure data set to be classified;

[0069] Input the graph structure data set to be classified into a graph convolutional network for node classification for processing, and output a classification result; wherein, the graph convolutional network for node classification is trained through a training data set and obtained according to the convolutional layer parameters obtained by training, and the training set is a graph structure data set.

[0070] Further, obtaining the graph convolutional network for node classification includes:

[0071] Step 1: Obtain source graph data (i.e., the training data set), and construct a graph structure data set G=(V, E, X, C); where V={v 1 , v 2 , …, v n} is a set containing n nodes; E={e 1,2 , e 1,3 , …, e n-1,n} is a set of edges that can be equivalent to an n×n adjacency matrix A. If e i,j ∈E, then A ij = 1, otherwise A ij = 0; is a matrix containing n nodes and their associated features. The graph G has multiple classes, and these classes divide G into |C| clusters, represented by C={C 1 , C 2 , …, C |C|}. Nodes in each cluster have the same label, and the node label information in G is represented as Y. The class distribution may be highly skewed because one or more classes contain more nodes than other classes, i.e., |C 1 | >> |C 2 |. In this case, C 1 belongs to the majority class, and C 2 belongs to the minority class.

[0072] Step 2: Input the data of G into a two-layer GCN to obtain the embedding vector O of the original nodes;

[0073] Step 3: For each node, calculate its uncertainty score and perform balance calibration according to the original embedding vector, obtain the node misclassification risk rate, and obtain the decision boundary node set;

[0074] Step 4: Input the boundary node set into a conditional generator. The conditional generator uses a multi-layer perceptron to encode the input nodes according to the input conditional labels and structures, and outputs the conditional generated embedding of the synthetic node g, so that the synthetic node g has the style features represented by the conditional label y' and the structural features of the input nodes;

[0075] Step 5: Input the original embedding of the original nodes and the conditional generated embedding of the synthetic nodes into the conditional discriminator. The conditional discriminator uses a multi-layer perceptron to perform conditional encoding on the input synthetic node g and real node v respectively according to the input conditional label and structure, and inputs two discriminator feedback values reflecting the true and false degrees of the synthetic node g and real node v.

[0076] Step 6: Set up a model update module to calculate the classification loss of the GCN; calculate the objective loss function of the generator, and update the parameters of the conditional generator according to the parameter gradients calculated backward from the loss function; calculate the objective loss function of the conditional discriminator, and update the parameters of the conditional discriminator and the GCN convolutional layer according to the parameter gradients calculated backward from the loss function.

[0077] Step 7: Obtain the GCN for node classification according to the trained convolutional layer parameters, and input the data to be classified into the GCN to get the prediction of the data to be classified.

[0078] Specifically, obtaining the embedding vector of the original nodes includes:

[0079] Input the graph structure dataset into the first graph convolutional layer of the GCN to obtain the first graph convolutional vector;

[0080] Input the first graph convolutional vector into the activation function to get the activation vector;

[0081] Input the activation vector into the second graph convolutional layer to obtain the embedding vector of the original nodes.

[0082] The calculation method is as follows:

[0083]

[0084] Among them, is the feature transformation matrix, X is the node feature matrix, W 0 and W 1 are the trainable weight matrices of the first and second layers respectively, σ is the activation function; A is the adjacency matrix, I is the identity matrix, and D is the degree matrix of A.

[0085] Specifically, in step 3, obtaining the decision boundary node set includes:

[0086] According to the embedding vector of each original node, calculate the node uncertainty score through the Kullback-Leibler divergence, and perform balance calibration on the node uncertainty score to obtain the final misclassification risk rate of the node;

[0087] Obtain the top K% nodes through the final misclassification risk rate of the node to get the decision boundary node set; where K is a hyperparameter.

[0088] Among them, calculating the uncertainty score of the node is as follows:

[0089]

[0090] In the formula, o v is the embedding vector of node v, o v (j) = P(y v = C j |G); C is the set of class label distributions, C j is the set of nodes of the j-th class, |C| is the total number of classes; U v is the uncertainty score of node v, D KL represents the function for KL calculation, represents the single-point distribution of the target class , represents the single-point distribution in the value of the class, is the predicted label of node v;

[0091] Obtaining the final misclassification risk rate of the node is as follows:

[0092]

[0093] In the formula, r v is the final misclassification risk rate, is the number of nodes of class imb in the training set; R

[0094] Step 4 specifically includes:

[0095] Step 4.1: Map the conditional label y to a conditional vector one through one-hot, and integrate the node context information as a conditional structure to map to a conditional vector two;

[0096] Step 4.2: Obtain the noise z through Gaussian perturbation, concatenate z with the conditional vector one and the conditional vector two, and after learning through a standard multi-layer perceptron, convert it into a pseudo-sample similar to the real one.

[0097] Specifically, referring to Figure 2 , in the conditional generator based on boundary nodes, map the label y v of the boundary node v to a conditional vector one: one-hot(y v ), and integrate the context information of the boundary node v as conditional structure information to map to a conditional vector two: TI v .

[0098] TI v The calculation method is as follows:

[0099]

[0100] Among them, N(v) represents the neighboring nodes of node v, and O u represents the original node embedding of node u;

[0101] Then, noise z is obtained according to Gaussian perturbation. After concatenating z with conditional vector one and conditional vector two, a realistic synthetic node is obtained through learning by a standard multi-layer perceptron as a fake sample. The formal expression for obtaining the fake sample for the input boundary node v is as follows:

[0102] g v = MLP(z || one-hot(y v ) || TI v ).

[0103] Obtaining the feedback value in step 5 includes:

[0104] Mapping the synthetic sample output by the conditional generator and the original node embedding corresponding to the original node to vectors h g and h r respectively;

[0105] Performing weighted summation using convolution and outputting the discriminator feedback values respectively reflecting the authenticity of the synthetic sample g output by the conditional generator and the real node.

[0106] Specifically, referring to Figure 3 , in the conditional discriminator based on the boundary node, mapping the fake sample g v obtained from the input boundary node v and the embedding vector o v of node v to vectors h g and h r respectively with the corresponding conditional label y; Performing weighted summation on h g and h r through one layer of convolution and outputting the discriminator feedback value of the authenticity of the node.

[0107] Furthermore, in step 6, it includes:

[0108] Setting the model optimizer, calculating the classification loss of the GCN; calculating the target loss function of the generator, and updating the parameters of the conditional generator according to the parameter gradient calculated backward from the target loss function of the generator; calculating the target loss function of the conditional discriminator, and updating the conditional discriminator according to the parameter gradient calculated backward from the target loss function of the conditional discriminator; updating the parameters of the GCN convolutional layer according to the parameter gradient calculated backward from the classification loss of the GCN.

[0109] Specifically, referring to Figure 4, in the model updater, the conditional generator hopes that the generated fake samples are close to the real data distribution and match the given class labels. Therefore, the loss function L of the conditional generator is calculated. G , and the parameters of the conditional generator are updated using the parameter gradients calculated by backpropagating the calculated loss function.

[0110] L G The calculation method is as follows:

[0111]

[0112] In the formula, L G is the target loss function of the conditional generator, V R is the set of decision boundary nodes, z i is the noise vector of the i-th node, y i is the conditional label of the i-th node, D(G(z i |(y i , TI i )|y i ) is the true discrimination probability of the discriminator for the generated nodes;

[0113] The conditional discriminator needs to classify the fake samples and real samples generated by the conditional generator. A well-trained classifier is required to distinguish between fake and real samples, and the loss function L of the conditional discriminator is calculated. D , and the parameters of the conditional discriminator are updated using the parameter gradients calculated by backpropagating its loss function. L D The calculation method is as follows;

[0114]

[0115] Among them, D(x i |y i ) is the discrimination probability of the discriminator for the real node being class y i , and 1 - D(G(z i |(y i , TI i )|y i ) is the "non-real" discrimination probability of the discriminator for the generated nodes.

[0116] GCN is the basic model for graph node classification. A well-trained classifier is required to distinguish between different classes of samples. In this embodiment, the expressive ability of GCN is strengthened in reverse through the generative adversarial training between the conditional generator and the conditional discriminator. Therefore, the classification loss L of GCN is calculated. gcn , and L gcn and L D are combined to form the GCN loss L, and its formal expression is as shown in the formula. The parameters of the GCN convolutional layer are updated using the parameter gradients calculated by backpropagating L.

[0117]

[0118] L = αL gcn +(1 - α)L D ;

[0119] Fix the parameters of the GCN model obtained through the above training, and input the data to be classified into the fixed GCN to obtain the classification results of the data to be classified.

[0120] This embodiment also provides an unbalanced graph node classification system based on boundary node conditional GAN, as Figure 5 , including:

[0121] Data acquisition module: used to construct a graph structure data set to be classified;

[0122] Model training module: used to input the graph structure data set to be classified into a graph convolutional network for node classification for processing, and output classification results; wherein, the graph convolutional network for node classification is trained through a training data set and obtained according to the convolutional layer parameters obtained by training, and the training set is a graph structure data set.

[0123] The model training module includes:

[0124] Boundary node evaluation unit: used to calculate the uncertainty score of each node according to the embedding vector of each node, and perform balance calibration to obtain the node misclassification risk rate, and obtain the decision boundary node set based on the node misclassification risk rate;

[0125] Conditional generator unit: used to input the boundary nodes into the conditional generator, and use a multi-layer perceptron to encode the input nodes according to the input conditional labels and structures, and output synthetic nodes;

[0126] Conditional discriminator unit: used to input the original node embeddings and synthetic nodes into the conditional discriminator, and use a multi-layer perceptron to perform conditional encoding on the input synthetic nodes and real nodes respectively according to the input conditional labels and structures, and input them into the discriminator to obtain feedback values;

[0127] Model update unit: used to update the conditional generative adversarial network model based on the feedback value until the loss function converges;

[0128] Model prediction unit: used to classify and predict the graph structure data set to be classified according to the trained classification model.

[0129] This embodiment provides an electronic device, as Figure 6 , including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the unbalanced graph node classification method based on boundary node conditional GAN are implemented.

[0130] The following is an example of implementing this technical solution in terms of financial risk classification:

[0131] Example 1

[0132] An unbalanced graph node classification method based on boundary node condition GAN, running on an electronic device and applied to financial risk classification, to solve the risk classification problem of uneven distribution of financial entity risk categories includes the following steps:

[0133] Step 1: Obtain the source data of financial risk graphs, obtain financial entities, including companies, individuals (natural persons), regulatory agencies, industries, etc.; obtain financial relationships, including company-company relationships, company-industry relationships, company-person relationships, etc.; obtain calculation features, including market data, financial statements, etc., which are relatively high-frequency and are attributes used for statistical calculations. Construct the financial entities, relationships, and features into the form of graph data. Construct a graph structure dataset G=(V, E, X, C); where V={v 1 ,v 2 ,…,v n} is a set containing n nodes; E={e 1,2 ,e 1,3 ,…,e n-1,n} is a set of edges that can be equivalent to the adjacency matrix A of n×n. If e i,j ∈E, then A ij =1, otherwise A ij =0; is a matrix containing n nodes and their associated features. The graph G has multiple classes, and these classes divide G into |C| clusters, represented by C={C 1 ,C 2 ,…,C |C|}. The nodes in each cluster have the same label, and the node label information in G is represented as Y. The class distribution may be highly skewed because one or more classes contain more nodes than other classes, that is, |C 1 | >> ||C 2 |. In this case, C 1 belongs to the majority class, and C 2 belongs to the minority class.

[0134] Step 2: Input the data of G into a two-layer GCN to obtain the embedding vectors of the original nodes;

[0135] Step 3: For each node, calculate its uncertainty score and perform balance calibration according to the original embedding vector, obtain the node misclassification risk rate, and obtain the decision boundary node set;

[0136] Step 4: Input the boundary node set into the condition generator. The condition generator uses a multi-layer perceptron to encode the input nodes according to the input condition labels and structures, and outputs the conditional generation embedding of the synthetic node g, so that the synthetic node g has the style features represented by the conditional label y' and the structural features of the input nodes;

[0137] Step 5: Input the original embedding of the original nodes and the conditional generation embedding of the synthetic nodes into the condition discriminator. The condition discriminator uses a multi-layer perceptron to perform conditional encoding on the input synthetic node g and the real node v respectively according to the input condition labels and structures, and inputs two discriminator feedback values reflecting the true and false degrees of the synthetic node g and the real node v;

[0138] Step 6: Set up a model update module to calculate the classification loss of the GCN; calculate the objective loss function of the generator, and update the parameters of the condition generator according to the parameter gradients calculated backward from the loss function; calculate the objective loss function of the condition discriminator, and update the parameters of the condition discriminator and the GCN convolutional layer according to the parameter gradients calculated backward from the loss function;

[0139] Step 7: Obtain the GCN for node classification according to the trained convolutional layer parameters, and input the data to be classified into the GCN to obtain the prediction of the data to be classified.

[0140] Specifically, obtaining the embedding vector of the original nodes includes:

[0141] Input the graph structure data set into the first graph convolutional layer of the GCN to obtain the first graph convolutional vector;

[0142] Input the first graph convolutional vector into the activation function to obtain the activation vector;

[0143] Input the activation vector into the second graph convolutional layer to obtain the embedding vector of the original nodes.

[0144] The calculation method is as follows:

[0145]

[0146] Among them, is the feature transformation matrix, X is the node feature matrix, W 0 、W 1 are the trainable weight matrices of the first and second layers respectively, σ is the activation function; A is the adjacency matrix, I is the identity matrix, and D is the degree matrix of A.

[0147] Specifically, in step 3, obtaining the decision boundary node set includes:

[0148] Based on the embedding vectors of each original node, calculate the node uncertainty score through the Kullback-Leibler divergence, and perform balanced calibration on the node uncertainty score to obtain the final misclassification risk rate of the node;

[0149] Through the final misclassification risk rate of the node, obtain the top K% nodes to get the decision boundary node set; where K is a hyperparameter.

[0150] Among them, the calculation of the node uncertainty score is:

[0151]

[0152] In the formula, o v is the embedding vector of node v, o v (j) =P(y v =C j |G); is the predicted label of node v, C is the set of class label distributions, C j is the set of nodes of the j-th class, |C| is the total number of classes; U v is the uncertainty score of node v, D KL represents the function for KL calculation. In this embodiment, for the class its value is 1, and for other classes, its value is 0; represents the value of the class in the single-point distribution ;

[0153] The final misclassification risk rate of the node is obtained as:

[0154]

[0155] In the formula, r v is the final misclassification risk rate, is the number of nodes of class imb in the training set; R

[0156] Step 4 specifically includes:

[0157] Step 4.1: Map the conditional label y into a conditional vector one through one-hot, and integrate the node context information as a conditional structure and map it into a conditional vector two;

[0158] Step 4.2: Obtain the noise z through Gaussian perturbation, concatenate z with the conditional vector one and the conditional vector two, and after learning through a standard multi-layer perceptron, convert it into a fake sample similar to the real one.

[0159] Specifically, referring to Figure 2 , in the conditional generator based on the boundary nodes, the label y of the boundary node vv Map it to conditional vector one through one - hot: one - hot(y v ), integrate the context information of boundary node v as conditional structure information and map it to conditional vector two: TI v .

[0160] TI v The calculation method is as follows:

[0161]

[0162] Among them, N(v) represents the neighborhood nodes of node v, o u represents the original node embedding of node u;

[0163] Then obtain noise z according to Gaussian perturbation, concatenate z with conditional vector one and conditional vector two, and obtain a synthetic node similar to the real one through learning by a standard multi - layer perceptron as a fake sample. The formal expression for obtaining the fake sample for the input boundary node v is as follows:

[0164] g v = MLP(z||one - hot(y v )||TI v ).

[0165] The steps to obtain the feedback value in step 5 include:

[0166] Map the synthetic sample output by the conditional generator and the original node embedding corresponding to the original node to vectors h g and h r respectively with the conditional label corresponding to the original node;

[0167] Use convolution for weighted summation and output the discriminator feedback values respectively reflecting the authenticity of the synthetic sample g output by the conditional generator and the real node.

[0168] Specifically, in the conditional discriminator based on boundary nodes, map the fake sample g v obtained from the input boundary node v and the embedding vector o v of node v to vectors h g and h r respectively with the corresponding conditional label y; perform weighted summation through one - layer convolution on h g and h r and output the discriminator feedback value of the authenticity of the node.

[0169] Furthermore, step 6 includes:

[0170] Set up a model optimizer to calculate the classification loss of the GCN; calculate the objective loss function of the generator, and update the parameters of the conditional generator according to the parameter gradients calculated backward from the objective loss function of the generator; calculate the objective loss function of the conditional discriminator, and update the conditional discriminator according to the parameter gradients calculated backward from the objective loss function of the conditional discriminator; update the parameters of the GCN convolutional layer according to the parameter gradients calculated backward from the classification loss of the GCN.

[0171] Specifically, referring to Figure 4 , in the model updater, the conditional generator hopes that the fake samples generated are close to the real data distribution and match the given class labels. Therefore, calculate the loss function L G of the conditional generator, and update the parameters of the conditional generator according to the parameter gradients calculated backward from the calculated loss function.

[0172] L G is calculated as follows:

[0173]

[0174] In the formula, L G is the objective loss function of the conditional generator, V R is the set of decision boundary nodes, z i is the noise vector of the i-th node, y i is the conditional label of the i-th node, D(G(z i | (y i , TI i ) | y i ) is the true discrimination probability of the discriminator for the generated nodes;

[0175] The conditional discriminator needs to classify the fake samples and real samples generated by the conditional generator. A well-trained classifier is needed to distinguish between real and fake samples. Calculate the loss function L D of the conditional discriminator, and update the parameters of the conditional discriminator according to the parameter gradients calculated backward from its loss function. L D is calculated as follows;

[0176]

[0177] Among them, D(x i | y i ) is the discrimination probability of the discriminator for the real node being the class y i , 1 - D(G(z i | (y i , TI i ) | y i ) is the "non-real" discrimination probability of the discriminator for the generated nodes.

[0178] GCN is a basic model for graph node classification, and a well-trained classifier is required to distinguish samples of different categories. In this embodiment, the expressive ability of GCN is enhanced in reverse through the generative adversarial training between the conditional generator and the conditional discriminator. Therefore, the classification loss L of GCN is calculated gcn , and L gcn and L D are combined to form the GCN loss L, and its formal expression is as shown in the formula. The parameters of the GCN convolutional layer are updated using the parameter gradients calculated in reverse by L.

[0179]

[0180] L = αL gcn +(1 - α)L D ;

[0181] Fix the parameters of the GCN model obtained through the above training, and input the data to be classified into the fixed GCN to obtain the classification results of the data to be classified.

[0182] Embodiment 2

[0183] An unbalanced graph node classification system based on boundary node conditional GAN runs on an electronic device and is applied to financial risk classification to solve the risk classification problem of uneven distribution of financial entity risk categories:

[0184] A data acquisition module is used to acquire the financial risk graph data set to be classified and preprocess the data; in this embodiment, financial entities are acquired, including companies, individuals (natural persons), regulatory agencies, industries, etc.; financial relationships are acquired, including company-company relationships, company-industry relationships, company-person relationships, etc.; calculation features are acquired, including market data, financial statements, etc., which are relatively high-frequency and are attributes used for statistical calculations. The financial entities, relationships, and features are constructed in the form of graph data.

[0185] The boundary node evaluation unit is used to evaluate the risk categories of misclassified financial entities and determine the decision boundary nodes between different categories;

[0186] The conditional generator unit is used to generate synthetic nodes of corresponding categories based on the boundary node categories and context structures;

[0187] The conditional discriminator unit is used to discriminate between the original nodes and the synthetic nodes, and inputs two discriminator feedback values reflecting the true and false degrees of the synthetic node g and the real node v;

[0188] The model update unit is used to update the conditional generative adversarial network model based on the feedback value until the loss function converges;

[0189] A model prediction unit is used for classifying and predicting a financial entity to be classified. According to the trained classification model, it predicts the category of the financial entity to be classified.

[0190] This embodiment provides an electronic device that can execute the foregoing method, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements the steps of the unbalanced graph node classification method based on the boundary node condition GAN.

[0191] In summary, the present application provides an unbalanced graph node classification method and system based on the boundary node condition GAN. The original embedding vectors of the original nodes are obtained through a two-layer GCN; the final misclassification risk rate of the nodes after node balance calibration is calculated based on the original embedding vectors; the nodes with a high misclassification risk rate are trained by a generative adversarial network based on conditional labels and conditional structures; the conditional discriminator and the conditional generator are updated until the loss function converges; the expression ability of the GCN model is strengthened in reverse based on the generative adversarial training. In the above manner, a GCN classifier that can effectively handle the imbalance of the category distribution is constructed.

[0192] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A node classification method for unbalanced graphs based on boundary node conditional GAN, characterized in that: include: Construct a graph structure dataset to be classified; The graph structure data set to be classified is input into a graph convolutional network for node classification for processing, and a classification result is output; wherein the graph convolutional network for node classification is trained by a training data set, and the convolutional layer parameters obtained by the training are obtained, and the training set is a graph structure data set.

2. The unbalanced graph node classification method based on boundary node condition GAN according to claim 1 is characterized in that: Obtaining the graph convolutional network for node classification, including: Input the training data set into the neural network GCN structure containing two layers of graph convolution to obtain the embedding vectors of the original nodes in the data set; Calculate the uncertainty score of each node according to the embedding vector of each node, perform balance calibration, obtain the node misclassification risk rate, and obtain the decision boundary node set based on the node misclassification risk rate; Input the boundary nodes into the condition generator, use the multi-layer perceptron to encode the input nodes according to the input condition labels and structures, and output the synthesized nodes; The original node embedding and the synthesized node are input into the conditional discriminator, and the input synthesized node and the original node are conditionally encoded according to the input conditional label and structure using a multi-layer perceptron, and two discriminator feedback values ​​reflecting the truth or falsehood of the synthesized node and the original node are output; Set up a model optimizer to calculate the classification loss of GCN; calculate the target loss function of the generator, and update the parameters of the conditional generator according to the parameter gradient calculated by the reverse calculation of the target loss function of the generator; calculate the target loss function of the conditional discriminator, and update the conditional discriminator according to the parameter gradient calculated by the reverse calculation of the target loss function of the conditional discriminator, and update the parameters of the GCN convolution layer in combination with the parameter gradient calculated by the reverse calculation of the classification loss; According to the convolutional layer parameters obtained through training, a trained graph convolutional network for node classification is obtained, that is, the graph convolutional network for node classification.

3. The unbalanced graph node classification method based on boundary node condition GAN according to claim 2 is characterized in that: Obtaining the embedding vector of the original node includes: Input the graph structure data set into the first graph convolution layer of GCN to obtain a first graph convolution vector; Inputting the first graph convolution vector into an activation function to obtain an activation vector; The activation vector is input into the second graph convolution layer to obtain the embedding vector of the original node.

4. The unbalanced graph node classification method based on boundary node condition GAN according to claim 3 is characterized in that: Obtaining the decision boundary node set includes: According to the embedding vector of each original node, the node uncertainty score is calculated by Kullback-Leibler divergence, and the node uncertainty score is balanced and calibrated to obtain the final misclassification risk rate of the node; The first K% nodes are obtained through the final misclassification risk rate of the nodes to obtain the decision boundary node set; The node uncertainty score is calculated as: In the formula, o v is the embedding vector of node v, o v (j) =P(y v =C j |G); C is the category label distribution set, C j is the node set of the jth class, |C| is the total number of classes; U v is the uncertainty score of node v, D KL represents the function of KL calculation, Represents a single point distribution The value of the category in is the predicted label of node v; The final misclassification risk rate of the node is obtained as follows: In the formula, r v is the final misclassification risk rate, for The number of class nodes in the training set; R imb is the imbalance ratio of the training set.

5. The unbalanced graph node classification method based on boundary node condition GAN according to claim 4 is characterized in that: Output the synthesis node, including: The conditional label is mapped into a conditional vector 1 through one-hot mapping, and the contextual information of the node is integrated as a conditional structure and mapped into a conditional vector 2; Obtain noise z through Gaussian perturbation, connect the noise z with the conditional vector 1 and the conditional vector 2 in series, and then learn through a standard multi-layer perceptron to convert it into a fake sample similar to the real one, and output the synthetic node; Among them, the conditional structure information TI of the node v Get the formal expression as: Where N(v) represents the neighboring nodes of node v, O u represents the original node embedding of node u.

6. The unbalanced graph node classification method based on boundary node condition GAN according to claim 5 is characterized in that: Obtaining the discriminator feedback value includes: The synthetic samples and original nodes output by the condition generator are respectively embedded into the conditional labels corresponding to the original nodes and mapped into vectors h g and h r ; Convolution is used to perform weighted summation, and the discriminator feedback value that reflects the degree of authenticity of the synthetic sample g output by the conditional generator and the real node is output.

7. The unbalanced graph node classification method based on boundary node condition GAN according to claim 2 is characterized in that: The objective loss function of the generator is: Where, L G is the target loss function of the conditional generator, V R is the set of decision boundary nodes, Z i is the noise vector of the ith node, y i is the conditional label of the i-th node, D(G(z i |(y i ,TI i )|y i )) is the true discrimination probability of the discriminator for the generated node; The objective loss function of the conditional discriminator is: Where, L D is the target loss function of the conditional discriminator, D(x i |y i ) is the discriminator for the real node as category y i The discriminant probability, 1-D(G(z i |(y i ,TI i )|y i )) is the probability of the discriminator discriminating the generated node as not being true; The loss of the GCN is: L=αL gcn +(1-α)L D ; Where L is the GCN loss, L gcn is the classification loss of GCN, and α is the weight value.

8. An unbalanced graph node classification system based on boundary node condition GAN, characterized in that: include: Data acquisition module: used to construct the graph structure data set to be classified; Model training module: used to input the graph structure data set to be classified into the graph convolution network for node classification for processing, and output the classification result; wherein, the graph convolution network for node classification is trained by the training data set, and the convolution layer parameters are obtained according to the training, and the training set is the graph structure data set.

9. The unbalanced graph node classification system based on boundary node condition GAN according to claim 8, characterized in that: The model training module includes: Boundary node evaluation unit: used to calculate the uncertainty score of each node according to the embedding vector of each node, perform balance calibration, obtain the node misclassification risk rate, and obtain the decision boundary node set based on the node misclassification risk rate; Condition generator unit: used to input boundary nodes into the condition generator, use the multi-layer perceptron to encode the input nodes according to the input condition labels and structures, and output the synthesized nodes; Conditional discriminator unit: used to embed the original node and the synthetic node into the conditional discriminator, use the multi-layer perceptron to conditionally encode the input synthetic node and real node according to the input conditional label and structure, input the discriminator, and obtain the feedback value; A model updating unit: used to update the conditional generative adversarial network model based on the feedback value until the loss function converges; Model prediction unit: used to perform classification prediction on the graph structure dataset to be classified based on the trained classification model.

10. An electronic device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein when the program or instruction is executed by the processor, the steps of the unbalanced graph node classification method based on boundary node condition GAN are implemented as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for semi-supervised learning of structured data

    CN109977094A

  • Imbalanced graph node neural network classification method based on graph data enhancement

    CN116756391A

  • System and method for structure learning for graph neural networks

    US20220101103A1

  • Scalable Self-Supervised Graph Clustering

    US20240176993A1

  • Classifier training using synthetic training data samples

    US20240256967A1

Cited By

  • Image classification adversarial training improvement method based on boundary sample enhancement

    CN121280812A

  • An image classification adversarial training improvement method based on boundary sample augmentation

    CN121280812B