Large-scale graph node classification model training method and system, computer equipment and computer program product

By cutting large-scale graph data and using invariance loss functions, training invariance promotes model IFM solves the problem of high computing resources requirements for large-scale graph node classification models, and achieves efficient training and generalization performance improvement under low computing resource conditions.

CN120219791APending Publication Date: 2025-06-27GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510180465.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, the large-scale graph node classification model has high requirements for computing resources, which limits its deployment in actual scenarios.

Method used

By cutting the target large-scale graph data, the target sub-graph set is obtained, and then the invariance promotion model IFM is trained based on the target sub-graph set. The loss is calculated using the invariance loss function and backpropagated to update the model parameters until the preset iteration training termination condition is reached.

Benefits of technology

The training of large-scale graph node classification model under low computing resource conditions is realized, which reduces the model's requirements for computing resources, and improves the generalization performance and interpretability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219791A_ABST
    Figure CN120219791A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of graph data processing, and provides a large-scale graph node classification model training method and system, computer equipment and a computer program product, and the method comprises the steps: cutting target large-scale graph data to obtain a target sub-graph set, and training an invariance promotion model (IFM) based on the target sub-graph set; calculating loss based on an invariance loss function and performing back propagation so as to update an invariance promotion model IFM parameter; returning to the step of training the IFM until a preset iterative training termination condition is reached; and outputting an invariance promotion model (IFM). According to the method, the target large-scale graph data is cut, and the iterative training is performed based on the IFM, so that the training of the large-scale graph node classification model under the condition of relatively low computing resources is realized, and the requirement of the IFM obtained by training on the computing resources is lower.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of graph data processing, and particularly relates to a training method, system, computer device, and computer program product for a large-scale graph node classification model. Background Art

[0002] Graph-structured data widely exists in complex systems in different fields, from social and biological networks to transportation and communication frameworks. The ubiquitous nature of graph-structured data has greatly promoted the application of graph neural networks (GNNs) and improved the modeling and analysis capabilities in many fields. However, the challenge of modeling large-scale graphs lies in their huge demand for computing resources, which limits the deployment of GNNs in many practical scenarios and becomes a major obstacle to fully realizing their potential in real-world applications.

[0003] Therefore, there is an urgent need for a training method for a large-scale graph node classification model that requires less computing resources. Summary of the Invention

[0004] In view of this, embodiments of this application provide a training method, system, computer device, and computer program product for a large-scale graph node classification model to solve the problem that the large-scale graph node classification model in the prior art requires high computing resources.

[0005] The first aspect of the embodiments of this application provides a training method for a large-scale graph node classification model, including:

[0006] Cut the target large-scale graph data to obtain a target sub-graph set, where the target sub-graph set includes multiple target sub-graphs;

[0007] Train an invariance promotion model IFM based on the target sub-graph set;

[0008] Calculate the loss based on an invariance loss function and perform backpropagation to update the parameters of the invariance promotion model IFM;

[0009] Return to the step of training the invariance promotion model IFM based on the target sub-graph set until a preset iterative training termination condition is reached;

[0010] Output the invariance promotion model IFM, where the invariance promotion model IFM is used to classify large-scale graph nodes.

[0011] In one implementation manner of the first aspect, before training the invariance promotion model IFM based on the target sub-graph set, it further includes:

[0012] Divide the target sub-graph set into a training sub-graph set, a validation sub-graph set, and a test sub-graph set;

[0013] Training the Invariance Promotion Model IFM based on the target sub - atlas set includes:

[0014] Training the Invariance Promotion Model IFM based on the training sub - atlas set;

[0015] The step of returning to train the Invariance Promotion Model IFM based on the target sub - atlas set until a preset iterative training termination condition is reached includes:

[0016] Returning the step of training the Invariance Promotion Model IFM based on the training sub - atlas set until a preset iterative training termination condition is reached.

[0017] In one implementation of the first aspect, after calculating the loss based on the invariance loss function and performing backpropagation to update the parameters of the Invariance Promotion Model IFM, it further includes:

[0018] Using the Invariance Promotion Model IFM to perform node classification on the validation sub - atlas set and evaluating the classification performance;

[0019] The step of returning to train the Invariance Promotion Model IFM based on the training sub - atlas set until a preset iterative training termination condition is reached includes:

[0020] Returning the step of training the Invariance Promotion Model IFM based on the training sub - atlas set until the classification performance no longer improves in consecutive iterations.

[0021] In one implementation of the first aspect, the Invariance Promotion Model IFM includes an Invariance Representation Encoder IRE and a Node Representation Encoder NRE;

[0022] Training the Invariance Promotion Model IFM based on the training sub - atlas set includes:

[0023] Learning the invariance representation of the class labels of the nodes in the training sub - atlas set based on the Invariance Representation Encoder IRE;

[0024] Learning the node representations in the training sub - atlas set based on the invariance representation and the Node Representation Encoder NRE.

[0025] In one implementation of the first aspect, learning the invariance representation of the class labels of the nodes in the training sub - atlas set based on the Invariance Representation Encoder IRE includes:

[0026] Learning the invariance representation of the class labels of the nodes in the training sub - atlas set based on the Invariance Attention InvarATT Implementing compressed long - range dependencies;

[0027] Among them, the InvarATT (Invariance Attention) is as follows:

[0028]

[0029] Among them, is the initial invariance representation of the class label c, and s c is the node with node class c, is obtained by splicing from

[0030] In one implementation manner of the first aspect, learning the node representations in the training sub-graph set based on the invariance representation and the node representation encoder NRE includes:

[0031] Capturing the global invariance of the class label from the invariance representation by means of the TeleATT (Teleportation Attention), and combining with the GNN (Graph Neural Network) to capture the local invariance of the neighbor nodes, so as to realize the learning of the node representations in the training sub-graph set;

[0032] Among them,

[0033] Among them, H is obtained by mapping the node representations in the training sub-graph set, and performing a stop-gradient operation on to obtain

[0034] In one implementation manner of the first aspect, the invariance loss function includes a contrast loss and an entropy loss

[0035]

[0036] Among them, is the invariance loss function, the contrast loss is used to minimize the distance between the node representations of the same class, and the entropy loss is used to maximize the entropy of the node representations. C is the total number of all class labels in the training sub-graph set, x is the node in the current iterative training sub-graph, and the invariance representation sharing the same label as the node x, is the reference node representation of the class label k, τ is the temperature coefficient, λ represents the regularization weight, represents the similarity between the node and the invariance representation , d is the dimension of the node representation, and x i is the value of the i-th dimensional feature of the node representation.

[0037] The second aspect of the embodiments of the present application provides a training system for a large-scale graph node classification model, including:

[0038] A graph cutting module for cutting target large-scale graph data to obtain a target sub-graph set, where the target sub-graph set includes multiple target sub-graphs;

[0039] A model training module for training an Invariance Facilitation Model (IFM) based on the target sub-graph set;

[0040] A parameter update module for calculating a loss based on an invariance loss function and performing backpropagation to update the parameters of the Invariance Facilitation Model (IFM);

[0041] An iteration module for returning the step of training the Invariance Facilitation Model (IFM) based on the target sub-graph set until a preset iteration training termination condition is reached;

[0042] A model output module for outputting the Invariance Facilitation Model (IFM), where the Invariance Facilitation Model (IFM) is used for classifying large-scale graph nodes.

[0043] A third aspect of the embodiments of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method described in the first aspect are implemented.

[0044] A fourth aspect of the embodiments of the present application provides a computer program product, including a computer program. When the computer program is run, the method described in the first aspect is executed.

[0045] The beneficial effect of the first aspect of the embodiments of the present application is that by cutting the target large-scale graph data to obtain a target sub-graph set, then training the Invariance Facilitation Model (IFM) based on the target sub-graph set, calculating a loss based on the invariance loss function and performing backpropagation to update the parameters of the Invariance Facilitation Model (IFM), returning the step of training the Invariance Facilitation Model (IFM) based on the target sub-graph set until a preset iteration training termination condition is reached, and finally outputting the Invariance Facilitation Model (IFM), the training of a large-scale graph node classification model is realized under low computing resource conditions, and the trained Invariance Facilitation Model (IFM) has lower requirements for computing resources.

[0046] It can be understood that the beneficial effects of the second to fourth aspects above can refer to the relevant descriptions in the first aspect above, and will not be repeated here. Description of the Drawings

[0047] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for use in the embodiments or the description of the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0048] Figure 1 It is a schematic flowchart of the implementation of the training method for the large-scale graph node classification model provided by the embodiments of the present application;

[0049] Figure 2 It is a schematic flowchart of the training process of the IFM model provided by the embodiments of the present application;

[0050] Figure 3 It is a schematic diagram of the training system for the large-scale graph node classification model provided by the embodiments of the present application;

[0051] Figure 4 It is a schematic diagram of the computer device provided by the embodiments of the present application;

[0052] Figure 5 It is a schematic diagram of the computer program product provided by the embodiments of the present application. Detailed implementation manners

[0053] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, the detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0054] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0055] It should also be understood that the term " / and" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0056] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]" depending on the context.

[0057] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0058] Reference to "one embodiment" or "some embodiments" or the like described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0059] An embodiment of the present application provides a training method for a large-scale graph node classification model, which is used to train a node classification model applicable to a large-scale graph under lower computing resource conditions. The training method for the large-scale graph node classification model provided by the embodiment of the present application cuts the target large-scale graph data to obtain a target sub-graph set, then trains the invariance promotion model IFM based on the target sub-graph set, calculates the loss based on the invariance loss function and performs backpropagation to update the parameters of the invariance promotion model IFM, returns to the step of training the invariance promotion model IFM based on the target sub-graph set until a preset iterative training termination condition is reached, and finally outputs the invariance promotion model IFM, realizing the training of the large-scale graph node classification model under lower computing resource conditions, and the trained invariance promotion model IFM has lower requirements for computing resources.

[0060] The training method of the large-scale graph node classification model provided by the embodiments of the present application can be applied to computer devices such as desktop computers, notebooks, palm computers, cloud servers, mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, notebook computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc. The embodiments of the present application do not impose any restrictions on the specific types of computer devices.

[0061] As Figure 1 shown, the first aspect of the embodiments of the present application provides a training method for a large-scale graph node classification model, including:

[0062] Step S10: Cut the target large-scale graph data to obtain a target sub-graph set, where the target sub-graph set includes a plurality of target sub-graphs.

[0063] In application, for the target large-scale graph G=(V, E), where V is the node set in the target large-scale graph and E is the edge set in the target large-scale graph. The target large-scale graph G is segmented by using the multilevel graph segmentation algorithm (Multilevel Edge and Vertex BISECTION, METIS) to obtain a target sub-graph set where G1,…,G m are m target sub-graphs. Each sub-graph G i ={X i ,A i ,Y i}, represents N in nodes with d i -dimensional features, represents the adjacency matrix, is the C-type node label.

[0064] It can be understood that the goal of sub-graph invariant learning can be regarded as finding an optimal predictor where is the graph space, is the node label space, making it perform well on all subgraphs. Note that subgraphs exhibit feature distribution shifts and structural distribution shifts because they may come from different environments in a large-scale graph. These distribution shifts pose significant challenges in finding the optimal predictor, mainly because they hinder the learning of effective feature representations. To obtain a robust feature representation, it is assumed that each node in the large-scale graph contains invariant information, ensuring that its relationship with the label (the class of the node) remains consistent across different subgraphs. This invariance principle guarantees that the model can achieve good out-of-distribution (OOD) performance by effectively identifying and exploiting this invariance, that is, satisfying the invariance assumption and the sufficiency assumption. The invariance property means that for any two subgraphs, if they have the same label, then their node representations should also be the same. The sufficiency property means that a predictor (classifier) should be able to be found such that the node label can be accurately predicted by the invariant feature representation, and the random noise is statistically independent of the invariant feature representation. The invariance property implies that the invariant information covers the entire large-scale graph, and the sufficiency property allows this invariant information to be fully used for predicting the label. Therefore, the problem of cultivating the optimal invariant feature encoder and predictor is: given a set of small-scale subgraphs, the task is designed to jointly capture the feature invariance in all subgraphs and learn a predictor through the invariant feature representation, so as to achieve robust OOD generalization performance.

[0065] Step S20, train the Invariance Facilitation Model (IFM) based on the target subgraph set.

[0066] In the application, train the Invariance Facilitation Model (IFM) based on multiple target subgraphs in the target subgraph set, so that the IFM can learn the node features in the target subgraph set.

[0067] Step S30, calculate the loss based on the invariance loss function and perform backpropagation to update the parameters of the Invariance Facilitation Model (IFM).

[0068] In the application, calculate the loss through the Invariance Loss function and perform backpropagation to update the parameters of the Invariance Facilitation Model (IFM) to improve the classification performance of the IFM. The AdamW algorithm is used to update the parameters of the model IFM in the application.

[0069] Step S40, return to the step of training the Invariance Facilitation Model (IFM) based on the target subgraph set until the preset iterative training termination condition is reached.

[0070] In the application, the invariance promotion model IFM is iteratively trained by training it multiple times based on the target sub-graph set to improve the accuracy of the IFM. By presetting the termination condition for iterative training, when the termination condition for iterative training is reached, the iterative training is stopped to avoid overfitting of the IFM to the target sub-graph set, thereby improving the generalization performance of the IFM.

[0071] Step S50: Output the invariance promotion model IFM, which is used for classifying large-scale graph nodes.

[0072] In the application, after the preset termination condition for iterative training is reached, the invariance promotion model IFM is output. This invariance promotion model IFM has low requirements for computing resources and can achieve accurate classification of large-scale graph nodes.

[0073] In the application, by training the node classification model based on the target large-scale graph data in different application fields, different large-scale graph data node classification models can be obtained. In the application, after outputting the invariance promotion model IFM, it further includes deploying the target classification model based on the invariance promotion model IFM, and inputting the large-scale graph to be classified into the target classification model to obtain the class labels of the nodes in the large-scale graph to be classified.

[0074] In applications, this application can be applied to social network analysis and user profile optimization. By constructing a user interaction graph in the social network, based on node classification technology, the user nodes are classified by features to predict features such as users' interests, hobbies, and occupations, and generate accurate personalized recommendations. For example, user nodes are classified with labels such as "technology enthusiast", "traveler", "fashionista", etc. to optimize the content recommendation algorithm. At the same time, the classification results are used to identify fake accounts or abnormal behaviors (such as zombie fans or malicious accounts) in the social network, thereby improving the security and user experience of the social platform. Accurate user profiles enhance the advertising delivery effect and strengthen user stickiness; anomaly detection reduces potential losses and reputation risks of the platform. This application can be applied to financial risk control and transaction fraud detection. In the financial transaction network, a fund flow graph or customer association graph is constructed, with accounts or customers as nodes, and high-risk customers or potential fraud behaviors are identified through node classification technology. For example, customer nodes are classified as "ordinary users", "high-net-worth users", "high-risk accounts", and the credit scoring model is optimized in combination with the classification results. In addition, abnormal behaviors (such as circular transactions or abnormal transfers) in the fund chain are detected through node classification technology to assist the anti-money laundering system in automatically identifying suspicious transactions. This helps financial institutions reduce operational risks, improve risk control efficiency, and ensure the security and compliance of the transaction system. This application can be applied to medical health and drug development. In the biomolecular network or patient relationship network, node classification technology is used to classify nodes such as genes, proteins, and patients. For example, gene nodes are classified as "disease-related" or "non-related" for disease prediction or target discovery; or patient nodes are classified as "high-risk patients", "medium-risk patients" to help doctors develop personalized treatment plans. In addition, classification technology can also be used in drug research and development. By identifying functional molecular nodes in the compound network, potential drug molecules are screened out. This application can be applied to intelligent transportation and urban planning. Based on the graph model constructed from the urban traffic network, nodes (such as intersections, bus stops) are classified as "high traffic", "medium traffic", or "low traffic" to predict the changing trend of traffic flow and assist in optimizing traffic signal scheduling. For example, the classification results are used to adjust the signal light time during peak hours or design diversion plans to relieve traffic congestion. Classification technology can also be applied to the public transportation network to optimize the station layout or route planning and improve the overall transportation efficiency.

[0075] In one embodiment, before training the invariance promotion model IFM based on the target sub-graph set in step S20, the following steps are further included:

[0076] Step S11, dividing the target sub-graph set into a training sub-graph set, a validation sub-graph set, and a test sub-graph set.

[0077] In applications, the target sub-graph set is divided into a training sub-graph set validation sub-graph set and the test subatlas Specifically, the training sub-atlas Include target subatlas 80% of the subgraphs in the validation set, including the labeled data Include target subatlas 10% of the subgraphs in the test set Include target subatlas 10% of the sub-graphs in the dataset have no labeled data.

[0078] The step S20, training the invariance promotion model IFM based on the target sub-graph set, comprises:

[0079] Step S21, training the invariance promotion model IFM based on the training sub-atlas.

[0080] In the application, the invariance promotion model IFM is trained based on the training subgraph set with labeled data, so that IFM can learn the node features in the training subgraph set.

[0081] The step S40, returning to the step of training the invariance promoting model IFM based on the target sub-graph set until a preset iterative training termination condition is reached, includes:

[0082] Step S41, returning to the step of training the invariance promoting model IFM based on the training sub-graph set until a preset iterative training termination condition is reached.

[0083] In one embodiment, after the step S30, calculating the loss based on the invariance loss function and performing back propagation to update the invariance-promoting model IFM parameters, further includes:

[0084] Step S31, using the invariance promotion model IFM to classify nodes on the verification subgraph set, and evaluate the classification performance.

[0085] In the application, after each iteration of training to update the parameters of IFM, the IFM classification performance of the invariance promotion model is evaluated by classifying the nodes in the validation subgraph set. In the application, the IFM classification performance can be evaluated by one or more of the indicators such as accuracy, precision, F1 score, confusion matrix, KS (Kolmogorov-Smirnov) value, etc.

[0086] The step S41, returning to the step of training the invariance promoting model IFM based on the training sub-graph set until a preset iterative training termination condition is reached, includes:

[0087] Step S42. Return to the step of training the invariance promotion model IFM based on the training sub - atlas until the classification performance no longer improves in consecutive multiple iterations.

[0088] In applications, the preset iteration training termination condition is that the classification performance of IFM no longer improves in consecutive multiple iterations. For example, the classification performance of IFM no longer improves in consecutive 10 iterations. Through the above - mentioned iteration training termination condition, early stopping can be achieved, avoiding over - fitting of IFM on the training sub - atlas, which is beneficial to improving the generalization performance of IFM.

[0089] In applications, the preset iteration training termination condition can also be set as the number of iteration training reaches the iteration threshold, and the iteration threshold can be specifically set according to the scale of the training sub - atlas.

[0090] As Figure 2 shown, in one embodiment, the invariance promotion model IFM includes an invariance representation encoder IRE and a node representation encoder NRE;

[0091] In applications, the invariance promotion model IFM includes an invariance representation encoder IRE (Invariance Representation Encoder) and a node representation encoder NRE (Node Representation Encoder).

[0092] The step S21 of training the invariance promotion model IFM based on the training sub - atlas includes:

[0093] Step S22. Learn the invariance representation of the class labels of the nodes in the training sub - atlas based on the invariance representation encoder IRE.

[0094] In applications, learning the invariance representation of the class labels of the nodes in the training sub - atlas based on IRE is invariance representation learning.

[0095] Step S23. Learn the node representations in the training sub - atlas based on the invariance representation and the node representation encoder NRE.

[0096] In applications, learning the node representations in the training sub - atlas based on the invariance representation and the node representation encoder NRE is node representation learning.

[0097] In applications, updating the parameters of the invariance promotion model IFM includes updating the parameters of IRE and updating the parameters of NRE.

[0098] In one embodiment, step S22 of learning the invariance representation of the class labels of the nodes in the training subset based on the invariance representation encoder IRE includes:

[0099] Step S221 of learning the invariance representation of the class labels of the nodes in the training subset based on the invariance attention InvarATT to achieve compression of long-range dependencies;

[0100] where the invariance attention InvarATT is:

[0101]

[0102] where is the initial invariance representation of the class label c, and s c is the node with node class c, is obtained by concatenation.

[0103] In an application, s c is a fixed number of randomly sampled node samples of the class label c from the set S c of nodes with label c in the training subset.

[0104] In one embodiment, step S23 of learning the node representations in the training subset based on the invariance representation and the node representation encoder NRE includes:

[0105] Step S231 of capturing the global invariance of the class labels from the invariance representation using the tele-attention TeleATT, and combining with the GNN to capture the local invariance of the neighboring nodes, to achieve learning of the node representations in the training subset;

[0106] where

[0107] where H is obtained by mapping the node representations in the training subset, and performing a stop-gradient operation on to obtain

[0108] In an application, the integration of the global invariance information is achieved through the tele-attention, and the capture of the local invariance of the adjacent nodes within the subgraph is achieved through the graph neural network GNN (Graph Neural Network); the combination of the tele-attention and the GNN is used to capture the long-range dependencies using the compressed invariance representation, and by combining the global and local structural information, a more comprehensive and detailed node representation is achieved.

[0109] In one embodiment, the invariance loss function includes a contrastive loss and an entropy loss

[0110] where is the invariance loss function, the contrastive loss is used to minimize the distance between node representations of the same class, and the entropy loss is used to maximize the entropy of the node representations. C is the total number of all class labels in the training subset of subgraphs, x is a node in the current iteratively trained subgraph, and the invariance representation shares the same label as node x, is the reference node representation of class label k, τ is the temperature coefficient, λ represents the regularization weight, represents the similarity between the node and the invariance representation , d is the dimension of the node representation, and x i is the value of the i-th dimensional feature of the node representation.

[0111] In applications, minimizing the contrastive loss can promote the representativeness of the node representations, making them more consistent with the invariance representations of the actual node labels, while remaining distinct from the invariance representations of other node labels, thereby enabling the invariance property to take effect and enhancing the interpretability of node prediction (classification) by making the node representations of the same node label closer to the invariance representations of the node labels. Maximizing the entropy loss can maximize the entropy of the node representations, which is beneficial for maintaining effective numerically balanced features, thereby avoiding the simplification of some features (tending to 0 after normalization), and encouraging the IFM model to retain more unique features, thus strengthening the sufficiency property.

[0112] In applications, comparing the performance of the invariance-promoting model IFM of the present application on small-scale subgraphs with the performance of traditional GNN models on large-scale graphs can lead to the following conclusions:

[0113] In terms of time complexity, the time complexity of the GNN model is O(n 2 d) per layer, where n is the number of nodes in the large-scale graph. In contrast, IFM operates on each subgraph, and the time complexity of the entire graph is O(ms 2 d), and the time complexity of a single subgraph is O(s 2 d), where m represents the number of subgraphs, and s << n represents the size of the largest subgraph. Therefore, the time complexity of IFM is significantly lower than that of GNN.

[0114] Model Capacity: According to the scaling law, the model capacity is positively correlated with the number of parameters. However, restricted by large-scale graph data, GNN models cannot scale the model parameters and cannot effectively capture long-range dependencies through a wide receptive field. Instead, due to the small scale of subgraphs, IFM can afford the overhead of scaling parameters. IFM effectively models long-range dependencies by introducing an invariance representation encoder as a long-range information compressor and using tele-attention TeeATT to deliver long-range information to each node.

[0115] Interpretability: According to the invariance property in the hypothesis, the parameters of IFM are optimized by the invariance loss, making the invariance representations of nodes with the same label more similar. Therefore, this method improves interpretability by pulling nodes with the same label closer to a fixed embedding. In contrast, GNN models apply the MLP function as a predictor and use the cross-entropy loss for parameter optimization, resulting in a lack of interpretability.

[0116] Sparsity: Irregular graphs or graphs without a hierarchy may lack clear clustering but exhibit sparsity, which may hinder the information flow of traditional GNNs and lead to the over-squashing problem. IFM solves this problem by dividing large-scale graphs into smaller subgraphs by removing sparse edges, thus improving scalability. In addition, the NRE module uses IRE and TeleATT to provide sufficient long-range information for nodes to achieve accurate label prediction. Therefore, IFM is applicable to various large-scale graph types, even those that do not possess the small-world property.

[0117] As shown in the following table, on the four datasets of Reddit, Amazon, Pascal VOC, and COCO, after adding the invariance loss function and IFM in the training method of this application to traditional GNN models (GCN, SAGE, SGC, GAT, GATv2, UniMP), the classification accuracy and classification balance and other performances evaluated by the weighted AUC-ROC (Area Under the Curve - Receiver Operating Characteristic Curve) score and the weighted F1 score have certain improvements.

[0118]

[0119] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0120] The embodiment of the present application also provides a training system for a large-scale graph node classification model, which is used to execute the steps in the embodiment of the training method for the large-scale graph node classification model. The training system for the large-scale graph node classification model can be a virtual appliance in a computer device, run by the processor of the computer device, or the computer device itself.

[0121] As Figure 3 shown, the second aspect of the embodiment of the present application provides a training system 300 for a large-scale graph node classification model, including:

[0122] A graph cutting module 301, configured to cut the target large-scale graph data to obtain a target sub-graph set, where the target sub-graph set includes a plurality of target sub-graphs;

[0123] A model training module 302, configured to train an invariance promotion model IFM based on the target sub-graph set;

[0124] A parameter update module 303, configured to calculate a loss based on an invariance loss function and perform backpropagation to update the parameters of the invariance promotion model IFM;

[0125] An iteration module 304, configured to return to the step of training the invariance promotion model IFM based on the target sub-graph set until a preset iteration training termination condition is reached;

[0126] A model output module 305, configured to output the invariance promotion model IFM, where the invariance promotion model IFM is used to classify large-scale graph nodes.

[0127] In one embodiment, the graph cutting module 301 is further configured to divide the target sub-graph set into a training sub-graph set, a validation sub-graph set, and a test sub-graph set;

[0128] The model training module 302 is configured to train the invariance promotion model IFM based on the training sub-graph set;

[0129] The iteration module 304 is configured to return to the step of training the invariance promotion model IFM based on the training sub-graph set until a preset iteration training termination condition is reached.

[0130] In one embodiment, the training system 30 for the large-scale graph node classification model further includes:

[0131] A performance evaluation module 306, configured to use the invariance promotion model IFM to perform node classification on the validation sub-graph set and evaluate the classification performance;

[0132] The iterative module 304 is configured to return the step of training the invariance promotion model IFM based on the training sub - atlas until the classification performance no longer improves in consecutive iterations.

[0133] In one embodiment, the invariance promotion model IFM includes an invariance representation encoder IRE and a node representation encoder NRE;

[0134] The model training module 302 includes:

[0135] The IRE unit 3021 is configured to learn an invariance representation of the class labels of the nodes in the training sub - atlas based on the invariance representation encoder IRE;

[0136] The NRE unit 3022 is configured to learn the node representations in the training sub - atlas based on the invariance representation and the node representation encoder NRE.

[0137] In one embodiment, the IRE unit 3021 is configured to learn an invariance representation of the class labels of the nodes in the training sub - atlas based on the invariance attention InvarATT to achieve compression of long - range dependencies;

[0138] where the invariance attention InvarATT is:

[0139]

[0140] where is the initial invariance representation of the class label c, s c is the node with node class c, is composed of concatenated.

[0141] In one embodiment, the NRE unit 3022 is configured to capture the global invariance of the class label from the invariance representation and combine with GNN to capture the local invariance of neighbor nodes to achieve learning of the node representations in the training sub - atlas;

[0142] where

[0143] where H is obtained by mapping the node representations in the training sub - atlas, and the stop - gradient operation is performed on to obtain

[0144] In one embodiment, the invariance loss function includes a contrastive loss and an entropy loss

[0145] Among them, is the invariance loss function, the contrastive loss is used to minimize the distance between node representations of the same class, and the entropy loss is used to maximize the entropy of the node representation. C is the total number of all class labels in the training subset of subgraphs, x is the node in the current iteration of the training subgraph, and the invariance representation sharing the same label as node x, is the reference node representation of class label k. τ is the temperature coefficient, and λ represents the regularization weight. represents the similarity between the node and the invariance representation d is the dimension of the node representation, and x i is the value of the i-th dimensional feature of the node representation;

[0146] The parameter update module 303 is used to calculate the loss based on the invariance loss function and perform backpropagation to update the parameters of the invariance promotion model IFM.

[0147] In applications, each module in the training system of the large-scale graph node classification model can be a software program module, can also be implemented by different logic circuits integrated in the processor, or can be implemented by multiple distributed processors.

[0148] Figure 4 is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 4 shown, the computer device 4 of this embodiment includes: at least one processor 40 ( Figure 4 only one is shown in the figure), a processor, a memory 41, and a computer program 42 stored in the memory 41 and executable on the at least one processor 40. When the processor 40 executes the computer program 42, the steps in any of the above-mentioned embodiments of the large-scale graph data node classification training method are implemented.

[0149] The computer device may include, but is not limited to, a processor 40 and a memory 41. Those skilled in the art can understand that Figure 4 is only an example of the computer device 4, and does not constitute a limitation on the computer device 4. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0150] The processor 40 may be a Central Processing Unit (CPU), or the processor 40 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0151] In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 4. Further, the memory 41 may also include both the internal storage unit and the external storage device of the computer device 4. The memory 41 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program, etc. The memory 41 may also be used to temporarily store data that has been output or is to be output.

[0152] It should be noted that for the content such as information interaction and execution process between the above-mentioned device / unit, since it is based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, reference may be specifically made to the method embodiment part, and details are not described herein again.

[0153] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0154] An embodiment of this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented.

[0155] An embodiment of this application provides a computer program product 5, including a computer program 50. When the computer program 50 is run, the steps in the foregoing method embodiments of various large-scale graph data node classification training methods are executed.

[0156] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the device / computer equipment, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0157] In the above embodiments, the descriptions of the respective embodiments each have their own emphasis. For parts not described in detail or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0158] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0159] In the embodiments provided in this application, it should be understood that the disclosed computer devices and methods can be implemented in other ways. For example, the computer device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0160] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0161] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of this application, and should all be included in the protection scope of this application.

Claims

1. A training method for a large-scale graph node classification model, characterized in that: include: Cutting the target large-scale graph data to obtain a target subgraph set, wherein the target subgraph set includes a plurality of target subgraphs; Training an invariance-enhancing model IFM based on the target sub-atlas; Calculate the loss based on the invariance loss function and perform back propagation to update the invariance promotion model IFM parameters; Returning to the step of training the invariance promoting model IFM based on the target sub-graph set until a preset iterative training termination condition is reached; The invariance promotion model IFM is output, and the invariance promotion model IFM is used to classify large-scale graph nodes.

2. The training method for a large-scale graph node classification model according to claim 1, characterized in that: Before training the invariance promotion model IFM based on the target sub-graph set, the method further includes: Dividing the target sub-atlas into a training sub-atlas, a validation sub-atlas and a test sub-atlas; The step of training the invariance promotion model IFM based on the target sub-graph set comprises: Training an invariance-enhancing model IFM based on the training sub-atlas; The step of returning to the step of training the invariance promoting model IFM based on the target sub-graph set until a preset iterative training termination condition is reached includes: Return to the step of training the invariance promoting model IFM based on the training sub-atlas until a preset iterative training termination condition is reached.

3. The training method of a large-scale graph node classification model as claimed in claim 2, characterized in that: After calculating the loss based on the invariance loss function and performing back propagation to update the invariance promotion model IFM parameters, it also includes: Using the invariance promotion model IFM to classify nodes on the validation subgraph set, and evaluating the classification performance; The step of returning to the step of training the invariance promoting model IFM based on the training sub-atlas until a preset iterative training termination condition is reached includes: Returning to the step of training the invariance-enhancing model IFM based on the training sub-atlas, until the classification performance no longer improves in a plurality of consecutive iterations.

4. The training method for a large-scale graph node classification model according to claim 2, characterized in that: The invariance promotion model IFM includes an invariance representation encoder IRE and a node representation encoder NRE; The step of training the invariance promotion model IFM based on the training sub-atlas comprises: Based on the invariant representation encoder IRE, learning the invariant representation of the category labels of the nodes in the training subgraph set; The node representations in the training sub-graph set are learned based on the invariant representation and the node representation encoder NRE.

5. The training method for a large-scale graph node classification model according to claim 4, characterized in that: The learning of the invariant representation of the category labels of the nodes in the training subgraph set based on the invariant representation encoder IRE includes: Based on the invariant attention InvarATT, we learn the invariant representation of the category labels of the nodes in the training subgraph set. Achieve compression of long-range dependencies; Among them, the invariant attention InvarATT is: in, is the initial invariant representation of the category label c, s c is a node of node category c, Depend on Splicing obtained.

6. The training method for a large-scale graph node classification model according to claim 5, characterized in that: The learning of the node representations in the training subgraph set based on the invariant representation and the node representation encoder NRE comprises: TeleATT based on remote attention is derived from the invariant representation The global invariance of the category label is captured in the training subgraph, and the local invariance of the neighboring nodes is captured by combining the GNN to achieve learning of the node representation in the training subgraph set; in, Where H is the node representation mapping in the training subgraph set. Perform a stop gradient operation to obtain 7. The training method for a large-scale graph node classification model according to any one of claims 1 to 6, characterized in that: The invariance loss function includes contrast loss and entropy loss in, is the invariance loss function, contrast loss Used to minimize the distance between node representations of the same category, entropy loss It is used to maximize the entropy of node representation, C is the total number of all category labels in the training subgraph set, x is the node in the current iteration training subgraph, and the invariance representation Share the same label as node x, is the reference node representation of category label k, τ is the temperature coefficient, λ represents the regularization weight, Representing nodes and invariant representation Similarity, d is the dimension of node representation, x i is the value of the i-th dimension feature represented by the node.

8. A training system for a large-scale graph node classification model, characterized in that: include: A graph cutting module, used for cutting the target large-scale graph data to obtain a target subgraph set, wherein the target subgraph set includes multiple target subgraphs; A model training module, used for training an invariance promotion model IFM based on the target sub-graph set; A parameter updating module, used for calculating the loss based on the invariance loss function and performing back propagation to update the invariance promotion model IFM parameters; An iteration module, used for returning to the step of training the invariance promotion model IFM based on the target sub-graph set until a preset iterative training termination condition is reached; The model output module is used to output the invariance promotion model IFM, and the invariance promotion model IFM is used to classify large-scale graph nodes.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product, characterized in that The invention comprises a computer program, which, when being executed, enables the method as claimed in any one of claims 1 to 7 to be performed.