Large-scale graph data processing method based on Meta-GNN model

Through the Meta-GNN model and meta-learning framework, large-scale graph data are sub-graphed and trained, which solves the efficiency problem when processing large-scale graph data on a computer in the existing technology, and realizes efficient graph data processing and learning.

CN119940402APending Publication Date: 2025-05-06HEFEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510044740.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When processing large-scale graph data, it is difficult to efficiently process and improve learning efficiency on a computer, and traditional methods have shortcomings in memory usage and computing efficiency.

Method used

Using the Meta-GNN model, by classifying nodes and dividing subgraphs of the data sets, selecting subgraphs with large edge degrees as meta-training subgraphs, using the meta-learning framework to train the meta-training subgraphs, obtaining parameters with strong generalization capabilities, and then training small batch training subgraphs to improve the learning efficiency of the entire graph data.

Benefits of technology

While not reducing the accuracy, the learning efficiency of graph neural networks when processing large-scale graph data is significantly improved, memory usage is reduced, and model generalization ability is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940402A_ABST
    Figure CN119940402A_ABST
Patent Text Reader

Abstract

The invention discloses a large-scale graph data processing method based on a Meta-GNN model, and relates to the technical field of graph data processing. For large-scale graph data, the storage problem can be effectively solved by dividing the large-scale graph data into a plurality of computers for learning or carrying out sampling learning on the large-scale graph data, but new problems of low transmission overhead or learning efficiency and the like are generated. Therefore, the Meta-GNN model is provided, the whole graph can be learned on one machine, and the learning efficiency is improved while the storage problem is solved and the precision is not reduced. The generalization and adaptability of the model are enhanced by identifying the general rule of the structure and the attribute in the graph data, so that the performance of the model during large-scale graph data processing is improved, and the efficiency of large-scale graph data processing is improved. Data set experiment results prove that the method has a remarkable effect on improving GNN training and efficiency on large-scale graph data, and a powerful solution is provided for solving the problem of complex graph data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of graph data processing, and in particular to a large-scale graph data processing method based on a Meta-GNN model. Background Art

[0002] Meta-learning and graph neural networks belong to two different fields, but there are certain connections and potential points of convergence. GNNs are good at processing structured data, such as social networks and knowledge graphs, and they can effectively capture the dependencies between nodes; while meta-learning focuses on quickly adapting to new tasks. If the new task is about graph-structured data, graph neural networks can be used as a powerful feature extractor to help meta-learning models better understand and process such data.

[0003] Meta-learning, or learning to learn, aims to address the shortcomings of neural networks in adaptability and generalization in new tasks. By learning a small number of tasks, it obtains parameters with strong generalization so that the model can quickly adapt to new tasks. It is mainly used in computer vision, natural language processing, autonomous driving, bioinformatics, reinforcement learning, recommendation systems and other fields. Because meta-learning can help solve challenges such as data scarcity, rapid adaptation, generalization ability and flexibility, thus making machine learning models more universal, adaptable and practical, meta-learning can currently be applied to multi-task and small sample scenarios in graph neural networks.

[0004] Multi-task learning requires the model to handle multiple related tasks at the same time. Given that different tasks may have different graph structures and different requirements for feature representation, traditional multi-task learning methods may encounter performance degradation in the field of graph representation learning. In the early days of multi-task learning, researchers focused on how to improve the performance of models on multiple related tasks by sharing representations, and introduced meta-learning to solve the problems of parameter sharing and task-specific parameter optimization in multi-task learning; in recent developments, researchers have begun to explore how to use meta-learning to optimize the multi-task learning framework to better share and utilize knowledge between tasks. Meta-learning can be used to simultaneously process text classification tasks in multiple languages ​​and improve cross-language generalization capabilities; in dialogue systems, meta-learning can help models simultaneously handle multiple subtasks, such as intent recognition, slot filling, etc.; meta-learning can be used to simultaneously identify multiple object categories and simultaneously perform tasks such as action recognition and scene classification; meta-learning can be used to train robots to master multiple skills at the same time, such as navigation, grasping, recognition, etc.; meta-learning can also be used to simultaneously predict the risks of multiple diseases and improve the accuracy of diagnosis, etc. The application of meta-learning in multi-task learning is gradually becoming a research hotspot. With the deepening of research, the application of meta-learning in multi-task learning will become more extensive and efficient.

[0005] Traditional machine learning methods usually require a large amount of data to train models, but in many practical applications, it is expensive and time-consuming to obtain a large amount of labeled data. Meta-learning has received widespread attention in small sample learning scenarios, aiming to solve the problem of performance degradation of traditional machine learning when the sample size is limited. It imitates the way humans learn and can quickly learn new concepts from a small number of examples, that is, quickly adapt to new tasks through a small number of samples. In robot learning, meta-learning can help robots quickly adapt to new tasks. For example, by using small sample learning methods, robots can learn how to manipulate new objects after only a few encounters with them. Data enhancement techniques such as FlipDA can be used to simulate different object perspectives, thereby improving the generalization ability of the model without increasing the actual data collection cost. In the field of natural language processing, small sample learning can be used to deal with text classification problems of a few categories, such as abnormal event detection on social media, and label hallucination methods can be used to generate more training samples to improve the model's ability to recognize rare events. In recommendation systems, small sample learning can be used to deal with the cold start problem of new users or new items. Through meta-learning, the system can quickly adapt to the small amount of behavior data of new users and provide personalized recommendations, etc. Small sample learning problems usually involve scenarios that require rapid adaptation to new tasks. Unlike traditional machine learning methods that require retraining the entire model to adapt to new tasks, meta-learning methods can extract generalized knowledge from a small number of training samples to quickly adapt to new tasks. It can improve the generalization ability of the model, reduce dependence on large amounts of labeled data, and achieve efficient learning and application.

[0006] Graph neural network is a deep learning model specially designed for processing graph structured data. It can effectively extract features and learn representations of nodes in the graph (such as individuals in social networks, atoms in molecular structures) and the associations between nodes (such as friendships, chemical bonds). Graph neural networks have shown great application potential in many fields such as social network analysis, bioinformatics, computer vision, natural language processing, recommendation systems, and traffic forecasting. As a deep learning model for processing graph structured data, graph neural networks have achieved rapid development in both theory and practice, especially in the field of scientific intelligence, showing its unique advantages and broad application prospects. Graph neural networks are of great significance for processing complex graph structured data, and improving the learning efficiency of graph neural networks is a research hotspot.

[0007] At present, there are many methods to improve the efficiency of graph neural networks, which are mainly divided into distributed and non-distributed. Distributed methods require multiple computers, and use multiple computing nodes or GPUs to accelerate the training and inference process to improve computing efficiency. Representative frameworks include Pregel, GraphX, JanusGraph, etc. In addition to distributed methods, non-distributed methods are also a research focus. In traditional GNNs, the feature update of each node requires the aggregation of the features of all its neighboring nodes. GraphSAGE improves the operating efficiency of GNN by randomly sampling neighboring nodes, but may cause incomplete node representation. FastGCN uses importance sampling to reduce the amount of calculation and increase the complexity of the algorithm. LightGCN simplifies the GCN structure to improve computing efficiency, which may reduce the model's expressiveness. APPNP combines graph neural networks and PageRank propagation mechanisms to accelerate convergence, which may amplify noise and affect robustness. Both distributed and non-distributed methods reduce memory usage on the basis of traditional graph neural networks, but distributed methods require multiple computers, resulting in complex implementation and maintenance, and large transmission overhead. Non-distributed methods improve learning efficiency to a certain extent, but are still not ideal. It is the purpose of this invention to reduce transmission overhead and improve learning efficiency to a greater extent.

[0008] Although the above methods all solve the problem of excessive memory usage in full batch processing, distributed graph processing requires multiple computers, and the graph sampling method is inefficient in processing graph data. Therefore, the goal of this invention is to process large-scale graph data on a single computer and improve learning efficiency, reduce storage space usage by graph partitioning, and speed up learning efficiency by first learning part of the graph and then generalizing to the entire graph. Summary of the invention

[0009] The present invention proposes a Meta-GNN model that can learn the entire graph on a single machine and improve learning efficiency while solving storage problems and not reducing accuracy. The generalization and adaptability of the GNN model is enhanced by identifying the universal laws of structure and attributes in graph data, thereby improving its performance in processing large-scale graph data and improving the efficiency of GNN in processing large-scale graph data. The present invention has conducted a large number of experiments on five public data sets, and the experimental results confirm that the method of the present invention has a significant effect in improving the training and efficiency of GNN on large-scale graph data, providing a powerful solution for processing complex graph data problems.

[0010] In order to achieve the above object, the technical solution adopted by the present invention is:

[0011] A large-scale graph data processing method based on the Meta-GNN model includes the following steps:

[0012] Step 1: Dataset processing

[0013] A node classification experiment was conducted on the datasets. Each dataset was divided into a training set, a validation set, and a test set. The training set was used for model training and optimization, and the test set was used to evaluate the performance of the model.

[0014] Step 2: Model parameter setting

[0015] The GNN on all datasets has 2 layers and 128 hidden units; the dropout rate is 0.5; the learning rate is 0.001, and the weight decay is 0.0005; the ratio of meta-training subgraphs to mini-batch subgraphs is 2:3; training is stopped when the accuracy of the validation set hardly changes;

[0016] Step 3: Model training

[0017] Subgraph division and selection: Subgraph division is to use Metis to divide the large model graph data into many local subgraphs. Subgraph selection is to select the subgraphs with large edge degrees as meta-training subgraphs, and the remaining subgraphs are small batch training subgraphs.

[0018] Meta-training: Meta-training is training the meta-training subgraphs. The meta-training model can learn more generalized rules and quickly adapt to small-batch training subgraphs. The parameters of the subgraphs obtained by meta-learning are used to learn small-batch training subgraphs to improve learning efficiency.

[0019] Mini-batch training: Mini-batch training is to train the remaining subgraphs of the meta-training subgraph. Its initial parameters are the parameters updated by the meta-training. Each mini-batch training subgraph is independent and interdependent.

[0020] Step 4: Prediction and evaluation

[0021] Prediction stage: predict the nodes in the test set, regard the output of GNN as the final representation of the node, and use the softmax function to calculate the probability of each label;

[0022] Evaluation metrics: The performance of the model is evaluated using evaluation metrics such as the accuracy of the test set, the average training time per cycle, and the GPU memory consumption during model training to reflect the classification accuracy and generalization ability of the model.

[0023] As a preferred technical solution of the present invention, the specific steps of dividing and selecting subgraphs in step 3 are as follows:

[0024] Subgraph partitioning uses Metis to partition the large model graph data into many local subgraphs, and selects the subgraph with the largest number of edges from the partitioned subgraphs as the important subgraphs. The selector of the important subgraphs calculates and sorts the number of edges of the local subgraphs, and selects them according to the ranking of the number of edges. In the partitioned local subgraphs, if there is an edge between vertices i and j, it is recorded as A[i][j]=1, and if there is no edge, it is recorded as A[i][j]=0. The number of edges E is:

[0025]

[0026] The number of edges of all local subgraphs is calculated and sorted, and the local subgraphs with a large number of edges are selected as meta-training subgraphs, and the local subgraphs with a small number of edges are selected as mini-batch training subgraphs.

[0027] As a preferred technical solution of the present invention, the specific steps of meta-training in step 3 are as follows:

[0028] The graph data is divided into N subgraphs. During meta-training, a subgraph is used as a task. Tasks are independent. Assume that the learning function of each task is f, the loss function is L(f), and the objective function is min L(f). The goal is to make it as small as possible.

[0029] The initial parameter of the task is θ0, and for task T1:

[0030]

[0031] Where α is the step size, is the gradient of the loss function of the support set in task T1, θ1 is the parameter obtained after one iteration update, and the iteration may be multiple times. This process is the inner loop; when it converges after n iterations, at this time:

[0032]

[0033] The updated parameter θ1 after convergence n The query set passed to task T1 calculates its loss L1;

[0034] In meta-learning, the tasks are independent, so the initial parameters of all tasks are θ0. The parameters of the task are updated in the support set of all tasks, and the verification calculation is performed in the query set. Repeat the above steps to obtain the query set loss of all tasks as L1, L2, L3...L N ;

[0035] All tasks have an impact on the model, and what is needed is the global optimal rather than the local optimal parameter. Therefore, it is necessary to find a parameter with good generalization ability for small batch training, that is, to minimize the sum of all task loss functions:

[0036]

[0037] Among them, β is the step size. This process can be iterated multiple times. Finally, the updated parameters are passed to the small batch subgraphs as initialization parameters. This process is the outer loop.

[0038] As a preferred technical solution of the present invention, the specific steps of small batch training in step 3 are as follows:

[0039] In the mini-batch training phase, each subgraph is treated as an independent batch for training. After the training of the previous subgraph is completed, the updated parameters are passed to the next subgraph. The training process is as follows:

[0040] First, the parameters of the mini-batch training are updated. The initialization parameters are the parameters updated at the end of the meta-training. The initial parameters of the mini-batch subgraph are recorded as Input the node features and edge information of each subgraph of the first mini-batch subgraph, as well as its initial parameters Conduct training;

[0041] For each node v, its new feature representation h v l The update process is determined by the feature representation of its neighbor nodes and its own feature representation.

[0042]

[0043] After the training of each subgraph is completed, the parameters of GNN are updated through the gradient descent algorithm according to the gradient of the loss function. GNN learns by minimizing the loss function. Its loss function L is as follows:

[0044]

[0045] where y v is the true label, is the predicted label; in each iteration, the gradient of the loss function with respect to the model parameters is calculated, and the parameters are updated accordingly to gradually approach the optimal solution; the training process will continue until the convergence condition is met, that is, the value of the loss function no longer decreases significantly or reaches the preset number of iterations.

[0046] Compared with the prior art, the beneficial effects of the present invention are mainly manifested in:

[0047] 1. Graph data is widely present in real life. The analysis and processing of graph data has always been the focus and difficulty of the research community. How to improve the efficiency of processing graph data has become one of the research focuses. In this paper, a meta-learning framework is introduced to process a part of the subgraph, thereby improving the efficiency and performance of GNN when processing graph data. Experimental results show that the method of the present invention improves learning efficiency while solving storage problems and not reducing accuracy on various real graph data sets, proving its feasibility and effectiveness in practical applications.

[0048] 2. Processing large-scale graph data often requires more memory space. It only takes a short time to learn a full batch of graph data, but when the graph data becomes larger and larger, the existing computer memory may not be able to meet the needs. Existing methods often solve the memory consumption problem through sampling methods, but this method consumes a lot of time. The present invention improves the efficiency and performance of GNN when processing graph data by generalizing to the entire graph after processing a part of the graph data. Experimental results show that the method of the present invention improves learning efficiency while solving storage problems and not reducing accuracy on various real graph data sets, proving its feasibility and effectiveness in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is the overall structure diagram of the model designed by the present invention.

[0050] Figure 2 This is a comparison chart of accuracy. DETAILED DESCRIPTION

[0051] The present invention proposes a large-scale graph data processing method based on the Meta-GNN model, comprising the following steps:

[0052] Step 1: Dataset processing

[0053] Node classification experiments were conducted on a total of five data sets, including AmazonCoBuyComputer, CoauthorCS, Flickr, Reddit, and Ogbn-products. Each data set was divided into training set, validation set, and test set. The training set was used for model training and optimization, and the test set was used to evaluate the performance of the model.

[0054] Step 2: Model parameter setting

[0055] The GNN on all datasets has 2 layers and 128 hidden units. The dropout rate is 0.5. ADAM is used as the optimizer. The learning rate is 0.001 and the weight decay is 0.0005. The ratio of meta-training subgraphs to mini-batch subgraphs is 2:3. When the accuracy on the validation set is almost unchanged, the training is stopped.

[0056] Step 3: Model training

[0057] Subgraph partitioning and selection: Subgraph partitioning is to use Metis to divide the large model graph data into many local subgraphs. Subgraph selection is to select the subgraphs with large edge degrees as meta-training subgraphs, and the remaining subgraphs are small batch training subgraphs.

[0058] Meta-training: Meta-training is the training of meta-training subgraphs. Through the meta-training model, more generalized rules can be learned, which can quickly adapt to small-batch training subgraphs. The parameters obtained by meta-learning some subgraphs are used to learn small-batch training subgraphs to improve learning efficiency.

[0059] Mini-batch training: Mini-batch training is to train the remaining sub-graphs of the meta-training sub-graph. Its initial parameters are the parameters updated by meta-training. Each mini-batch training sub-graph is independent and interdependent.

[0060] Step 4: Prediction and evaluation

[0061] Prediction stage: predict the nodes in the test set, regard the output of GNN as the final representation of the node, and use the softmax function to calculate the probability of each label.

[0062] Evaluation metrics: The performance of the model is evaluated using evaluation metrics such as the accuracy of the test set, the average training time per cycle, and the GPU memory consumption during model training to reflect the classification accuracy and generalization ability of the model.

[0063] The present invention is further explained in detail below in conjunction with Examples 1 and 2:

[0064] Example 1

[0065] The present invention provides a method for accelerating the sampling efficiency of graph neural network subgraphs based on meta-learning, with the purpose of accelerating the learning efficiency of the remaining graph data by learning some local subgraphs. The processing of large-scale graph data is divided into three parts: subgraph division and selection, meta-training and small-batch training. After the large-scale graph data is divided into graphs, more important subgraphs are selected as meta-training subgraphs for meta-training, and the remaining subgraphs are small-batch training subgraphs for small-batch training. By learning the meta-training subgraphs, the model of the present invention can learn a global optimal parameter. Based on this parameter, the small-batch training subgraphs can be learned faster and better, thereby improving the learning efficiency. The overall structure diagram of the model designed by the present invention is shown as follows: Figure 1 shown.

[0066] Figure 1The working mechanism of the model includes the following parts: (1) First, a large graph is divided into several subgraphs; (2) For the divided subgraphs, the more important subgraphs are selected for meta-training. In meta-training, one subgraph represents one task, and the subgraphs other than the meta-training subgraphs are used for small batch training; (3) For each task T i The dataset in is divided into a support set and a query set. The initial model parameters θ0 and the support set in the T1 task are input into the GNN. (4) The updated node features can be obtained and the loss L of the support set can be calculated. support ; (5) According to the task L support Update the parameters of the GNN in the support set and iterate until the loss function converges; (6) After the loss function of the support set converges, pass the updated parameters θ1 to the GNN in the query set; (7) Calculate the loss of the query set of task T1 and record it as L query Repeat steps (3)-(7) for all tasks, and the initial parameters of all tasks are θ0. After calculating the query set loss of all tasks, (8) add these losses and average them as meta-loss. The initial model parameters are gradient updated through meta-loss. Repeat steps (4)-(8). Stop the loop when the meta-loss function converges. The updated parameters θ n is the initial parameter of mini-batch training. (9-12) In mini-batch training, one subgraph is trained each time, the first subgraph is passed to GNN, and its loss update parameter is calculated as is the initialization parameter of the second subgraph, and so on. After training all small batch subgraphs, iterate again until convergence. The updated node features can be used for downstream tasks.

[0067] 1. Division and selection of subgraphs

[0068] The present invention first divides the large model graph and then selects it. Subgraph selection is to divide a large graph data into many local subgraphs, and subgraph selection is to select the divided subgraphs as meta-training subgraphs and small-batch training subgraphs. Dividing the graph data can make GCN require less memory when processing large-scale graph data. Selecting important subgraphs for meta-training first and then performing small-batch training on the remaining subgraphs can speed up learning efficiency.

[0069] Subgraph partitioning can reduce memory requirements, reduce computational complexity, and improve model performance. Using traditional GCN to process large-scale graph data requires huge memory overhead. Dividing the graph data before processing can effectively reduce memory requirements. Traditional GCN needs to consider the structural information of the entire graph in the calculation of each layer. After subgraph partitioning, only the local structural information of the subgraph needs to be processed, which reduces the computational complexity. Subgraph partitioning can help the model capture local structural information that may be ignored from a global perspective. By training the model on subgraphs, richer feature representations can be learned to improve the performance of downstream tasks. Therefore, subgraph partitioning has an important impact on experiments.

[0070] Use Metis to partition the graph, which is an efficient and flexible graph partitioning method. Metis (MinimumCut) is an efficient graph partitioning library that can quickly partition a large graph into multiple subgraphs that are relatively balanced in size, number of edges, and load. Metis provides a variety of parameter options to adjust the partitioning strategy as needed. In addition, it can also generate subgraphs with high locality, that is, there are more edges directly between vertices within the subgraph, and fewer edges across subgraphs.

[0071] Select subgraphs after the graph is divided, and learn the remaining subgraphs based on the parameters obtained by meta-learning some subgraphs to improve learning efficiency. Therefore, the selection of meta-training subgraphs is particularly important, and more important subgraphs are selected for meta-training.

[0072] The present invention selects the subgraph with the largest number of edges from the divided subgraphs as the important subgraph. The larger the edge degree, the higher the connectivity, which means that the nodes in the subgraph can be connected to each other through more paths. In large-scale graph data, subgraphs with large edge degrees often bear important information and other functions, usually contain more information and complex structures, and therefore have higher value in training.

[0073] The selector of the important subgraph calculates and sorts the number of edges of the local subgraph and selects according to the ranking of the number of edges. In the divided local subgraph, if there is an edge between vertices i and j, it is recorded as A[i][j]=1, and if there is no edge, it is recorded as A[i][j]=0. The number of edges E is:

[0074]

[0075] Calculate the number of edges of all local subgraphs and sort them. Select local subgraphs with more edges as meta-training subgraphs, and local subgraphs with fewer edges as small batch training subgraphs. Meta-training subgraphs are more important for large-scale graph data. By training important subgraphs, the learning speed of the model can be accelerated.

[0076] 2-tuple training

[0077] Meta-training refers to the training of meta-training subgraphs. Large-scale graph data is divided and important graph data is selected as meta-training subgraphs. These subgraphs may contain more representative structures in the data. Meta-training can enable the model to learn more generalized rules. The model can quickly adapt to the remaining subgraphs and can converge quickly during small batch training, saving training time and improving learning efficiency.

[0078] Previous experiments have divided the graph data into N subgraphs. During meta-training, a subgraph is used as a task, and the tasks are independent. Assume that the learning function of each task is f, the loss function is L(f), and the objective function is min L(f). The goal is to make it as small as possible.

[0079] The initial parameter of the task is θ0, and for task T1:

[0080]

[0081] Where α is the step size, is the gradient of the loss function for the support set in task T1, θ1 is the parameter obtained after one iteration update, and the iteration may be multiple times. This process is the inner loop. When convergence occurs after n iterations, at this time:

[0082]

[0083] The updated parameter θ1 after convergence n The query set passed to task T1 is used to calculate its loss L1.

[0084] In meta-learning, the tasks are independent, so the initial parameters of all tasks are θ0. The parameters of the task are updated in the support set of all tasks, and the verification calculation is performed in the query set. Repeat the above steps to obtain the query set loss of all tasks as L1, L2, L3...L N .

[0085] All tasks have an impact on the model, and what is needed is the global optimal rather than the local optimal parameter. Therefore, it is necessary to find a parameter with good generalization ability for small batch training, that is, to minimize the sum of all task loss functions:

[0086]

[0087] Among them, β is the step size. This process can be iterated multiple times. Finally, the updated parameters are passed to the small batch subgraphs as initialization parameters. This process is the outer loop.

[0088] 3. Mini-batch training

[0089] The subgraphs for mini-batch training are the remaining subgraphs except the meta-training subgraphs. After the important subgraphs are trained through meta-training, the last updated parameters are passed to mini-batch training as their initial parameters to train the mini-batch subgraphs.

[0090] In the mini-batch training phase, each subgraph is treated as an independent batch for training. After the training of the previous subgraph is completed, the updated parameters are passed to the next subgraph. The training process is as follows:

[0091] First, the parameters of the mini-batch training are updated. The initialization parameters are the parameters updated at the end of the meta-training. The initial parameters of the mini-batch subgraph are recorded as Input the node features and edge information of each subgraph of the first mini-batch subgraph, as well as its initial parameters Conduct training.

[0092] For each node v, its new feature representation h v l Determined by the feature representation of its neighbor nodes and its own feature representation, this update process can be expressed as:

[0093]

[0094] After the training of each subgraph is completed, the parameters of GNN are updated through the gradient descent algorithm according to the gradient of the loss function. GNN learns by minimizing the loss function. Its loss function L is as follows:

[0095]

[0096] where y v is the true label, is the predicted label. In each iteration, the gradient of the loss function with respect to the model parameters is calculated, and the parameters are updated accordingly to gradually approach the optimal solution. The training process will continue until the convergence condition is met, that is, the value of the loss function no longer decreases significantly or reaches the preset number of iterations. In this way, the present invention ensures the stability and generalization ability of the model during the training process.

[0097] The small batch training stage not only optimizes the representation of the entire graph through fine training of each subgraph, but also improves the training efficiency, especially when processing large-scale graph data. Through the initial parameters obtained by meta-training, the present invention can accelerate the convergence speed of small batch training and ensure that the model can capture more subtle structural information in the graph.

[0098] Example 2 Experiment

[0099] 1 Dataset

[0100] In order to verify the effectiveness and efficiency of meta-learning for training local subgraphs, node classification experiments were conducted on five real-world public datasets. For a detailed introduction of each dataset, please see Table 1.

[0101] Table 1 Dataset statistics

[0102]

[0103] The sources of the data sets in Table 1 are:

[0104] AmazonCoBuyComputer: A co-purchase graph dataset extracted from Amazon. In this dataset, nodes represent products, edges represent co-purchase relationships between products, and features are bag-of-words vectors extracted from product reviews. This dataset is often used for research and development in the field of machine learning and data science.

[0105] CoauthorCS: is a co-authorship network dataset based on Microsoft Academic Graph, which is extracted from the KDD Cup Challenge in 2016. This dataset is specifically for the field of Computer Science (CS), where nodes represent authors, and if two authors co-author a paper, there will be an edge connecting them. The features of the nodes represent the keywords of each author's paper, and the class labels indicate the most active research field of each author.

[0106] Flickr: is a large social network dataset based on the Flickr online social platform. It contains millions of user nodes and rich social relationship edges. Each node represents a user, and the edge represents the social connection between users, such as friend relationships. In addition, the dataset may also include photos uploaded by users and their attribute information, such as tags, descriptions, etc. It is often used in social network analysis, graph data mining and graph neural network research.

[0107] Reddit: is a popular graph dataset derived from posts on the Reddit forum in September 2014. The characteristic of this dataset is that it uses posts on the Reddit forum as nodes. If two posts are commented on by the same user, then the two posts are considered to be related in the graph. The Reddit dataset is very valuable for studying community detection, social network analysis, and graph neural networks.

[0108] ogbn-products: is part of the Open Graph Benchmark (OGB) dataset, a collection of benchmark datasets for graph machine learning (ML). ogbn-products is an unweighted, undirected graph based on the Amazon product joint purchasing network. In this dataset, nodes represent products on Amazon, and the presence of edges indicates that two products are often purchased together. The OGB dataset is a series of challenging, realistic, large-scale, and diverse benchmark datasets designed to promote scalability, robustness, and reproducibility research in graph machine learning. These datasets cover a variety of fields, from social and information networks to biological networks, molecular graphs, source code abstract syntax trees (ASTs), and knowledge graphs.

[0109] 2 Evaluation indicators

[0110] In the experimental evaluation, the present invention adopts three key evaluation indicators to comprehensively measure the performance and efficiency of the model: the accuracy of the test set, the average training time per cycle, and the GPU memory consumption during model training.

[0111] The accuracy of the test set is an important indicator for measuring the generalization ability of the model. It reflects the performance of the model on unseen data, and is usually expressed as the ratio of the number of correctly predicted samples to the total number of test samples. The average training time per cycle is an indicator for evaluating the efficiency of model training, which measures the time required for the model in a single training cycle. The memory consumption of the GPU during model training is an indicator for evaluating the efficiency of model resource utilization, which can understand the model's demand for hardware resources during training. Through the evaluation of these three dimensions, the present invention can fully understand the performance, training speed and resource consumption of the model.

[0112] 3 Experimental setup

[0113] Experimental Environment: All experiments are performed on a machine equipped with an NVIDIA GeForce RTX3090 GPU (24GB GPU memory), an Intel Xeon Silver 4214R CPU (12 cores, 2.40GHz), and 256GB RAM.

[0114] Parameter settings: GNN on all datasets has 2 layers and 128 hidden units. Dropout rate is 0.5. ADAM is used as optimizer. Learning rate is 0.001 and weight decay is 0.0005. The ratio of meta-training subgraph to mini-batch subgraph is 2:3. Stop training when the accuracy of validation set hardly changes.

[0115] 4 Experimental results

[0116] The model of the present invention is tested on five datasets together with Cluster-GCN and full batch. The experimental results are as follows: Figure 2 Table 2 is a graph and table of node classification accuracy, Table 3 is a time comparison table, and Table 4 is a memory consumption comparison table.

[0117] Table 2. Node classification accuracy (%) on different datasets. The best and second-best results are highlighted in bold and underlined.

[0118]

[0119] Table 3. Comparison of average training time (s) per batch on different datasets. The best and second-best results are highlighted in bold and underlined.

[0120]

[0121] Table 4. Comparison of GPU memory (MB) consumption when training on different datasets. The best and second best results are highlighted in bold and underlined.

[0122]

[0123]

[0124] Compared with the full batch: In terms of accuracy, the accuracy difference between Meta-GNN and the full batch in the above five data sets is -1.2% to +4.36%. It can be seen that the accuracy difference of Meta-GNN is not large, and it is even 4.36% higher in Ogbn-products; in terms of training time, the full batch is the fastest; but the memory space required for the full batch is very large, especially when the data is large. It can be seen in Table 4 that the memory of Meta-GNN in Ogbn-products is only 4.63% of the full batch. When using the full batch to process graph data with billions of points, it is obviously unable to meet the memory requirements.

[0125] Compared with Cluster-GCN: In terms of accuracy, the accuracy difference between Meta-GNN and Cluster-GCN in the above five data sets is -0.96% to +2.68%, and the accuracy of four of the data sets is not worse than Cluster-GCN; in terms of training time, although Meta-GNN is much slower than the full batch, it is about half faster than Cluster-GCN; because Meta-GNN needs to train the meta-training subgraph, its memory requirement is greater than Cluster-GCN.

[0126] In summary, we can see that Meta-GNN is a good choice when processing memory requirements and needing to learn large-scale graph data faster. In addition to ensuring accuracy, it can save a lot of memory space compared to a full batch and save a lot of time compared to Cluster-GCN.

[0127] The above contents are merely examples and explanations of the concept of the present invention. The technicians in this technical field may make various modifications or additions to the specific embodiments described or replace them in a similar manner. As long as they do not deviate from the concept of the invention or exceed the scope defined by the claims, they should all fall within the protection scope of the present invention.

Claims

1. A large-scale graph data processing method based on the Meta-GNN model, characterized in that: The steps include: Step 1: Dataset processing A node classification experiment was conducted on the datasets. Each dataset was divided into a training set, a validation set, and a test set. The training set was used for model training and optimization, and the test set was used to evaluate the performance of the model. Step 2: Model parameter setting The GNNs on all datasets have 2 layers and 128 hidden units; The dropout rate is 0.5; the learning rate is 0.001, and the weight decay is 0.0005; the ratio of meta-training subgraphs to mini-batch subgraphs is 2:3; training is stopped when the accuracy of the validation set has hardly changed; Step 3: Model training Subgraph division and selection: Subgraph division is to use Metis to divide the large model graph data into many local subgraphs. Subgraph selection is to select the subgraphs with large edge degrees as meta-training subgraphs, and the remaining subgraphs are small batch training subgraphs. Meta-training: Meta-training is training the meta-training subgraphs. The meta-training model can learn more generalized rules and quickly adapt to small-batch training subgraphs. The parameters of the subgraphs obtained by meta-learning are used to learn small-batch training subgraphs to improve learning efficiency. Mini-batch training: Mini-batch training is to train the remaining subgraphs of the meta-training subgraph. Its initial parameters are the parameters updated by the meta-training. Each mini-batch training subgraph is independent and interdependent. Step 4: Prediction and evaluation Prediction stage: predict the nodes in the test set, regard the output of GNN as the final representation of the node, and use the softmax function to calculate the probability of each label; Evaluation metrics: The performance of the model is evaluated using evaluation metrics such as the accuracy of the test set, the average training time per cycle, and the GPU memory consumption during model training to reflect the classification accuracy and generalization ability of the model.

2. The large-scale graph data processing method according to claim 1, characterized in that: The specific steps of dividing and selecting subgraphs in step 3 are as follows: Subgraph partitioning uses Metis to partition the large model graph data into many local subgraphs, and selects the subgraph with the largest number of edges from the partitioned subgraphs as the important subgraphs. The selector of the important subgraphs calculates and sorts the number of edges of the local subgraphs, and selects them according to the ranking of the number of edges. In the partitioned local subgraphs, if there is an edge between vertices i and j, it is recorded as A[i][j]=1, and if there is no edge, it is recorded as A[i][j]=0. The number of edges E is: The number of edges of all local subgraphs is calculated and sorted, and the local subgraphs with a large number of edges are selected as meta-training subgraphs, and the local subgraphs with a small number of edges are selected as mini-batch training subgraphs.

3. The large-scale graph data processing method according to claim 2, characterized in that: The specific steps of meta-training in step 3 are as follows: The graph data is divided into N subgraphs. During meta-training, a subgraph is used as a task. Tasks are independent. Assume that the learning function of each task is f, the loss function is L(f), and the objective function is min L(f). The goal is to make it as small as possible. The initial parameter of the task is θ0, and for task T1: Where α is the step size, is the gradient of the loss function of the support set in task T1, θ1 is the parameter obtained after one iteration update, and the iteration may be multiple times. This process is the inner loop; when it converges after n iterations, at this time: The updated parameter θ1 after convergence n The query set passed to task T1 calculates its loss L1; In meta-learning, the tasks are independent, so the initial parameters of all tasks are θ0. The parameters of the task are updated in the support set of all tasks, and the verification calculation is performed in the query set. Repeat the above steps to obtain the query set loss of all tasks as L1, L2, L3...L N ; All tasks have an impact on the model, and what is needed is the global optimal rather than the local optimal parameter. Therefore, it is necessary to find a parameter with good generalization ability for small batch training, that is, to minimize the sum of all task loss functions: Among them, β is the step size. This process can be iterated multiple times. Finally, the updated parameters are passed to the small batch subgraphs as initialization parameters. This process is the outer loop.

4. The large-scale graph data processing method according to claim 3, characterized in that: The specific steps of small batch training in step 3 are as follows: In the mini-batch training phase, each subgraph is treated as an independent batch for training. After the training of the previous subgraph is completed, the updated parameters are passed to the next subgraph. The training process is as follows: First, the parameters of the mini-batch training are updated. The initialization parameters are the parameters updated at the end of the meta-training. The initial parameters of the mini-batch subgraph are recorded as Input the node features and edge information of each subgraph of the first mini-batch subgraph, as well as its initial parameters Conduct training; For each node v, its new feature representation h v l The update process is determined by the feature representation of its neighbor nodes and its own feature representation. After the training of each subgraph is completed, the parameters of GNN are updated through the gradient descent algorithm according to the gradient of the loss function. GNN learns by minimizing the loss function. Its loss function L is as follows: where y v is the true label, is the predicted label; In each iteration, the gradient of the loss function with respect to the model parameters is calculated, and the parameters are updated accordingly to gradually approach the optimal solution; the training process will continue until the convergence condition is met, that is, the value of the loss function no longer decreases significantly or reaches the preset number of iterations.