A method for solving the problem of unbalanced data partitioning in heterogeneous graph neural networks for massive data

By dividing heterogeneous graph data into type subgraphs and optimizing using METIS algorithm, the problem of data imbalance in graph neural network training is solved, load balancing and low communication frequency are achieved, and model training efficiency and accuracy are improved.

CN116432738BActive Publication Date: 2025-07-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310211374.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-07
Publication Date
2025-07-25
Estimated Expiration
2043-03-07

AI Technical Summary

Technical Problem

When training large-scale heterogeneous graph data, existing graph neural network models have unbalanced data division problems, resulting in imbalance in traffic and computational volume, affecting training time and model accuracy.

Method used

Edge-based interaction relationships are used to divide the entire graph data into various types of subgraphs, quickly divide it using the METIS algorithm, and determine the best merger solution by calculating the maximum benefits between the matrices, ultimately achieving balanced partitioning of large-scale data and reducing cross-partition communication.

Benefits of technology

Implement the equalization of the number and type of nodes under large-scale data, reduce cross-partition communication frequency, and improve model training efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116432738B_ABST
    Figure CN116432738B_ABST
Patent Text Reader

Abstract

The present invention relates to a processing method for solving the unbalanced data partitioning of heterogeneous graph neural networks for massive data, belonging to the field of deep learning. This method first evenly distributes the full graph data to each type of subgraph according to the interaction relationship of the edges, and splits the full graph G into G 1 , G 2 , …, G k ; uses the METIS algorithm based on the balancing strategy for rapid partitioning to solve the weighted k-way graph partitioning problem; calculates the maximum benefit between matrices according to the results of each type of subgraph after partitioning to obtain the optimal merging scheme; queries the original data set according to the merging scheme to combine the node and edge sets, and finally realizes the balanced partitioning of large-scale data to obtain the corresponding partitioning results. The present invention can be used for large-scale data partitioning. The partitioning results have roughly the same number and type of nodes in each partition, have good load balancing characteristics, have a low critical node replication ratio, effectively reduce the cross-partition communication volume, can better support the training of heterogeneous graph neural network models, and reduce the overall training time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a processing method for solving the problem of unbalanced data partitioning of heterogeneous graph neural networks for massive data, and belongs to the field of deep learning. Background Art

[0002] As of June 2022, the number of Internet users in China reached 1.051 billion, and the Internet penetration rate reached 74.4%. The mobile Internet presents a new development trend. With the popularization of the Internet, user data is constantly increasing, and higher-level user behavior analysis and user category differentiation are required. To improve the classification and prediction results, the number of model parameters increases, resulting in longer training time and higher computing resource requirements.

[0003] Graph neural network technology has been proven to be an effective tool for processing non-Euclidean graph data and is widely used in many fields such as search, recommendation, and risk control. However, due to the large scale of training data and long training time of the graph neural network model, distributed training has become a good choice, which can utilize multiple machines for parallel training to solve the problem that a single machine cannot train quickly and independently.

[0004] In the main process of distributed graph neural network training, the first step is to partition the data into each partition to distribute the "subgraphs" of the graph data on different computing nodes. Graph partitioning helps to decompose a large-scale graph dataset into multiple subsets that can be computed in parallel, thereby realizing the distributed training process. Most existing studies have shown that due to problems such as unbalanced data distribution and low data correlation during the machine training process, the imbalance between communication volume and computing volume will occur. On the one hand, due to the increased parameter synchronization waiting time caused by unbalanced data allocation, the training time of the model will also increase. On the other hand, improper data partitioning will also cause other impacts. For example, during the model training process of each computing node, there is likely to be a large difference in training results, which will introduce a greater error in the parameter synchronization process, thereby affecting the accuracy of the overall model. Therefore, when performing graph partitioning, the computing volume and communication volume must be carefully considered to balance the two, and further ensure the efficiency and accuracy of model training.

[0005] Currently, most distributed graph neural network frameworks use METIS for fast graph partitioning when dividing data. METIS is a simple partitioning algorithm that, based on the idea of multilevel encoding, aims to decompose the partitioning task of large sparse matrices into a series of multilevel partitioning sub-tasks that are easier to solve, making it have a relatively low space complexity. DistDGL optimizes the above partitioning method. By optimizing densely connected nodes, it assigns a unique partition to highly frequently accessed hot nodes and replicates vertices that are not core nodes but have a certain frequency of cross-partition access, thereby ensuring that the neighbors of local vertices in each partition are accessible. This can reduce cross-partition communication access during subsequent training sampling. In addition, it also reduces the retention of low-quality edges and retains high-weight edges to improve overall performance. A constraint mechanism is used to support balanced partitioning of edges and nodes to make their numbers similar. However, this method has good performance results for graph partitioning with a small number of relationships. In heterogeneous graph datasets with complex relationships, the combination of node and edge types results in multiple combinations of meta-paths, which may cause the partitioning results to not meet expectations.

[0006] Generally speaking, most graph partitioning methods for distributed graph neural networks consider homogeneous graphs, while there are relatively few studies on the partitioning of heterogeneous graphs. In heterogeneous graphs, due to the existence of heterogeneous nodes and differences in the number of nodes, using existing partitioning methods often leads to unbalanced graph partitioning results, resulting in more frequent data exchange between different computing nodes, thereby affecting the accuracy and training efficiency of graph neural networks. Therefore, this paper proposes a large-scale heterogeneous graph balanced partitioning method. From the perspective of splitting and aggregation, it reasonably partitions large-scale heterogeneous graph data to make it have data balance, improve data correlation, and reduce the cross-partition interaction frequency of subsequent models. Summary of the Invention

[0007] To overcome the deficiencies of the prior art, the present invention provides a method for solving the unbalanced processing of heterogeneous graph neural network data partitioning for massive data. This method can be applied to heterogeneous graph neural network data in scenarios such as social networks, academic networks, and commodity trading networks. Among them, this type of data has heterogeneous nodes, heterogeneous relationships, and feature vectors of nodes and themselves. First, according to the interaction relationships of edges, the whole graph data is evenly divided into subgraphs of each type. The whole graph G is split into G1, G2, …, G according to type dependency relationships. k; then use the METIS algorithm based on the balanced strategy for fast partitioning to solve the weighted k-way graph partitioning problem; next, calculate the maximum revenue between matrices according to the results of each type of subgraph obtained by partitioning to obtain the best merging scheme; finally, query the original dataset according to the merging scheme to combine the node and edge sets, and finally achieve the balanced partitioning of large-scale data to obtain the corresponding partitioning results. The present invention can be used for large-scale data partitioning, and the partitioning results have roughly the same number and type of nodes in each partition, with good load balancing characteristics, a low replication ratio for critical nodes, and effectively reduce the cross-partition communication volume, and can better support the training of heterogeneous graph neural network models and reduce the overall training time.

[0008] The technical solution adopted by the present invention to solve its technical problems includes the following steps:

[0009] Step 1: Load the heterogeneous graph neural network data G into the memory for subsequent processing. The graph neural network data G is academic paper data or social network data or commodity trading network data, and there are heterogeneous nodes, heterogeneous relationships, and the feature vectors of the nodes and themselves.

[0010] Step 2: Divide the whole graph data evenly into each type of subgraph according to the interaction relationship of the edges, and split the whole graph G into G1, G2,..., G k according to the type dependence relationship, which is used to simplify the partitioning difficulty and reduce the possible imbalance problem in the initial partitioning of the original graph, and provide the dependence relationship of data calculation for the subsequent step of merging blocks;

[0011] Step 2-1: For the edge type φ(e i ) belonging to the graph G, obtain the node types τ(n i ), τ(n j ) on both sides of a certain type of edge φ(e k );

[0012] Step 2-2: Obtain all the relationship edges of the type τ(n j ), τ(n k ) obtained in Step 2-1 and their related edges in the whole graph G;

[0013] Step 2-3: Split the whole graph G type into multiple heterogeneous type split graphs G1, G2,..., G k such that for any split graph G i , there exists a unique τ(n k ) ∧ φ(n) ≠ φ(k), and for any split graph G i in all split graphs, there exists a graph G j such that there is the same node type between the two graphs;

[0014] Step 3: Use the METIS algorithm with a multi-constraint equilibrium strategy to partition the final G1, G2, …, G obtained in Step 2 k individually for graph partitioning. For each split graph G i , the partitioned results G i,j of each partition can be obtained, where G i,j represents the j-th partition in the i-th split graph;

[0015] Step 3-1: Use the METIS partitioning method for graph G i . When setting parameters, specify according to the characteristics of the heterogeneous dataset. For heterogeneous graphs, it is recommended to set edge balance and vertex balance;

[0016] Step 3-2: To reduce communication overhead, use the edge replication method in DistDGL to perform a certain proportion of replication processing on edge nodes to retain the information of other partition nodes of the connected edges.

[0017] Step 4: Calculate the optimal heterogeneous graph merging interval scheme according to the graph partitioning results in Step 3;

[0018] Step 4-1: In Step 3, the partitioned results G i,j of each partition can be obtained, where G i,j represents the j-th partition in the i-th split graph;

[0019] Step 4-2: According to the correlation between nodes in Step 2-3, for G i and G j in all split graphs that have the relationship described in Step 2-3, where the partitions of the two graphs are denoted as G i,k and G j,l . According to G i,k and G j,l , the matrix A can be calculated, where each element a i,j in the matrix A represents the specific number of overlapping nodes between each partition in G i,k and G j,l .

[0020] Step 4-3: Use the constraint function to calculate each partition to obtain the optimal merging scheme, where the constraint function is shown as follows. Maximize means to obtain the maximum value, and subject to represents the constraint condition.

[0021]

[0022]

[0023] Step 4-4: According to the maximum benefit z value obtained in Step 4-3, record the calculation matrix X at that time, and each of its elements represents x i,j, whose value is represented as 1 for taking and 0 for not taking, that is, the scheme for merging intervals.

[0024] Step 5: Construct a partitioned subgraph according to the scheme for merging intervals obtained in Step 4.

[0025] Step 5-1: According to the merging scheme obtained in Step 4-4, query each node in the original graph to obtain the node sets V of each subgraph i

[0026] Step 5-2: According to the node set V obtained in Step 5-1 i , query the dependency relationships between each node in the original graph to obtain the subgraph edge set E i

[0027] Step 5-3: According to the node set V in Step 5-1 i and the edge set E in Step 5-2 i , a partitioned subgraph G can be constructed i =(V i , E i ).

[0028] The beneficial effects of the present invention are as follows:

[0029] The present invention can be used under large-scale data. The method splits various types of nodes and quickly partitions each graph, and then statistically calculates the overlap degree of each type of node set after partitioning to combine the node sets of each type, so as to obtain the partitioning result. It has a certain effectiveness. Due to reasonable partitioning, in distributed computing, the amount of data read across computing nodes is reduced, thereby reducing the communication frequency between partitions. On this basis, the partition replication rate is not increased significantly, and finally the partitioning result can be provided to different distributed nodes and combined with different heterogeneous graph neural network models for subsequent training. Description of the Drawings

[0030] Figure 1 is a schematic diagram of the meaning of nodes in the embodiment;

[0031] Figure 2 is a schematic diagram of splitting the entire graph by type in the embodiment;

[0032] Figure 3 is a schematic diagram of partitioning each type of split graph in the embodiment;

[0033] Figure 4 is a schematic diagram of calculating the partition repetition degree in the embodiment;

[0034] Figure 5 is a schematic diagram of generating a partitioned subgraph according to points querying the original data set in the embodiment; Detailed Embodiments

[0035] The present invention will be further described below with reference to the drawings and embodiments.

[0036] Specifically, as Figure 1 shown, the heterogeneous graph dataset in this embodiment is an academic dataset, and the data it contains are papers, authors, institutions, affiliated journals, etc.; there are corresponding relationships between the edges and nodes, such as <paper, affiliated field, field>, <paper, affiliated journal, journal>, <author, writes, paper>, <author, affiliated with, institution>, etc.

[0037] Step 1: Assume that the entire graph of the heterogeneous graph dataset is G, the graph G contains the corresponding node set N, edge set E, and corresponding feature information. Assume that a certain node among them is n, and the two connected nodes around it are n1 and n2. The node type of n is τ(n), and the node types of the similar n1 and n2 are τ(n1) and τ(n2). Assume that the connected edges between n1, n2 and the node n are e1 and e2 respectively, where the edge types of e1 and e2 are φ(e1) and φ(e2) respectively, and e1 = (n1, t), e2 = (n2, t). The relevant structure starting from the target node can be called a meta-relationship, simply denoted as <τ(n1), φ(n), τ(n2)>. Load the above graph data G into memory for subsequent processing. Specifically, the node set N in the example is a specific node set such as papers, authors, institutions, affiliated journals, etc., and the relationships included in the edge set E in the example are <paper, affiliated field, field>, <paper, affiliated journal, journal>, <author, writes, paper>, <author, affiliated with, institution>, etc.

[0038] Step 2: Divide the entire graph data evenly into subgraphs of each type according to the interaction relationship of the edges. Split the entire graph G into G1, G2,..., G according to the type dependency relationship k to simplify the division difficulty and reduce the possible imbalance problem in the initial division of the original graph, and provide a dependency relationship for data calculation for the subsequent step of merging blocks;

[0039] Step 2-1: For the edge type φ(e i ), obtain the node types τ(n i ), τ(n j ) on both sides of a certain type of edge φ(e k ). Specifically, assume that the obtained edge type is "writes", then it is necessary to obtain the node types "author" and "paper" around the "writes" relationship;

[0040] Step 2-2: In the entire graph G, obtain all the relationship edges of the types τ(n j ), τ(n k ) obtained in Step 2-1 and their related edges;

[0041] Step 2-3: Split the entire graph G type into multiple heterogeneous type split graphs G1, G2,..., G ksuch that for any split graph G i , there exists a unique τ(n k ) ∧ φ(n) ≠ φ(k), and for any split graph G i in all split graphs, there exists a graph G j such that there is the same node type between the two graphs;

[0042] Specifically, assume that there is a certain meaning represented by node colors as Figure 1 shown, and there are corresponding relationships between the edges and nodes, such as <paper, field of study, field>, <paper, journal, journal>, <author, write, paper>, <author, affiliated with, institution>, etc. As Figure 2 shown, the graph obtained by decomposing the whole graph according to the edge relationship is called a homogeneous split graph. According to this split method, three split graphs can be obtained, namely, graph G1 is the "field of study - paper" split graph, graph G2 is the "paper - author" split graph, and graph G3 is the "author - research institution" split graph. Among them, graph G1 and graph G2 have the same node "paper", graph G2 and graph G3 have the same node "author", there is a unique relationship between heterogeneous nodes in each graph, and there is only one node as the connection between the two graphs;

[0043] Step 3: Use the METIS algorithm with a multi - constraint equilibrium strategy to separately partition the G1, G2,..., G k finally obtained in Step 2. For each split graph G i , the partition results of each partition G i,j can be obtained, where G i,j represents the j - th partition in the i - th split graph;

[0044] Step 3 - 1: Use the METIS partitioning method for graph G i . When setting parameters, specify according to the characteristics of the heterogeneous data set. For a heterogeneous graph, it is recommended to set edge balance and vertex balance;

[0045] Step 3 - 2: To reduce communication overhead, use the edge replication method in DistDGL to perform a certain proportion of replication processing on the edge nodes to retain the information of other partition nodes of the connected edges. The partitioning schematic diagram is as Figure 3 shown.

[0046] Step 4: Calculate the best heterogeneous graph merging interval scheme according to the graph partitioning results in Step 3;

[0047] Step 4 - 1: In Step 3, the partition results of each partition G i,j can be obtained, where G i,j represents the j - th partition in the i - th split graph;

[0048] Step 4-2: According to the relevance between nodes in Step 2-3, for Gs in all split graphs that have the relationship described in Step 2-3 i and G j , where the partitions of the two graphs are denoted as G i,k and G j,l , G i,k is the k-th partition of the i-th graph, and G j,l is the l-th partition of the j-th graph. According to G i,k and G j,l , the matrix A can be calculated, where each element a i,j in the matrix A represents the specific number of overlapping nodes between the respective partitions in G i,k and G j,l .

[0049] Step 4-3: Use the constraint function to calculate each partition to obtain the optimal merging scheme, where the constraint function is shown as follows, maximize means to obtain the maximum value, and subject to represents the constraint conditions.

[0050]

[0051]

[0052] Specifically in this step, assume there is an N×N matrix a, and the objective function is given as where x i,j is a binary variable indicating whether the number in the i-th row and j-th column of the matrix is taken. Let 1 represent taken and 0 represent not taken, and a i,j is the number in the i-th row and j-th column of the matrix. It is necessary to obtain the maximum value of z when satisfying the objectives and . Intuitively, the higher the calculated repetition degree in the matrix, the higher the overall correlation between nodes, and thus a beneficial fusion scheme can be provided for the split results of heterogeneous graph types.

[0053] Step 4-4: According to the maximum benefit z value obtained in Step 4-3, record the calculation matrix X at that time, and each of its elements represents x i,j , whose value represents taken as 1 and not taken as 0, that is, the scheme for merging intervals, and its calculation result is as shown in Figure 4 .

[0054] Step 5: Construct subgraph partitions according to the scheme for merging intervals obtained in Step 4.

[0055] Step 5-1: According to the merging scheme obtained in Step 4-4, query each node in the original graph to obtain the node sets V of each subgraph i

[0056] Step 5-2: According to the node set V obtained in Step 5-1 i , query the dependency relationships between the nodes in the original graph to obtain the subgraph edge set E i

[0057] Step 5-3: According to the node set V in Step 5-1 i and the edge set E in Step 5-2 i , the partitioned subgraph G i =(V i , E i ) can be constructed, that is Figure 5 the partitioning result in

[0058] It can be seen from this example that the academic dataset is reasonably partitioned, which can reduce the single-round training duration and communication frequency for downstream distributed model training.

Claims

1. A method for solving the problem of uneven data partitioning in heterogeneous graph neural networks for massive data, characterized in that, It includes the following steps: Step 1: Load heterogeneous graph neural network data into memory for subsequent processing. The graph neural network data is academic paper data. The graph contains the corresponding point set and edge set , and the corresponding feature information. Assume that a certain node among them is , and the two connected points around it are and ; The node type is , and similarly the node type is ; Assume and the connected edges of node are respectively , where the edge types of are respectively , , and ; The relevant structure starting from the target node can be called a meta-relationship, simply denoted as ; Load the above graph data into memory for subsequent processing; specifically, the point set is the specific node set of papers, authors, institutions, and affiliated journals, and the relationships contained in the edge set are <paper, affiliated field, field>, <paper, affiliated journal, journal>, <author, writes, paper>, <author, affiliated with, institution>;​​​ Step 2: Divide the whole graph data evenly into each type of subgraph according to the edge interaction relationship, and split the whole graph according to the type dependency relationship into , which is used to simplify the partitioning difficulty and reduce the possible imbalance problems in the initial partitioning of the original graph, and provide the dependency relationship for data calculation in the subsequent step of merging blocks; Step 2-1: For the edge type belonging to Figure , obtain the node types on both sides of a certain type of edge ; Step 2-2: In the whole graph obtain all the relationship edges of the type obtained in Step 2-1 and its related edges; Step 2-3: Split the whole graph type into multiple heterogeneous type split graphs such that for any split graph there exists a unique and in all split graphs, for any split graph there exists a graph such that there is a same node type between the two graphs; Step 3: Use the METIS algorithm with a multi-constraint equilibrium strategy to perform graph partitioning on the final result obtained in Step 2 separately for each graph, and for each split graph , the split results of each partition can be obtained , where represents the th partition in the th split graph; Step 3-1: For the graph Use the METIS partitioning method and specify parameters according to the characteristics of the heterogeneous dataset. For heterogeneous graphs, it is recommended to set edge balance and vertex balance. Step 3-2: To reduce communication overhead, use the edge replication method in DistDGL to perform a certain proportion of replication processing on edge nodes to retain the information of other partition nodes of the connected edges; Step 4: According to the graph partitioning results in Step 3, calculate the optimal heterogeneous graph merging interval scheme; Step 4-1: The split results of each partition can be obtained in Step 3 , where represents the th partitioning partition in the th split diagram; Step 4-2: According to the relevance between nodes in Step 2-3, for those in all split graphs that have the relationship described in Step 2-3 and , where the partitions of the two graphs are denoted as and ; According to and , matrix can be calculated, where each element in matrix represents and the specific number of overlapping nodes between the respective partitions; Step 4-3: Use the constraint function to calculate each partition to obtain the optimal merging scheme, where the constraint function is shown as follows, maximize means to obtain the maximum value, and subject to represents the constraint conditions; Step 4-4: According to the maximum profit obtained in Step 4-3 value, record the calculation matrix at that time , each of its elements represents , and its value is expressed as take, do not take, that is, the scheme of merging intervals; Step 5: Construct a partitioned subgraph according to the scheme of the merging interval obtained in Step 4; Step 5-1: According to the merging scheme obtained in Step 4-4, query each node in the original graph to obtain each sub-graph node set ; Step 5-2: According to the node set obtained in Step 5-1 , query the dependency relationships between the nodes in the original graph to obtain the subgraph edge set ; Step 5-3: According to the node set in Step 5-1 and the edge set in Step 5-2 , the subgraph can be constructed .

Citation Information

Patent Citations

  • Link prediction method based on topic perception heterogeneous graph neural network

    CN113672735A

  • Graph neural network sampling method for large-scale heterogeneous graph

    CN115423073A