A distributed graph convolutional neural network implementation method for feature segmentation

Through the distributed graph convolution neural network method of feature slicing, the input features are segmented into multiple feature fragments and allocated to multiple computing nodes, solving the problem of inefficient training of large-scale graph data, and achieving the effect of improving training efficiency while maintaining or improving model accuracy.

CN119106702BActive Publication Date: 2025-06-06YIZHENG POWER SUPPLY OF JIANGSU ELECTRIC POWER +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411137148.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2025-06-06
Estimated Expiration
2044-08-19

AI Technical Summary

Technical Problem

When processing large-scale graph data, the training efficiency of the GNN model is inefficient, and due to insufficient memory, it is difficult to effectively train on a single computing device, which affects the scalability and practical application convenience of the model.

Method used

A distributed graph convolutional neural network method using feature slicing is used to generate feature slicing strategies, and the input features are segmented into multiple feature fragments and allocated to multiple computing nodes. Each computing node independently performs forward propagation calculations, and finally performs splicing and slice encoding on the main node to obtain global node characterization.

Benefits of technology

While maintaining or improving the accuracy of the model, this method can improve training efficiency on large-scale graph data, reduce communication needs between computing nodes, and effectively utilize complete full batch graph data for training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119106702B_ABST
    Figure CN119106702B_ABST
Patent Text Reader

Abstract

The present invention provides a distributed graph convolutional neural network method for feature segmentation. First, the input features are divided into multiple parts and distributed to different computing nodes. The complete graph structure and a small part of node features are loaded on each computing node, thereby reducing memory requirements; then, the GCN slicing model on each computing node performs forward propagation on its feature fragments, and the communication between computing nodes is only performed in the input and output stages; finally, the output representations of all computing nodes are transmitted to the master node for splicing, and the output representations of each block of GCN are uniformly adjusted through slice coding, ultimately improving the training efficiency of graph neural networks on large-scale graph data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of distributed deep learning technology, and in particular to a distributed graph convolutional neural network method for feature segmentation. Background Art

[0002] At present, the application of graph neural network (GNN) is becoming more and more widespread, among which the model represented by GCN (graph convolutional neural network) is particularly popular and has become the de facto standard in the field of GNN. GCN is based on the message passing mechanism and mainly includes two steps: aggregation and update. GCN first aggregates the features of a node and all its neighboring nodes, and then updates the representation of the node through learnable weights. By stacking several layers, GCN is able to capture the dependencies between distant nodes, thereby demonstrating excellent performance in various tasks such as node classification, graph classification, and link prediction. However, the GNN model still faces the problem of low training efficiency when processing large-scale graph data. The aggregation process of message passing is the bottleneck operation in GNN, which limits the scalability of GNN. As the number of layers increases, the dependencies of nodes become more complex, requiring more neighborhood node information. Therefore, the naive training method of GCN is to input all node features in the graph, that is, the complete graph, to provide complete dependencies and ensure the best accuracy of the model. However, on a single computing device, the complete graph data will lead to an out of memory problem (OOM), which results in a large amount of computing and memory resources required for training large-scale graphs. This limits the convenience of the GNN method in practical applications and also limits the exploration of more complex GNN-based graph representation learning methods.

[0003] To solve these problems, the above-mentioned full batch learning can be changed to small batch learning. For example, small batches can be created by sampling neighbors to reduce computing and memory requirements. However, sampling-based methods will bring feature approximation errors, which will affect the performance of the model. Therefore, for GNN, the complete graph, that is, full batch learning, is the most ideal. Through distributed training, more computing resources can be called to process larger-scale graph data, and the training process can also be accelerated. Existing distributed GNN training methods often consider dividing a large graph into multiple subgraphs, each of which is assigned to a computing node. In order to ensure the accuracy of model training, it is necessary to exchange information about neighbor nodes between computing nodes to ensure the integrity of neighbor node information. However, this exchange operation will cause additional communication overhead, which will affect training efficiency; while the use of asynchronous updates can make computing and communication overlap as much as possible, but it will result in a loss of accuracy.

[0004] Therefore, how to design an efficient distributed training framework to improve the training efficiency of GNN while ensuring that the accuracy does not decrease is a challenging problem. Summary of the invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art.

[0006] The technical solution of the present invention is: a distributed graph convolutional neural network method for feature segmentation, comprising the following steps:

[0007] Step 1: Generate a feature segmentation strategy, divide the input features into multiple feature segments according to the feature segmentation strategy, and distribute them to multiple computing nodes;

[0008] Step 2: GCN sharding is initialized. The GCN sharding model on each computing node independently performs forward propagation calculations on the feature fragments.

[0009] Step 3: The output representation of each computing node is transmitted to the master node for splicing, and the output representation of each block of GCN is uniformly adjusted through slice encoding, and input into the classifier for downstream tasks to obtain the global node representation.

[0010] The process of generating a feature segmentation strategy is:

[0011] First, initialize an empty slice list and calculate the initial size of each slice; then, for each computational node, calculate the slice start and end positions corresponding to the node in turn, and adjust the slice size considering the scaling factor. If the slice end position exceeds the maximum value of the feature dimension, adjust the start and end positions accordingly;

[0012] The slicing strategy of each computing node is finally added to the slice list; finally, the generated slicing strategy and the size of each slice are returned.

[0013] When the input features are split into multiple parts according to the generated feature segmentation strategy, it includes: direct segmentation or feature fusion.

[0014] The direct segmentation includes: directly segmenting by columns, and segmenting the input features by columns into the same number of parts as the number of computing nodes.

[0015] The feature fusion includes: using a learnable neural network to map the features to a low-dimensional space in a scaled manner, and the number of copies is also the same as the number of computing nodes.

[0016] The process of GCN shard initialization is:

[0017] The input dimension size of the GCN shard on the computing node is set to the dimension of the feature fragment, the intermediate dimension size of the GCN shard is set according to one-half of the number of computing nodes of the intermediate dimension size of the single-machine GCN, the output dimension is set to one-half of the number of computing nodes of the final node representation dimension size, and then the weight parameters of the GCN shard are initialized.

[0018] The process of forward propagation calculation is:

[0019] For a computing node, the feature fragment it has is used to input the small parameter amount GCN fragment on the computing node, and the output of the small parameter amount GCN is obtained.

[0020] The design process of slice encoding is:

[0021] Initialize a learnable matrix with the number of rows equal to the number of computing nodes. The matrix is ​​slice encoding, and each row is the encoding vector corresponding to a computing node.

[0022] In the work, the present invention divides the initial input of GCN in the feature dimension, so that more graph nodes can be stored on a single computing node to reduce or avoid distributed computing node communication. At the same time, the complete batch of graph data can be used for training, and the exchange of training information between computing nodes is avoided during GCN layer propagation. The communication demand only exists at the beginning and the end, and feature fusion and slice encoding are designed at the beginning and the end of GCN to ensure model accuracy. The present invention can improve training efficiency on large-scale graph data while maintaining or even improving model accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a flow chart of the method of the present invention;

[0024] Figure 2 It is a principle block diagram of the present invention. DETAILED DESCRIPTION

[0025] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0026] The present invention is mainly carried out on large-scale graph data. Figure 2 This is a general framework diagram of a distributed graph convolutional neural network method for feature segmentation in the present invention, which includes the following parts: feature segmentation, GCN segmentation and slice encoding.

[0027] like Figure 1 As shown, the present invention provides a general framework diagram of a distributed graph convolutional neural network method for feature segmentation, comprising the following steps:

[0028] The framework diagram of the present invention is as follows Figure 1As shown in the figure, the input features are divided into p parts, corresponding to p distributed computing nodes or devices (such as GPUs). Each computing node is loaded with a complete graph structure and a small piece of features of all nodes, and the model on each computing node is a GCN sharding model, and the dimension size of each layer is 1 / p of the dimension size of each layer of the corresponding single-machine GCN, which greatly reduces the memory requirement. In each round of iteration, the GCN sharding model on each computing node will forward propagate its own feature fragments, and the communication between computing nodes only occurs before the input of the graph data and after the output of the GCN, that is, at the beginning and end of the GCN.

[0029] The present invention comprises the following steps:

[0030] Step 1: At the beginning of GCN, such as Figure 1 The figure shows two feature segmentation methods of the feature segmentation module: (a) simply and evenly segmenting the input feature X into p parts, and the communication overhead is mainly to copy the segmented features to each computing node; (b) in order to obtain higher accuracy and enable the GCN on each GPU to see richer feature information, the input feature X is subjected to feature fusion, that is, the feature dimension is reduced to 1 / p of the original by using a neural network to achieve the purpose of reducing the memory requirement of a single card. The present invention adopts MLP, namely "FF" (Feedforward Neural Network) in the figure. This method has relatively more computational overhead in addition to feature replication;

[0031] Step 2: If Figure 1 As shown in Figure 1, each GCN shard completes its own calculations during the parallel computing process without the need for communication. On the i-th computing node, there is a smaller GCN shard model parameter Since fewer node features need to be processed, the number of required parameters is adjusted accordingly;

[0032] Step 3: At the end of GCN, Figure 1 As shown, the GCN sharding model on each computing node will represent its own output H i The master node will concatenate the output representations on each GPU and input them into the classifier responsible for downstream tasks to obtain the global node representation H. In addition, considering each group of GCN parameters θ i Each of them is relatively independent, so before aggregation, the present invention designs slice encoding to uniformly adjust the output representation of each GCN block to improve the quality of splicing representation.

[0033] The present invention splits the initial input of GCN in the feature dimension, so that more graph nodes can be stored on a single computing node to reduce or avoid communication. At the same time, a complete full batch of graphs can be used for training, and the exchange of training information between computing nodes is avoided during GCN layer propagation.

[0034] Specifically, step 1 includes:

[0035] For input features Processing to obtain feature slices with smaller storage space Where n is the number of graph nodes, d′ is the feature dimension, and p is the number of computational nodes.

[0036] Directly split by column:

[0037] X i =X[:,start:end]

[0038] The start and end segmentation strategies are generated as follows:

[0039]

[0040] Feature fusion method:

[0041] X i =XW FF

[0042] in, is the learnable parameter matrix of feature fusion.

[0043] Specifically, step 2 includes:

[0044] After processing the input features, the GCN sharding model is initialized accordingly: for the GCN sharding model on the i-th computing node, The learnable aggregate weights at layer l Update weights with representation Then for the node at layer l on the i-th computing node , calculated as follows:

[0045]

[0046] like Figure 1 As shown, after p computing nodes complete the above calculations in parallel, p groups of node representations H can be obtained. i,(L) , where L is the number of layers in each GCN slice.

[0047] Specifically, step 3 includes:

[0048] Considering that there is no communication or interaction between the p GCN shard models, the learned representation patterns may be different. Therefore, at the end of GCN, Figure 1 As shown, the present invention designs a slice coding module for uniformly adjusting the output representation of each GCN block. The calculation process of slice coding to adjust the representation output of the i-th computing node is as follows:

[0049] H i =H i,(L) +E 1,:

[0050] in is the parameter matrix of slice encoding. Finally, the master node represents the output of each GCN block H i After splicing, it is input to the MLP classifier responsible for downstream tasks to complete the forward propagation process and backpropagation training:

[0051] H=CONCAT(H 1 , H 2 , …, H p )

[0052]

[0053] The present invention solves the problem that the traditional GCN model cannot adapt to the balance of efficiency and accuracy in large-scale graph data training, thereby improving the ability of GCN to be applied in real graph analysis task scenarios.

[0054] The present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention. These corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.

Claims

1. A distributed graph convolutional neural network implementation method for feature segmentation, characterized in that: The following steps are involved: Step 1: Generate a feature segmentation strategy, divide the input features into multiple feature fragments according to the feature segmentation strategy, and distribute them to multiple computing nodes, which are GPUs; The process of generating a feature segmentation strategy is: First, initialize an empty slice list and calculate the initial size of each slice; then, for each computational node, calculate the slice start and end positions corresponding to the node in turn, and adjust the slice size considering the scaling factor; if the slice end position exceeds the maximum value of the feature dimension, adjust the start and end positions accordingly; The slicing strategy of each computing node is finally added to the slicing list; finally, the generated slicing strategy and the size of each slice are returned; Step 2: GCN sharding is initialized. The GCN sharding model on each computing node independently performs forward propagation calculations on the feature fragments. The process of GCN shard initialization is: Set the input dimension size of the GCN shard on the computing node to the dimension of the feature fragment, set the intermediate dimension size of the GCN shard to one-half the number of computing nodes of the intermediate dimension size of the single-machine GCN, set the output dimension to one-half the number of computing nodes of the final node representation dimension size, and then initialize the weight parameters of the GCN shard; Step 3: The output representation of each computing node is transmitted to the master node for splicing, and the output representation of each block of GCN is uniformly adjusted through slice encoding, and input into the classifier for downstream tasks to obtain the global node representation.

2. The method for implementing a distributed graph convolutional neural network for feature segmentation according to claim 1, characterized in that: When the input features are split into multiple parts according to the generated feature segmentation strategy, it includes: direct segmentation or feature fusion.

3. The method for implementing a distributed graph convolutional neural network for feature segmentation according to claim 2, characterized in that: The direct segmentation includes: directly segmenting by columns, and segmenting the input features by columns into the same number of parts as the number of computing nodes.

4. The method for implementing a distributed graph convolutional neural network for feature segmentation according to claim 2, characterized in that: The feature fusion includes: using a learnable neural network to map the features to a low-dimensional space in a scaled manner, and the number of copies is also the same as the number of computing nodes.

5. The method for implementing a distributed graph convolutional neural network for feature segmentation according to claim 1, characterized in that: The process of forward propagation calculation is: For a computing node, the feature fragment it has is used to input the small parameter amount GCN fragment on the computing node, and the output of the small parameter amount GCN is obtained.

6. The method for implementing a distributed graph convolutional neural network for feature segmentation according to claim 1, characterized in that: The design process of slice encoding is: Initialize a learnable matrix with the number of rows equal to the number of computing nodes. The matrix is ​​slice encoding, and each row is the encoding vector corresponding to a computing node.

Citation Information

Patent Citations

  • Topological segmentation graph neural network training mechanism based on breadth-first segmentation

    CN113918319A

  • KR20230039290A