Graph neural network stand-alone training method based on hybrid dynamic caching

By adopting a hybrid dynamic cache strategy and a ensemble similarity reordering algorithm in graph neural network training, the problems of large overhead and low static cache hit rate in mini-batch preparation stage are solved, and more efficient graph neural network training and GPU resource utilization are achieved.

CN120087454APending Publication Date: 2025-06-03EAST CHINA NORMAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510200440.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In the training of graph neural networks, the overhead of mini-batch preparation stage is high, and the static cache strategy has a low hit rate under the load of a multi-sampling algorithm, resulting in low training efficiency.

Method used

The hybrid dynamic cache strategy is adopted to dynamically adjust the proportion of node feature cache and static adjacency matrix cache by pre-executing the access frequency of node features and adjacency matrix items in the statistical graph data, and to accelerate mini-batch generation using set similarity reordering algorithm and first-in first-out queues, and accelerate data transmission with direct memory access (DMA) technology.

Benefits of technology

The hit rate of node feature cache during mini-batch generation is improved, mini-batch preparation time is reduced, more efficient GNN training is achieved, and the utilization rate of GPU devices is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087454A_ABST
    Figure CN120087454A_ABST
Patent Text Reader

Abstract

The invention discloses a graph neural network stand-alone training method based on hybrid dynamic cache, which is characterized by comprising the following steps of: obtaining the proportion of dynamic node feature cache to static adjacent matrix item cache by using graph nodes and adjacent matrix item access frequency information in the pre-training stage; in the training stage, a set similarity reordering algorithm is used for rearranging a sampling node sequence, a first-in first-out queue is used as a cache strategy, the training stage is divided into four sub-stages, the four sub-stages of the training stage are parallel by using a pipeline technology, and the four sub-stages are connected by using a shared queue; and super-large graph neural network single-machine training of hybrid dynamic caching is realized. Compared with the prior art, the method has the advantages that the end-to-end training time is short, the utilization rate of GPU equipment is high, the cache hit rate can be effectively increased under different sampling algorithm loads, the mini-batch preparation time is shortened, more efficient GNN training is achieved, the method is simple and convenient, the use effect is good, and the method has good application prospects and commercial development value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of graph neural networks, and in particular to a single-machine training method for graph neural networks based on hybrid dynamic caching. Background Art

[0002] In recent years, with the continuous development of deep learning technology, graph learning has increasingly become a hot topic of concern. Graph Neural Networks (GNNs) are a new type of network structure that has emerged in the fields of machine learning and graph theory in recent years. Different from traditional neural networks, GNNs are particularly suitable for processing graph-structured data, such as social networks, molecular structures, traffic networks, etc. With the increase in the complexity and relevance of data, GNNs have gradually become a research hotspot due to their powerful representation learning and reasoning capabilities.

[0003] Existing graph neural network training systems are divided into two types: whole-graph training and mini-batch training. Whole-graph training uses the full graph structure data and graph node feature data, and it is necessary to keep the full data in memory. For large-scale graph data, it is often impossible to maintain the full graph structure data and graph node feature data in the video memory. To address the drawbacks of whole-graph training, the mini-batch-based training method uses sampling techniques to divide the whole graph into multiple mini-batches for training. The video memory occupancy of a mini-batch is often one-thousandth of the size of the whole graph and can be entirely placed in the GPU video memory for calculation. The training of a mini-batch-based graph neural network training system is divided into two steps: the preparation of mini-batches and the training of the model. The mini-batch preparation stage is further divided into the neighbor node sampling step and the feature transfer step. The neighbor node sampling step refers to performing a neighbor sampling algorithm on a given batch of training nodes to obtain a subgraph, and the feature transfer step refers to transferring the sampled subgraph from the host memory to the GPU video memory. Existing research has found that in the mini-batch training method, the mini-batch preparation stage has a greater overhead than the model training stage. To reduce the overhead of mini-batch preparation, some systems use the method of static feature caching to accelerate the neighbor node sampling and feature transfer processes. These static feature caching methods determine the data items cached in the GPU video memory by counting the access frequencies of node features and adjacency matrices during the pre-execution process. The researchers deeply counted the hit rate of node feature caching during mini-batch training under the loads of different sampling algorithms. After performing the pre-execution process using a certain sampling algorithm, when the obtained node feature cache accounts for 10% of the total node feature data, only a 68% hit rate can be achieved during the subsequent mini-batch preparation process, and the end-to-end benefit of the mini-batch preparation time is only about 7%. The static cache plan obtained in the pre-execution stage cannot achieve the best performance under the training loads of different sampling algorithms, which is due to the limitations of the static cache strategy.

[0004] In summary, for the whole-graph training and mini-batch training of the existing technology, although the video memory utilization rate can be improved, there are problems such as a relatively large overhead in the preparation stage, poor performance of the static cache strategy under diverse sampling algorithm loads, and low static cache strategy hit rates. Therefore, a new dynamic cache strategy is needed to optimize the graph neural network training system under different sampling method loads. Summary of the Invention

[0005] The object of the present invention is to provide a single-machine training method for graph neural networks based on hybrid dynamic caching in view of the deficiencies of the prior art. By adopting a hybrid dynamic node feature caching method and utilizing the characteristics of the graph data structure, hybrid dynamic caching is achieved, the hit rate of node feature caching in the mini-batch generation process is improved, and the mini-batch generation and the model training stage are pipelined and parallelized, reducing the end-to-end training time and improving the utilization rate of GPU devices. It effectively solves the problems of too high mini-batch preparation cost and low hit rate of static caching strategies during the training of graph neural networks. This method maintains dynamic node feature caching and static adjacency matrix caching in video memory, accelerates data transmission through direct memory access (DMA) technology, and the adopted hybrid dynamic caching strategy can effectively improve the cache hit rate under different sampling algorithm loads, reduce the mini-batch preparation time, and achieve more efficient GNN training. The method is simple, has good use effects, and has good application prospects and commercial development value.

[0006] The object of the present invention is achieved as follows: A single-machine training method for graph neural networks based on hybrid dynamic caching, characterized in that the method uses the access frequency information of graph nodes and adjacency matrix entries in the pre-training stage to obtain the ratio of dynamic node feature caching and static adjacency matrix entries caching, uses the set similarity reordering algorithm to reorder the sampling node order in the training stage, uses the first-in-first-out queue as the caching strategy, divides the training stage into four sub-stages, parallelizes the four sub-stages using pipeline technology, and uses a shared queue to connect the four sub-stages to achieve single-machine training of ultra-large graph neural networks with hybrid dynamic caching.

[0007] In the pre-training stage, by pre-executing to count the access frequencies of node features and adjacency matrix entries in the graph data, the cache allocation algorithm is used to obtain the ratio of dynamic node feature caching and static adjacency matrix caching, which is used in the training stage.

[0008] In the training stage, a Mini-batch generation component and a model training component are adopted. The Mini-batch generation component uses dynamic node feature caching and static adjacency matrix entries caching to accelerate the mini-batch generation speed; the model training component uses the GCN or GraphSAGE model to perform mini-batch training on the single-machine graph neural network, and traverses the entire graph data once in each round of epoch training until the model converges.

[0009] In the training stage, the Mini-batch generation component and the model training component are further divided into four sub-parts: reordering, sampling, transmission, and model training. These four sub-parts are processed in pipeline parallelization, and the mini-batch preparation stage and the model training stage are overlapped in the pipeline to better utilize the computing resources of the GPU.

[0010] The dynamic node feature cache uses the set similarity reordering algorithm and the first-in-first-out queue to obtain the topological information and node feature vectors of the sampled subgraph, and migrates the graph topological information and node feature vectors from the host memory to the video memory.

[0011] The mini-batch generation component includes: a node neighbor set similarity reordering algorithm that reorders the sampling order of randomly generated training nodes so that the set similarity between the neighbor sets of each adjacent sampled node is the highest, which can result in a higher hit rate of the first-in-first-out queue.

[0012] The mini-batch generation component also includes: a first-in-first-out queue on the video memory that maintains a first-in-first-out queue in the video memory and maintains a first-in-first-out queue of node feature vectors in the video memory. The static adjacency matrix cache on the video memory maintains a static adjacency matrix cache in the video memory. The ratio of the node feature cache size to the adjacency matrix item cache size is obtained from the pre-training stage.

[0013] The mini-batch generation component also includes: using direct memory access (DMA) to transfer node features and adjacency matrix data in the host memory to the video memory, reducing the overhead of data transfer between the host memory and the video memory.

[0014] Compared with the prior art, the present invention reduces the end-to-end training time, greatly improves the utilization rate of the GPU device, parallelizes the mini-batch generation and model training stages in the pipeline, improves the hit rate of node feature caching during mini-batch generation, maintains a dynamic node feature cache and a static adjacency matrix cache in the video memory, accelerates data transfer through direct memory access (DMA) technology, and the adopted hybrid dynamic cache strategy can effectively improve the cache hit rate under different sampling algorithm loads, reduce the mini-batch preparation time, achieve more efficient GNN training, with a simple method, good use effect, and having good application prospects and commercial development value. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic diagram of the framework of the graph neural network training system of the present invention; Figure 2 It is a schematic diagram of the set similarity reordering algorithm of Embodiment 1; Figure 3 Schematic diagram of the four-stage pipeline execution of Example 1. Specific implementation mode

[0016] The present invention utilizes the inherent characteristics of graph data to implement a single-machine training of a super-large graph neural network using a hybrid dynamic cache, aiming to solve the problem of low hit rate of static node feature caches under diverse sampling method loads. The present invention will be further described and explained in detail below in conjunction with embodiments and the accompanying drawings.

[0017] Example 1 Refer to Figure 1 , a single-machine training system of a super-large graph neural network using a hybrid dynamic cache adopting the architecture of the present invention. This system consists of a cache size allocation part before training and a part during training. The cache size allocation stage before training is responsible for determining the ratio of the node feature cache size and the adjacency matrix cache size under a given cache size according to the node access frequency of the training graph data and the access frequency of the adjacency matrix items. The determination of the node feature cache and the adjacency matrix size is carried out by using the node feature access frequency and the access frequency of the adjacency matrix items as the weights of the knapsack problem. Specifically, this problem is transformed into a knapsack problem and solved by a greedy algorithm. During the training stage, the randomly generated training node mini-batch input is re-ordered. The reason for re-ordering is to improve the hit rate of the sampled neighbor nodes in the first-in-first-out cache in the video memory.

[0018] Refer to Figure 2 a, the original graph data used in training, the first-in-first-out queue maintained in the video memory. After the random seed node re-ordering, the improvement in the hit rate of the sampled nodes in the first-in-first-out queue can reach a higher hit rate compared to the static cache, and still bring higher benefits than the static cache considering the dynamic cache update overhead.

[0019] Refer to Figure 2 b, the execution order of the first-in-first-out cache adopting the random seed node sampling order is: 17, 15, 11, 7, 9, 1. One-hop neighbor nodes are selected in each round of sampling and added to the first-in-first-out cache queue. During the sampling process of this mini-batch sequence, there are 8 cache hits in total, and the hit rate is 25%.

[0020] Refer to Figure 2c. After reordering by set similarity, the sampling order is: 17, 9, 1, 9, 7, 11, 15. During the sampling process of the mini-batch sequence after reordering, there are 12 cache hits in total, and the hit rate is 38%, which is 13% higher than that in the case of random sampling order. The specific reordering method determines the sampling order in the mini-batch by calculating the similarity of the k-hop neighbor sets between randomly input seed nodes. The calculation method of the similarity of the k-hop neighbor sets of nodes is as follows: 。

[0021] Among them, represents the k-hop neighbor set of node u.

[0022] In large-scale graph data, the k-hop neighbor sets of nodes are often extremely large. To calculate the similarity of the k-hop neighbor sets between pairwise nodes more quickly, the present invention uses the locality-sensitive hashing algorithm to quickly calculate the similarity of the k-hop neighbor sets of two nodes. In this embodiment, a k-hop neighbor set table is first generated for the randomly generated seed node sequence, and the locality-sensitive hashing algorithm implemented on the GPU is used to calculate the similarity values of the k-hop neighbor sets between different nodes pairwise, forming a similarity query table. When the reordering algorithm is executed, the algorithm randomly selects a training node as the initial node, then selects the node with the highest similarity of the k-hop neighbor set to this node and adds it to the reordered sequence. Then, using this node as the query node, it queries the node with the highest similarity of the k-hop neighbor set to it in the similarity query table and adds it to the reordered sequence. If there is no similar node to it in the table or all similar nodes to it have been added to the reordered sequence, a node is randomly selected from the set of nodes that have not been added to the reordered sequence, added to the reordered sequence, and used as the query node for the next round of loop until all the training nodes in the mini-batch are added to the reordered sequence. The time complexity of the k-hop neighbor set similarity query is O(1). In the algorithm, all N randomly selected seed nodes need to be queried, so the time complexity of this algorithm is O(N).

[0023] In the stage of generating mini-batch training data, sampling operations are sequentially performed according to the reordered seed node order. The sampled node IDs and the feature vectors of these nodes are migrated from the host memory to the first-in-first-out queue maintained in the video memory, and the first-in-first-out queue maintained in the video memory is updated. At the same time, when performing node neighbor sampling, if the node ID is less than the maximum ID of the static adjacency matrix items cached in the video memory by the system, the static adjacency matrix items cached in the video memory can be directly used to accelerate the sampling time.

[0024] During the model training phase, the system pipelines the model training phase and the mini-batch preparation phase by parallelizing the process of preparing small batches of data during training and the model training process through a pipeline mechanism. The system pipelines the seed node reordering, neighbor sampling, node feature transmission, and model training phases. By pipelining these phases, the longest time-consuming phase in each batch becomes the total time-consuming one.

[0025] See Figure 3 , in the model training phase of this embodiment, the pipeline is divided into four phases to achieve the overlap of different computing nodes and I / O phases, and each phase is processed by a CPU thread. The threads execute consecutive batches of data in different phases in parallel, using a lock-free shared queue in the producer-consumer pattern.

[0026] See Figure 3 a, the four pipeline phases in this system are: 1) the seed node reordering phase; 2) the subgraph sampling phase; 3) the node feature transmission phase; 4) the model training phase. Each phase in the four pipeline paths depends on the execution of the previous phase. The seed reordering phase starts from a random sampling sequence of seed nodes and reorders the sampling sequence using the set similarity sorting algorithm.

[0027] See Figure 3 b, the reordered seed nodes are put into the shared queue for the next pipeline phase. The sampling thread will perform neighbor sampling according to the reordered seed sequence and add the sampled node feature vectors to the shared queue. The transmission thread will obtain the node features from the video memory and the host memory, transfer them to the video memory, and update the first-in-first-out cache in the video memory to form a small batch of data for training. In the system, direct memory access to the video memory and UVA access to the host memory data are used to accelerate the feature data transmission speed. Finally, the model training phase obtains the node features obtained from the node transmission phase and the graph structure data obtained from the sampling phase, and performs GNN training calculations on this small batch.

[0028] The above is only a further explanation of the present invention and is not intended to limit the present invention. Equivalent implementations without departing from the spirit and scope of the present invention's concept should be included within the scope of the claims of the present invention.

Claims

1. A graph neural network stand-alone training method based on hybrid dynamic cache, characterized in that: The method uses the access frequency information of graph nodes and adjacency matrix items in the pre-training stage to obtain the ratio of dynamic node feature cache and static adjacency matrix item cache. In the training stage, the set similarity reordering algorithm is used to reorder the order of sampled nodes, a first-in-first-out queue is used as a cache strategy, and the training stage is divided into four sub-stages. The four sub-stages of the training stage are parallelized using pipeline technology, and the four sub-stages are connected using a shared queue to realize the single-machine training of a hybrid dynamic cached ultra-large graph neural network. The pre-training stage obtains the ratio of the node feature cache size and the adjacency matrix cache size under the premise of a given video memory size by pre-statistically analyzing the access frequency of each node feature and adjacency matrix item in the graph data; in the training stage, random seed nodes are reordered using the set similarity reordering algorithm to obtain a new seed node sequence, and the sorted seed node sequence sequence is used to sample neighbor nodes, and the sampled sub-graph is transferred from the host memory to the video memory, and finally the model is calculated; the four sub-stages are reordering, sampling, transmission and model training, and the four divided sub-stages are parallelized using pipeline technology, hiding the overlapping calculations and I / O between the sub-stages.

2. The graph neural network stand-alone training method based on hybrid dynamic cache according to claim 1 is characterized in that: The pre-training stage obtains the node feature access frequency and the adjacency matrix item access frequency in the graph data through pre-training pre-execution. According to the node feature access frequency, the adjacency matrix item access frequency, the node feature size and the adjacency matrix size, a greedy algorithm is used to resolve the ratio of the node feature cache size to the adjacency matrix cache size into a 0 / 1 knapsack problem.

3. The graph neural network stand-alone training method based on hybrid dynamic cache according to claim 1 is characterized in that: The set similarity reordering algorithm in the training phase uses the k-hop neighbor set of each seed node to establish a local sensitive hash table and calculate the set similarity between the k-hop neighbor sets of each node.

4. The graph neural network stand-alone training method based on hybrid dynamic cache according to claim 1 is characterized in that: The training phase simultaneously maintains a first-in-first-out queue for caching node features and a static graph topology data cache for caching graph topology data in the video memory, wherein the node feature cache updates cache items in the cache during the sampling and transmission phases.

5. The graph neural network stand-alone training method based on hybrid dynamic cache according to claim 1 is characterized in that: The training phase uses pipeline technology to parallelize the four sub-phases of training. Each phase is executed in parallel using one thread. The input of each phase depends on the output of another phase. The various phases communicate with each other using a shared queue.

Citation Information

Cited By

  • AVX-based large-scale GNN distributed training optimization method

    CN120762760A