Graph neural network load balancing method and device based on greedy algorithm
By generating the sliding window width and window threshold through a greedy algorithm and sorting and sliding operations based on node computing load, the problem of computational load imbalance in graph neural networks is solved, hardware-friendly and efficient load balancing distribution is achieved, and the computational efficiency of graph neural networks is improved.
Patent Information
- Application Number
- CN202510803510.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-26
AI Technical Summary
The computational load imbalance problem in graph neural networks leads to the inability to fully utilize hardware computing resources, affecting the overall reasoning efficiency. Existing methods have high computational complexity and poor hardware adaptability, making it difficult to meet real-time reasoning requirements.
A greedy algorithm is used to generate the sliding window width and window threshold. Based on the node computing load sorting and sliding operation, computing tasks are dynamically allocated to hardware computing units to achieve load balancing.
It reduces the complexity of computing load distribution, improves hardware computing efficiency, optimizes resource utilization, adapts to dynamic graph structure characteristics, and improves the computing efficiency of graph neural networks.
Smart Images

Figure CN120706466A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of graph neural network technology, and in particular to a graph neural network load balancing method and device based on a greedy algorithm. Background Art
[0002] During the inference calculation process based on graph neural networks, the aggregation operation requires distributed calculation of graph node features through the adjacency matrix. Due to the topological characteristics of the graph structure, the non-zero elements in the adjacency matrix are randomly and sparsely distributed, resulting in significant non-uniform load in the calculation rows / columns allocated by parallel computing units. This load imbalance phenomenon makes it impossible to fully utilize hardware computing resources. Some computing units will experience execution delays due to excessive assigned tasks, which directly affects the overall inference efficiency. Therefore, how to achieve dynamic load balancing while maintaining the advantages of parallel computing architecture has become a key requirement for improving the inference performance of graph neural networks.
[0003] To address these issues, relevant technical solutions optimize load distribution by introducing multi-level task scheduling algorithms or customized hardware resource allocation strategies. These approaches are typically based on complex optimal solution search models and achieve load balancing through computational task remapping or dynamic hardware resource partitioning. Some solutions also employ pre-compilation optimization techniques to generate static load distribution plans before task execution to reduce runtime overhead.
[0004] However, these methods suffer from high computational complexity and poor hardware adaptability. Algorithms based on optimal solution search consume significant time and resources for task allocation decisions, making it difficult to meet the low-latency requirements of real-time reasoning. Static precompilation schemes, on the other hand, cannot adapt to the dynamically changing characteristics of graph structures. Furthermore, existing load distribution mechanisms lack parallel generation capabilities and are unable to simultaneously generate balanced task flows for multiple computing units, resulting in a coordination gap between the allocation process and the hardware computing pipeline. These shortcomings severely restrict the application effectiveness of graph neural networks in edge computing devices and real-time reasoning scenarios. Summary of the Invention
[0005] This application provides a graph neural network load balancing method and device based on a greedy algorithm to solve the problem of unbalanced computing load in graph neural networks.
[0006] In a first aspect, the present application provides a graph neural network load balancing method based on a greedy algorithm, the method comprising:
[0007] Generate sliding window width and window threshold based on the number of hardware computing units, the number of nodes, and the total computing load;
[0008] The sliding window width is the number of matrix rows / columns that need to be calculated and allocated to each hardware computing unit; the window threshold is the threshold of the equalization effect that the sliding window is expected to achieve;
[0009] Sort the nodes in the graph neural network according to the size of the computational load;
[0010] Performing a sliding operation along the sorted node sequence based on the sliding window width;
[0011] When the window height reaches the window threshold, the sliding of the current sliding window is stopped;
[0012] Allocate the nodes in the current sliding window to the hardware computing unit to perform aggregation calculation.
[0013] By generating the sliding window width and window threshold based on the number of hardware computing units, the number of nodes and the total computing load, sorting the nodes according to the computing load and performing the sliding operation based on the sliding window width, and stopping the sliding when the window height reaches the window threshold and allocating the nodes in the current sliding window to the hardware computing unit, the complexity of allocating the computing load can be reduced, and it is easy to expand and parallelize, and easy to execute with the computational pipeline of the hardware computing unit, thereby improving the execution efficiency of the hardware calculation, and having good hardware friendliness, it can solve the problem of unbalanced computing load in graph neural networks, thereby improving the computing efficiency of graph neural networks.
[0014] Optionally, the step of generating the sliding window width includes:
[0015] The number of matrix rows / columns allocated to each hardware computing unit is determined based on the ratio of the number of nodes to the number of hardware computing units to serve as the sliding window width.
[0016] By determining the sliding window width based on the ratio of the number of nodes to the number of hardware computing units, the number of matrix rows / columns allocated to each hardware computing unit is made more reasonable, which helps to optimize the distribution of computing tasks and thus improve load balancing and computing resource utilization.
[0017] Optionally, the step of generating the window threshold includes:
[0018] Determine the balancing target value based on the ratio of the total computing load to the number of nodes;
[0019] A window threshold is generated based on the equalization target value, the number of nodes in the window, and a preset floating value.
[0020] By determining the balancing target value based on the ratio of the total computing load to the number of nodes, and generating the window threshold based on the number of nodes in the window and the preset floating value, the load distribution standard can be dynamically adjusted to make the task division more in line with the actual computing needs, thereby optimizing the load balancing effect and improving computing efficiency.
[0021] Optionally, the step of performing a sliding operation along the sorted node sequence based on the sliding window width includes:
[0022] Sort the nodes from largest to smallest according to their computing load to generate a sorted node sequence;
[0023] The sliding window width is used as the range of the number of nodes covered by the sliding window each time it slides on the node sequence, and the sliding window starts from the starting end of the sorted node sequence.
[0024] By sorting the nodes from large to small according to the computing load to generate a node sequence, and sliding from the beginning of the sequence with the sliding window width as the sliding coverage range, high-load nodes can be allocated first, making the computing resource allocation more reasonable, thereby improving the load balancing effect and improving computing efficiency.
[0025] Optionally, when the window height reaches the window threshold, the step of stopping the sliding of the current sliding window includes:
[0026] The sum of the computational loads of all nodes within the sliding window is dynamically and real-timely obtained during the sliding window process;
[0027] When the sum of the computational loads of all nodes in the sliding window reaches the window threshold, the sliding of the current sliding window is stopped.
[0028] By dynamically obtaining the sum of the computing loads of the nodes in the sliding window in real time during the sliding process and stopping sliding when the window threshold is reached, the task allocation range can be dynamically adjusted according to the actual load situation, making the computing resource allocation more reasonable, thereby optimizing the load balancing effect and improving computing efficiency.
[0029] Optionally, after allocating the nodes in the current sliding window to the hardware computing unit, the method further includes:
[0030] The sliding window starts sliding again from the starting position of the sorted node sequence;
[0031] When the window height reaches the window threshold, the sliding of the current sliding window is stopped to obtain a second set of nodes within the sliding window;
[0032] Allocating the second group of sliding window nodes to the hardware computing unit to perform aggregation calculations.
[0033] By re-starting sliding from the starting position of the node sequence after the current sliding window allocation is completed, and obtaining a second set of nodes in the sliding window for allocation when the window threshold is reached, multiple rounds of task allocation can be achieved, making the computing load distribution of the hardware computing unit more even, thereby improving the overall load balancing effect and increasing the utilization of computing resources.
[0034] Optionally, the method further includes:
[0035] Sort the nodes in the graph neural network according to the size of the computational load in multiple parallel spaces at the same time;
[0036] In each parallel space, a sliding operation is performed along the sorted node sequence based on the sliding window width;
[0037] Among them, the sliding operations of multiple parallel spaces can simultaneously generate multiple balanced sliding window nodes and distribute them to corresponding hardware computing units respectively.
[0038] By executing node sorting and sliding operations simultaneously in multiple parallel spaces, multiple balanced sliding window node groups can be generated in parallel and allocated to corresponding hardware computing units, thereby improving task allocation efficiency, shortening load balancing processing time, and optimizing the load distribution between multiple computing units, thereby improving overall computing resource utilization.
[0039] Optionally, in each parallel space, the step of performing a sliding operation along the sorted node sequence based on the sliding window width includes:
[0040] Synchronously execute multiple sliding window traversal operations on independent node sequences;
[0041] Each sliding window corresponds to the allocation queue of a different hardware computing unit.
[0042] By synchronously executing multiple sliding window traversal operations on independent node sequences, each sliding window corresponds to a different hardware computing unit allocation queue, which can realize the parallel task allocation process and reduce the allocation waiting time between different computing units, thereby improving the load balancing processing efficiency and optimizing the computing resource scheduling effect.
[0043] Optionally, the method further includes:
[0044] The sliding window allocation phase and the aggregation calculation phase of the hardware computing unit are executed in a pipeline manner; the pipeline execution includes: when the hardware computing unit executes the currently allocated calculation load, the sliding window allocation operation of the next calculation cycle is synchronously executed.
[0045] By executing the sliding window allocation stage and the aggregation calculation stage of the hardware computing unit in a pipeline manner, the computing load distribution and the computing processing process are parallelized, which can reduce the idle time of the computing unit and improve the continuity of task processing, thereby improving the overall system throughput and optimizing the efficiency of computing resource utilization.
[0046] A second aspect of the present application provides a graph neural network load balancing device based on a greedy algorithm, which is applicable to the graph neural network load balancing method based on a greedy algorithm described in the first aspect, and the device includes:
[0047] A parameter generation module is configured to generate a sliding window width and a window threshold based on the number of hardware computing units, the number of nodes, and the total computing load; wherein the sliding window width is the number of matrix rows / columns required for calculation allocated to each hardware computing unit; and the window threshold is the threshold of the equalization effect expected to be achieved by the sliding window;
[0048] The node sorting module is used to sort the nodes in the graph neural network according to the size of the computational load;
[0049] The dynamic sliding window module is used to perform sliding window traversal and node allocation based on window threshold triggering.
[0050] The above-mentioned device determines the sliding window width and window threshold through the parameter generation module, the node sorting module sorts the computing load, and the dynamic sliding window module performs sliding window traversal and node allocation based on the window threshold trigger, which can realize the dynamic allocation of computing tasks, reduce the complexity of allocating computing load, and is easy to expand and parallelize. It is also easy to execute with the computational pipeline of the hardware computing unit, thereby improving the execution efficiency of hardware computing and having good hardware friendliness. It can solve the problem of unbalanced computing load in graph neural networks, thereby improving the computing efficiency of graph neural networks.
[0051] It can be seen from the above technical solution that the present application provides a graph neural network load balancing method and device based on a greedy algorithm, wherein the method generates a sliding window width and a window threshold according to the number of hardware computing units, the number of nodes and the total computing load; wherein the sliding window width is the number of matrix rows / columns that need to be calculated allocated to each hardware computing unit; the window threshold is the threshold of the balancing effect that the sliding window expects to achieve; the nodes in the graph neural network are sorted according to the size of the computing load; a sliding operation is performed along the sorted node sequence based on the sliding window width; when the window height reaches the window threshold, the sliding of the current sliding window is stopped; the nodes in the current sliding window are allocated to the hardware computing unit to perform aggregation calculations to solve the problem of unbalanced computing load in the graph neural network. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0053] Figure 1 A schematic diagram of a flow chart of a graph neural network load balancing method based on a greedy algorithm provided in an embodiment of the present application;
[0054] Figure 2 A schematic diagram of the sliding window allocation phase process in the graph neural network load balancing method based on the greedy algorithm provided in an embodiment of the present application;
[0055] Figure 3 A schematic diagram of the parallel multiple spatial sliding window allocation phase process in the graph neural network load balancing method based on the greedy algorithm provided in an embodiment of the present application;
[0056] Figure 4 Schematic diagram of the pipeline execution process in the graph neural network load balancing method based on the greedy algorithm provided in the embodiment of the present application. DETAILED DESCRIPTION
[0057] The following embodiments are described in detail, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following embodiments are not intended to represent all possible implementations consistent with the present application. They are merely examples of systems and methods consistent with certain aspects of the present application.
[0058] To facilitate understanding of the technical solutions of the embodiments of the present application, before describing the specific implementation methods of the embodiments of the present application, some technical terms in the technical field to which the embodiments of the present application belong are briefly explained:
[0059] A greedy algorithm is an algorithmic strategy that takes the best (locally optimal) decision in each step, hoping to reach a global optimal solution. Its core characteristic is that the accumulation of local optimal choices eventually approaches the global optimal solution.
[0060] Graph neural network is a type of deep learning model specifically designed to process graph-structured data. It can represent, learn, and predict nodes, edges, and global graph information in the graph.
[0061] During graph neural network inference, we optimize the load imbalance caused by the randomness and sparsity of graph adjacency relationships. The aggregation part of graph neural network inference requires aggregating the features of all nodes to the nodes connected to them. Due to the inherent randomness and sparsity of graph adjacency relationships, the distribution of non-zero elements in the adjacency matrix of graph nodes exhibits random and sparse distribution characteristics. During the multiplication of the adjacency matrix and the feature matrix, rows / columns are evenly and parallelly assigned to specific parallel hardware computing units to perform calculations. The number of non-zero elements in these rows / columns determines the total computational load. Therefore, the randomness and sparsity of the adjacency matrix will cause an imbalance in the computational load assigned to the parallel hardware computing units, thereby affecting the efficiency of network inference.
[0062] In a related embodiment, load distribution is optimized by introducing a multi-level task scheduling algorithm or a customized hardware resource allocation strategy. Such methods are usually based on a complex optimal solution search model, and load balancing is achieved by computing task remapping or dynamic partitioning of hardware resources. Some solutions also use pre-compilation optimization technology to generate a static load distribution scheme before task execution to reduce runtime overhead. However, the above methods have the defects of high computational complexity and poor hardware adaptability. The algorithm based on optimal solution search consumes a lot of time resources for task allocation decisions, which is difficult to meet the low latency requirements of real-time reasoning; the static pre-compilation scheme cannot adapt to the dynamically changing graph structure characteristics. In addition, the existing load distribution mechanism lacks parallel generation capabilities and cannot generate balanced task flows for multiple computing units at the same time, resulting in a coordination gap between the allocation process and the hardware computing pipeline. These defects seriously restrict the application efficiency of graph neural networks in edge computing devices and real-time reasoning scenarios.
[0063] To solve the above problems, see Figure 1 Some embodiments of the present application provide a graph neural network load balancing method based on a greedy algorithm, the method comprising:
[0064] S100: Generate a sliding window width and a window threshold according to the number of hardware computing units, the number of nodes, and the total computing load.
[0065] The sliding window width is the number of matrix rows / columns required for calculation allocated to each hardware computing unit; and the window threshold is the threshold of the equalization effect expected to be achieved by the sliding window.
[0066] In some embodiments, the step of generating the sliding window width includes: determining the number of matrix rows / columns allocated to each hardware computing unit based on the ratio of the number of nodes to the number of hardware computing units, as the sliding window width.
[0067] By determining the sliding window width based on the ratio of the number of nodes to the number of hardware computing units, the number of matrix rows / columns allocated to each hardware computing unit is made more reasonable, which helps to optimize the distribution of computing tasks and thus improve load balancing and computing resource utilization.
[0068] In some embodiments, the step of generating the window threshold includes: determining a balancing target value based on a ratio of the calculated total load to the number of nodes; and generating a window threshold based on the balancing target value, the number of nodes in the window, and a preset floating value.
[0069] It should be understood that the preset floating value is used to enable the window threshold to achieve the desired equalization effect. The range of the window threshold is less than or equal to the product of the equalization target value and the number of nodes in the window.
[0070] By determining the balancing target value based on the ratio of the total computing load to the number of nodes, and generating the window threshold based on the number of nodes in the window and the preset floating value, the load distribution standard can be dynamically adjusted to make the task division more in line with the actual computing needs, thereby optimizing the load balancing effect and improving computing efficiency.
[0071] S200: Sort the nodes in the graph neural network according to the size of the computational load.
[0072] S300: Performing a sliding operation along the sorted node sequence based on the sliding window width.
[0073] In some embodiments, the step of performing a sliding operation along the sorted node sequence based on the sliding window width includes:
[0074] Sort the nodes from largest to smallest according to their computing load to generate a sorted node sequence;
[0075] The sliding window width is used as the range of the number of nodes covered by the sliding window each time it slides on the node sequence, and the sliding window starts from the starting end of the sorted node sequence.
[0076] By sorting the nodes from large to small according to the computing load to generate a node sequence, and sliding from the beginning of the sequence with the sliding window width as the sliding coverage range, high-load nodes can be allocated first, making the computing resource allocation more reasonable, thereby improving the load balancing effect and improving computing efficiency.
[0077] S400: When the window height reaches the window threshold, the sliding of the current sliding window is stopped.
[0078] In some embodiments, when the window height reaches the window threshold, the step of stopping the sliding of the current sliding window includes:
[0079] The sum of the computational loads of all nodes within the sliding window is dynamically and real-timely obtained during the sliding window process;
[0080] When the sum of the computational loads of all nodes in the sliding window reaches the window threshold, the sliding of the current sliding window is stopped.
[0081] By dynamically obtaining the sum of the computing loads of the nodes in the sliding window in real time during the sliding process and stopping sliding when the window threshold is reached, the task allocation range can be dynamically adjusted according to the actual load situation, making the computing resource allocation more reasonable, thereby optimizing the load balancing effect and improving computing efficiency.
[0082] S500: Allocate the nodes in the current sliding window to the hardware computing unit to perform aggregation calculation.
[0083] By generating the sliding window width and window threshold based on the number of hardware computing units, the number of nodes and the total computing load, sorting the nodes according to the computing load and performing the sliding operation based on the sliding window width, and stopping the sliding when the window height reaches the window threshold and allocating the nodes in the current sliding window to the hardware computing unit, the complexity of allocating the computing load can be reduced, and it is easy to expand and parallelize, and easy to execute with the computational pipeline of the hardware computing unit, thereby improving the execution efficiency of the hardware calculation, and having good hardware friendliness, it can solve the problem of unbalanced computing load in graph neural networks, thereby improving the computing efficiency of graph neural networks.
[0084] In some embodiments, after allocating the nodes in the current sliding window to the hardware computing unit, the method further includes:
[0085] The sliding window starts sliding again from the starting position of the sorted node sequence;
[0086] When the window height reaches the window threshold, the sliding of the current sliding window is stopped to obtain a second set of nodes within the sliding window;
[0087] Allocating the second group of sliding window nodes to the hardware computing unit to perform aggregation calculations.
[0088] It should be understood that the above method should be repeated until the sliding window traverses the sorted node sequence. Since the sliding window stops each time it reaches the window threshold, the computing load distributed to the hardware computing unit is balanced.
[0089] By re-starting sliding from the starting position of the node sequence after the current sliding window allocation is completed, and obtaining a second set of nodes in the sliding window for allocation when the window threshold is reached, multiple rounds of task allocation can be achieved, making the computing load distribution of the hardware computing unit more even, thereby improving the overall load balancing effect and increasing the utilization of computing resources.
[0090] See also Figure 2 As shown in the figure, NNZ represents the total number of computational loads; thr represents the balancing target value; Row ID is the node row identifier; Sort represents sorting; and Slide represents sliding. Nodes are sorted from largest to smallest by computational load, and the starting point of the sorted node sequence is the node with the largest computational load. Therefore, the initially generated window height must be higher than the window threshold. As the window slides, the window height gradually decreases until it reaches the window threshold. The sliding of the current window is stopped, and the computational load within the window is extracted and allocated to the hardware computing units. This is superior to not needing to solve the optimal allocation during the sliding process. Instead, the window stops at the window threshold to ensure that all window heights reach the window threshold, thus ensuring load balancing and significantly reducing the time overhead of solving the problem.
[0091] In some embodiments, the method further comprises:
[0092] Sort the nodes in the graph neural network according to the size of the computational load in multiple parallel spaces at the same time;
[0093] In each parallel space, a sliding operation is performed along the sorted node sequence based on the sliding window width;
[0094] Among them, the sliding operations of multiple parallel spaces can simultaneously generate multiple balanced sliding window nodes and distribute them to corresponding hardware computing units respectively.
[0095] By executing node sorting and sliding operations simultaneously in multiple parallel spaces, multiple balanced sliding window node groups can be generated in parallel and allocated to corresponding hardware computing units, thereby improving task allocation efficiency, shortening load balancing processing time, and optimizing the load distribution between multiple computing units, thereby improving overall computing resource utilization.
[0096] In some embodiments, in each parallel space, the step of performing a sliding operation along the sorted node sequence based on the sliding window width includes:
[0097] Synchronously execute multiple sliding window traversal operations on independent node sequences;
[0098] Each sliding window corresponds to the allocation queue of a different hardware computing unit.
[0099] See also Figure 3 As shown in the figure, NNZ represents the total number of computational loads; thr represents the balancing target value; Row ID is the node row identifier; Sort represents sorting; and Slide represents sliding. Users can perform sliding window load distribution in multiple spaces in parallel to ensure balanced computational loads are distributed to multiple parallel hardware computing units. Multiple sliding windows can simultaneously slide across corresponding spaces, generating balanced computational loads for distribution to multiple hardware computing units, improving execution efficiency.
[0100] By synchronously executing multiple sliding window traversal operations on independent node sequences, each sliding window corresponds to a different hardware computing unit allocation queue, which can realize the parallel task allocation process and reduce the allocation waiting time between different computing units, thereby improving the load balancing processing efficiency and optimizing the computing resource scheduling effect.
[0101] In some embodiments, the method further comprises:
[0102] The sliding window allocation phase and the aggregation calculation phase of the hardware computing unit are executed in a pipeline manner; the pipeline execution includes: when the hardware computing unit executes the currently allocated calculation load, the sliding window allocation operation of the next calculation cycle is synchronously executed.
[0103] See also Figure 4 As shown in the figure, Latency (optimized) represents the optimized latency; Latency (original) represents the original latency; Time represents time; Preprocessing Stage represents the preprocessing stage; Computing Stage represents the computing stage; and Tile represents the hardware unit. Among them, the load distribution stage of the sliding window and the computing stage of the hardware unit can be executed in a pipeline manner. After the sliding window distributes the load, the hardware computing unit obtains a more balanced computing load, thereby shortening the execution time of the original Tile0 to the optimized Tile0 execution time in the figure. Because the next balanced computing load can be generated while the hardware unit is calculating, the hardware unit can directly perform the calculation of the optimized Tile1 without waiting, so it has good software and hardware synergy and hardware friendliness.
[0104] By executing the sliding window allocation stage and the aggregation calculation stage of the hardware computing unit in a pipeline manner, the computing load distribution and the computing processing process are parallelized, which can reduce the idle time of the computing unit and improve the continuity of task processing, thereby improving the overall system throughput and optimizing the efficiency of computing resource utilization.
[0105] Some embodiments of the present application further provide a graph neural network load balancing device based on a greedy algorithm, which is applicable to the graph neural network load balancing method based on a greedy algorithm described in the above embodiments. The device includes:
[0106] A parameter generation module is configured to generate a sliding window width and a window threshold based on the number of hardware computing units, the number of nodes, and the total computing load; wherein the sliding window width is the number of matrix rows / columns required for calculation allocated to each hardware computing unit; and the window threshold is the threshold of the equalization effect expected to be achieved by the sliding window;
[0107] The node sorting module is used to sort the nodes in the graph neural network according to the size of the computational load;
[0108] The dynamic sliding window module is used to perform sliding window traversal and node allocation based on window threshold triggering.
[0109] The above-mentioned device determines the sliding window width and window threshold through the parameter generation module, the node sorting module sorts the computing load, and the dynamic sliding window module performs sliding window traversal and node allocation based on the window threshold trigger, which can realize the dynamic allocation of computing tasks, reduce the complexity of allocating computing load, and is easy to expand and parallelize. It is also easy to execute with the computational pipeline of the hardware computing unit, thereby improving the execution efficiency of hardware computing and having good hardware friendliness. It can solve the problem of unbalanced computing load in graph neural networks, thereby improving the computing efficiency of graph neural networks.
[0110] In some embodiments, the device further includes: a parallel allocation module for generating multiple parallel computing load distribution schemes; and a pipeline control module for coordinating the timing execution relationship between the sliding window allocation and the hardware computing unit.
[0111] The device generates multiple parallel computing load distribution schemes through a parallel allocation module, and combines with a pipeline control module to coordinate the timing execution relationship between sliding window allocation and hardware computing units, thereby realizing parallel allocation and pipeline processing of computing tasks, reducing the idle time of hardware computing units, thereby improving overall computing efficiency and optimizing resource utilization.
[0112] It can be seen from the above technical solution that an embodiment of the present application provides a graph neural network load balancing method and device based on a greedy algorithm, wherein the method generates a sliding window width and a window threshold according to the number of hardware computing units, the number of nodes and the total computing load; wherein the sliding window width is the number of matrix rows / columns that need to be calculated allocated to each hardware computing unit; the window threshold is the threshold of the balancing effect that the sliding window expects to achieve; the nodes in the graph neural network are sorted according to the size of the computing load; a sliding operation is performed along the sorted node sequence based on the sliding window width; when the window height reaches the window threshold, the sliding of the current sliding window is stopped; the nodes in the current sliding window are allocated to the hardware computing unit to perform aggregation calculations to solve the problem of unbalanced computing load in the graph neural network.
[0113] Similar parts between the embodiments provided in this application can be referenced to each other. The specific implementation methods provided above are only a few examples under the overall concept of this application and do not constitute a limitation on the scope of protection of this application. For those skilled in the art, any other implementation methods expanded based on the scheme of this application without expending creative work shall fall within the scope of protection of this application.
Claims
1. A graph neural network load balancing method based on a greedy algorithm, characterized in that: The method comprises: Generate sliding window width and window threshold based on the number of hardware computing units, the number of nodes, and the total computing load; The sliding window width is the number of matrix rows / columns that need to be calculated and allocated to each hardware computing unit; the window threshold is the threshold of the equalization effect that the sliding window is expected to achieve; Sort the nodes in the graph neural network according to the size of the computational load; Performing a sliding operation along the sorted node sequence based on the sliding window width; When the window height reaches the window threshold, the sliding of the current sliding window is stopped; Allocate the nodes in the current sliding window to the hardware computing unit to perform aggregation calculation.
2. The graph neural network load balancing method based on the greedy algorithm according to claim 1 is characterized in that: The step of generating the sliding window width includes: The number of matrix rows / columns allocated to each hardware computing unit is determined based on the ratio of the number of nodes to the number of hardware computing units to serve as the sliding window width.
3. The graph neural network load balancing method based on the greedy algorithm according to claim 1 is characterized in that: The step of generating the window threshold comprises: Determine the balancing target value based on the ratio of the total computing load to the number of nodes; A window threshold is generated based on the equalization target value, the number of nodes in the window, and a preset floating value.
4. The graph neural network load balancing method based on the greedy algorithm according to claim 3 is characterized in that: The step of performing a sliding operation along the sorted node sequence based on the sliding window width includes: Sort the nodes from largest to smallest according to their computing load to generate a sorted node sequence; The sliding window width is used as the range of the number of nodes covered by the sliding window each time it slides on the node sequence, and the sliding window starts from the starting end of the sorted node sequence.
5. The graph neural network load balancing method based on the greedy algorithm according to claim 4 is characterized in that: When the window height reaches the window threshold, the step of stopping the sliding of the current sliding window includes: The sum of the computational loads of all nodes within the sliding window is dynamically and real-timely obtained during the sliding window process; When the sum of the computational loads of all nodes in the sliding window reaches the window threshold, the sliding of the current sliding window is stopped.
6. The graph neural network load balancing method based on the greedy algorithm according to claim 4 is characterized in that: After allocating the nodes in the current sliding window to the hardware computing unit, the method further includes: The sliding window starts sliding again from the starting position of the sorted node sequence; When the window height reaches the window threshold, the sliding of the current sliding window is stopped to obtain a second set of nodes within the sliding window; Allocating the second group of sliding window nodes to the hardware computing unit to perform aggregation calculations.
7. The graph neural network load balancing method based on the greedy algorithm according to claim 4 is characterized in that: The method further comprises: Sort the nodes in the graph neural network according to the size of the computational load in multiple parallel spaces at the same time; In each parallel space, a sliding operation is performed along the sorted node sequence based on the sliding window width; Among them, the sliding operations of multiple parallel spaces can simultaneously generate multiple balanced sliding window nodes and distribute them to corresponding hardware computing units respectively.
8. The graph neural network load balancing method based on the greedy algorithm according to claim 7 is characterized in that: In each parallel space, the step of performing a sliding operation along the sorted node sequence based on the sliding window width includes: Synchronously execute multiple sliding window traversal operations on independent node sequences; Each sliding window corresponds to the allocation queue of a different hardware computing unit.
9. The graph neural network load balancing method based on the greedy algorithm according to claim 8 is characterized in that: The method further comprises: The sliding window allocation phase and the aggregation calculation phase of the hardware computing unit are executed in a pipeline manner; the pipeline execution includes: when the hardware computing unit executes the currently allocated calculation load, the sliding window allocation operation of the next calculation cycle is synchronously executed.
10. A graph neural network load balancing device based on a greedy algorithm, characterized in that: The method for load balancing a graph neural network based on a greedy algorithm according to any one of claims 1 to 9, wherein the device comprises: A parameter generation module is configured to generate a sliding window width and a window threshold based on the number of hardware computing units, the number of nodes, and the total computing load; wherein the sliding window width is the number of matrix rows / columns required for calculation allocated to each hardware computing unit; and the window threshold is the threshold of the equalization effect expected to be achieved by the sliding window; The node sorting module is used to sort the nodes in the graph neural network according to the size of the computational load; The dynamic sliding window module is used to perform sliding window traversal and node allocation based on window threshold triggering.