Graph division method and system for graph neural network reasoning in heterogeneous edge scene

By evaluating device capabilities in heterogeneous edge scenarios and using graph division methods with descending vertex arrangement and seed allocation strategies, the load imbalance caused by device heterogeneity is solved, and the inference efficiency and resource utilization of graph neural networks are improved.

CN120371541AActive Publication Date: 2025-07-25NAT UNIV OF DEFENSE TECH

Patent Information

Application Number
CN202510865703.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

In heterogeneous edge scenarios, the existing graph neural network inference tasks are load imbalanced and inferred inference efficiency due to device heterogeneity and defects in traditional graph division methods. It is difficult for existing evaluation methods to accurately evaluate the computing capabilities of the device and consider the message delivery characteristics of graph neural networks.

Method used

The proportion of graph neural network inference capabilities of edge devices is evaluated through benchmark tests, and the vertex descending order arrangement and seed allocation strategy of graph data are used, and the graph division is divided by combining deep expansion trees and streaming greedy allocation strategies to optimize load balancing and message delivery efficiency among devices.

Benefits of technology

It effectively improves the inference efficiency of graph neural networks in heterogeneous edge environments, realizes load balancing and resource utilization, and adapts to diversified data sets and edge device scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371541A_ABST
    Figure CN120371541A_ABST
Patent Text Reader

Abstract

The invention discloses a graph division method and system for graph neural network reasoning in a heterogeneous edge scene, and the method comprises the steps: evaluating the proportion of the graph neural network reasoning capability of edge equipment in a distributed environment through a benchmark test, carrying out the graph division of input graph data into partitions corresponding to # imgabs0 edge equipment, comprising the following steps: arranging vertexes in graph data in a descending order according to degrees of the vertexes, and screening first-order seeds and second-order seeds; randomly allocating partitions for the first-order seeds and the second-order neighbor blocks; constructing a depth extension tree for second-order neighbor blocks of the second-order seeds, and distributing partitions according to blocks according to the proportion of the inference capability of the graph neural network of edge equipment in a distributed environment; and adopting a streaming greedy distribution strategy to distribute partitions to the remaining undistributed vertexes. The invention aims to solve the problems of load imbalance and low reasoning efficiency caused by equipment heterogeneity and defects of a traditional graph division method in an edge graph neural network reasoning scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of edge deployment of graph neural networks, and particularly relates to a graph partitioning method and system for graph neural network inference in heterogeneous edge scenarios. Background Art

[0002] Edge-deployed graph neural networks (GNNs) have become the mainstream paradigm for processing graph inference tasks at the device side, showing significant efficacy in applications such as point cloud processing, autonomous driving, and privacy information processing. This benefits from the excellent graph-structured data representation ability of graph neural networks, the continuous improvement of edge hardware performance, and the rapid development of open-source graph neural network computing libraries (such as PyG and DGL, etc.). However, due to the limited computing resources and memory capacity of edge devices, many graph neural network inference tasks still cannot meet the requirements of real-time inference. Therefore, accelerating edge graph neural network inference through distributed technology has become a research hotspot.

[0003] As the current mainstream graph neural network processing libraries, PyG and DGL are widely used in distributed edge graph neural network inference scenarios. In these two libraries, METIS, as the only default graph partitioning algorithm, is responsible for partitioning the input graph into multiple subgraphs, and then PyG or DGL assigns each subgraph to different devices for graph neural network inference. However, both PyG and DGL default to using an equal partitioning strategy and do not fully consider the inherent characteristic of device heterogeneity in the edge environment - its heterogeneity stems from multiple factors such as hardware generation differences, application requirement differences, cost-oriented design limitations, and dynamic deployment conditions, resulting in unbalanced graph neural network inference loads on each device in the heterogeneous edge environment and seriously affecting the overall inference efficiency.

[0004] Currently, graph partitioning for graph neural network inference in heterogeneous edge environments faces two core challenges. First, it is difficult to evaluate the computing power of edge devices for graph neural network inference. This is because the hardware heterogeneity of edge devices (such as processors, memory, etc.) makes it difficult to determine standardized evaluation metrics, and most existing evaluation methods rely on a single metric (such as FLOPs), which cannot comprehensively evaluate the complex mathematical calculations of graph neural networks. At the same time, the randomness of graph neural network cross-device sampling and the dynamic nature of the communication network in a distributed environment lead to unpredictable delays, which also increase the difficulty of accurate evaluation. Second, it is still challenging to identify and utilize the characteristics of graph neural network inference to guide the design of graph partitioning algorithms. Existing graph partitioning algorithms do not consider the specific operations (such as neighbor sampling, message passing) and other characteristics of graph neural network inference during the execution of distributed graph neural networks, resulting in subgraphs partitioned by traditional graph partitioning algorithms not achieving the optimal effect of message passing during inference. Summary of the Invention

[0005] Technical problems to be solved by the present invention: Aiming at the above problems of the prior art, a graph partitioning method and system for graph neural network inference in heterogeneous edge scenarios are provided. The present invention aims to solve the problems of load imbalance and low inference efficiency caused by device heterogeneity and defects of traditional graph partitioning methods in the edge graph neural network inference scenario.

[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A graph partitioning method for graph neural network inference in heterogeneous edge scenarios, including the following steps: S101, evaluating the proportion of the graph neural network inference ability of edge devices in a distributed environment through benchmark testing; S102, according to the proportion of the graph neural network inference ability of edge devices in a distributed environment, partitioning the input graph data into partitions corresponding to edge devices, including: sorting the vertices in the graph data in descending order of vertex degree and screening first-order seeds and second-order seeds; randomly assigning partitions to the first-order seeds and their second-order neighbor blocks; constructing a depth expansion tree for the second-order neighbor blocks of the second-order seeds and allocating partitions by block according to the proportion of the graph neural network inference ability of edge devices in a distributed environment; using a streaming greedy allocation strategy to allocate partitions to the remaining unallocated vertices.

[0007] Optionally, step S101 includes: S201, constructing a benchmark test dataset, where the benchmark test dataset includes different topological structures and different types of graph datasets, and the different topological structures refer to some or all of the number of vertices, the number of edges, the average degree of vertices, and the feature dimension of vertices being the same; S202, for each graph dataset, using the edge cut graph partitioning method METIS to partition the graph dataset into non-overlapping partitions, where is the total number of edge devices in the distributed environment; S203, performing graph neural network inference on the benchmark test dataset through the distributed environment, and calculating the ability of any edge device to infer any graph dataset for the partition according to the following formula : , where, is the number of vertices of the non-overlapping partition of the dataset partitioned to the edge device , is the time for the edge device to infer the partition of the dataset ; S204. Calculate any edge device according to the following formula The proportion of the graph neural network inference ability on any graph dataset is as follows : ; S205. Construct a graph neural network inference ability matrix according to the proportion of the graph neural network inference ability, and calculate any edge device according to the following formula The proportion of the graph neural network inference ability of any edge device in a distributed environment is as follows , where is the number of graph datasets in the benchmark dataset

[0008] Optionally, among different types of graph datasets in step S201, the graph data types involved include some or all of knowledge graphs, citation networks, e-commerce networks, open-source community networks, social networks, financial transaction networks, and academic cooperation networks

[0009] Optionally, sorting the vertices in the graph data by the degree of the vertices in descending order and screening the first-order seeds and second-order seeds in step S102 includes S301. Sort the vertices in the graph data by the degree of the vertices in descending order S302. Select the vertices with the top preset proportion from the sorted vertex set to form a seed vertex set ; S303. Select the top highest-degree vertices from the seed vertex set as the first-order seeds, and select the remaining vertices as the second-order seeds, where is the total number of edge devices in the distributed environment, and is the number of seed vertices in the seed vertex set

[0010] Optionally, when randomly assigning partitions to the first-order seeds and their second-order neighbor blocks in step S102, it includes: preferentially assigning the first-order seeds to the partition with the largest capacity according to the degree; for the second-order neighbor blocks of the first-order seeds, using the proportion of the graph neural network inference ability of the edge device in the distributed environment as the probability of assigning each second-order neighbor block to the edge device, so as to randomly assign each second-order neighbor block of the first-order seeds to the partitions corresponding to edge devices in the distributed environment

[0011] ​Optionally, when constructing a depth expansion tree for the second-order neighbor blocks of the second-order seeds in step S102 and allocating partitions by block according to the proportion of the graph neural network inference ability of the edge device in the distributed environment, for each allocated second-order seed in the second-order neighbor blocks of the first-order seeds , collect the set composed of the unallocated vertices in the second-order neighbor block of this second-order seed . If the size of the set is 0, exit the processing of the second-order neighbor blocks of the second-order seeds and enter the stage of allocating partitions to the remaining unallocated vertices using the streaming greedy allocation strategy; otherwise, allocate partitions to the set according to the following formula: , where represents the allocated partition, is the m-th partition, is the number of vertices that the m-th partition can accommodate, is the number of existing vertices in the m-th partition, is the number of vertices that the j-th partition can accommodate, is the j-th partition, is the number of existing vertices in the j-th partition, and the calculation function expression for the number of vertices that any partition can accommodate is: , where is the number of vertices that the m-th partition can accommodate, is the number of vertices in the graph dataset, is the edge device 's proportion of the graph neural network inference ability in the distributed environment. The m-th partition is the partition corresponding to the edge device .

[0012] Optionally, when allocating partitions to the remaining unallocated vertices using the streaming greedy allocation strategy in step S102, allocating partitions to any vertex using the streaming greedy allocation strategy includes: S401, collect the set of partitions that contain the allocated first-order neighbors of the vertex , where is the j-th partition; S402, if the set of partitions is not empty, jump to step S403; otherwise, jump to step S4O5; S403, for the set of partitions , calculate the number of neighbors of each partition according to the following formula: , Among them, is the number of neighbors of the j-th partition, is the vertex assigned first-order neighbor blocks, is the vertex; S404. Determine whether the device capacity of all partitions where the assigned first-order neighbors of vertex is saturated. If it is, jump to step S405. Otherwise, assign a partition to vertex according to the following formula: , where represents the assigned partition, is the j-th partition, is the number of existing vertices in the j-th partition, is the number of vertices that the j-th partition can accommodate; S405. Assign a partition to vertex according to the following formula: ; And the calculation function expression of the number of vertices that any partition can accommodate is: , where is the number of vertices that the m-th partition can accommodate, is the number of vertices in the graph dataset, is the edge device in the proportion of the graph neural network inference ability in the distributed environment. The m-th partition is the partition corresponding to the edge device .

[0013] In addition, the present invention also provides a graph partitioning system for graph neural network inference in a heterogeneous edge scenario, including a microprocessor and a memory connected to each other. The microprocessor is programmed or configured to execute the graph partitioning method for graph neural network inference in the heterogeneous edge scenario.

[0014] In addition, the present invention also provides a computer-readable storage medium. A computer program or instruction is stored in the computer-readable storage medium. The computer program or instruction is programmed or configured to execute the graph partitioning method for graph neural network inference in the heterogeneous edge scenario through a processor.

[0015] In addition, the present invention also provides a computer program product, including a computer program or instruction. The computer program or instruction is programmed or configured to execute the graph partitioning method for graph neural network inference in the heterogeneous edge scenario through a processor.

[0016] Compared with the prior art, the present invention can mainly achieve the following beneficial effects: In order to solve the problems of load imbalance and low inference efficiency caused by device heterogeneity and defects of traditional graph partitioning methods in the edge graph neural network inference scenario, as well as the three core defects of the existing METIS-based balanced graph partitioning strategy in the edge environment - ignoring the impact of the message passing mechanism of the graph neural network on the inference efficiency, high memory overhead being inapplicable to resource-constrained devices, and lacking the ability to perceive the inference ability of heterogeneous graph neural networks, the present invention proposes a novel framework for heterogeneous edge graph neural network inference for the graph partitioning method of graph neural network inference in heterogeneous edge scenarios. By evaluating the computing power of heterogeneous devices and performing graph partitioning that combines computing power perception and message passing optimization, the inference efficiency of graph neural networks in heterogeneous edge environments is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of the basic process of the method in an embodiment of the present invention.

[0018] Figure 2 It is a schematic diagram of the process of graph data partitioning in an embodiment of the present invention.

[0019] Figure 3 It is a schematic diagram of the module structure in an embodiment of the present invention.

[0020] Figure 4 It is a schematic diagram of the process of the proportion of graph neural network inference ability in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The present invention aims to achieve four core objectives: (1) Support graph partitioning in heterogeneous environments to achieve load balancing; (2) Optimize the time and memory costs of the partitioning process, as this process itself needs to be executed on resource-constrained edge devices; (3) Be able to integrate with mainstream graph neural network processing libraries (such as PyG); (4) Adapt to diverse data sets, edge devices, and application scenarios. To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings in the embodiments of the present invention.

[0022] As Figure 1 shown, the graph partitioning method for heterogeneous edge graph neural network inference in this embodiment includes the following steps: S101, evaluating the proportion of the graph neural network inference ability of edge devices in a distributed environment through benchmark testing; S102, according to the proportion of the graph neural network inference ability of edge devices in a distributed environment, performing graph partitioning on the input graph data to the partitions corresponding to the edge devices.

[0023] AsFigure 2 As shown, in this embodiment, step S102 includes: Step 1, arranging the vertices in the graph data in descending order of the vertex degrees and screening the first-order seeds and second-order seeds; Step 2, randomly assigning partitions to the first-order seeds and their second-order neighbor blocks; Step 3, constructing a depth expansion tree for the second-order neighbor blocks of the second-order seeds and allocating partitions to the blocks according to the proportion of the graph neural network inference ability of the edge devices in the distributed environment; Step 4, using a streaming greedy allocation strategy to allocate partitions to the remaining unallocated vertices.

[0024] The graph data can be graph data in fields such as knowledge graphs, citation networks, e-commerce networks, open-source community networks, social networks, financial transaction networks, and academic collaboration networks. In these fields, the meanings of nodes and edges are defined according to specific application scenarios and data characteristics, and they together constitute the basic structure of graph data for describing complex entity relationships and interaction patterns. For example: (1) In the graph data of a knowledge graph, nodes usually represent entities such as person names, place names, organizations, concepts, events, etc. Edges represent the relationships between nodes, such as "belong to", "be located in", "participate in", and "cause", etc. (2) In the graph data of a citation network, nodes represent documents such as papers, books, research reports, etc. Each academic document can be a node. Edges represent the citation relationships between documents. If document A cites document B, then there is an edge in the graph pointing from A to B. Such an edge reflects the inheritance of knowledge and the context of academic research. (3) In the graph data of an e-commerce network, nodes can be products, users, merchants, stores, etc. For example, a product (such as a mobile phone), a user (consumer), and a merchant (brand owner) can all be nodes. Edges represent the interaction relationships between nodes. For example, there can be edges such as "purchase", "browse", "favorite" between a user and a product; there is an edge representing the "sale" relationship between a merchant and a product; there is an edge representing the "settle in" relationship between a merchant and a store. (4) In the graph data of an open-source community network, nodes can be developers, code repositories, projects, programming languages, etc. For example, a developer, an open-source project, and a programming language can all be nodes. Edges represent the associations between nodes. For example, there is an edge representing the "contribute" relationship between a developer and a project; there is an edge representing the "use" relationship between a project and a programming language; there can be an edge representing the "cooperate" relationship between developers. (5) In the graph data of a social network, nodes represent users (individuals or organizations). Edges represent the social relationships between users. For example, if user A follows user B, then there is an edge in the graph pointing from A to B. (6) In the graph data of a financial transaction network, nodes can be financial institutions (such as banks, securities companies), customers (individuals or enterprises), trading accounts, financial products, etc. For example, a bank, a stock account, and a financial product can all be nodes. Edges represent transaction or fund flow relationships. For example, if customer A transfers money from bank B to customer C, then there is an edge in the graph pointing from A's account to C's account, and the weight of the edge can be the transaction amount. (7) In the graph data of an academic collaboration network, nodes represent researchers, research institutions, laboratories, etc. For example, a university professor, a research team, and a research institute can all be nodes. Edges represent cooperation relationships. For example, if two professors jointly publish a paper, then there is an edge connecting their two nodes in the graph; if a research team collaborates with another research institution on a project, there will also be an edge representing this cooperation relationship.

[0025] Such as Figure 3As shown, as an alternative implementation, the method of this embodiment consists of a heterogeneous computing power evaluation module HETER and a graph partitioning algorithm module GSD (Graph Seed Division). Among them, the heterogeneous computing power evaluation module HETER is used to execute step S101, and the graph partitioning algorithm module GSD is used to execute step S102. On this basis, for the convenience of distinction, the method of this embodiment is named HETER-GSD.

[0026] In this embodiment, the heterogeneous computing power evaluation module HETER constructs a benchmark test suite covering 8 types of graph datasets with topological features, and provides quantitative guidance for the graph partitioning algorithm by accurately measuring the computing power of heterogeneous devices in the GNN inference task. The graph partitioning algorithm module GSD is designed based on the message passing characteristics of GNN inference, and adopts a hierarchical seed progressive allocation strategy to effectively reduce the peak memory consumption and partitioning time overhead while balancing the cross-partition communication caused by neighbor sampling operations and optimizing the message passing efficiency. As Figure 4 shown, step S101 in this embodiment includes: S201, constructing a benchmark test dataset, where the benchmark test dataset includes different graph datasets with different topological structures and different types The different topological structures refer to partial or all of the number of vertices, the number of edges, the average degree of vertices, and the feature dimension of vertices being the same; S202, for each graph dataset, using the edge cut graph partitioning method METIS to divide the graph dataset into non-overlapping partitions, where is the total number of edge devices in the distributed environment; the heterogeneous computing power evaluation module HETER uses the widely used edge cut graph partitioning method METIS to perform balanced partitioning into 𝑀 non-overlapping partitions, and its vertex cardinality is , and the edge cut graph partitioning method METIS is selected because the edge cut graph partitioning method METIS is a widely used graph partitioning method worldwide, and its characteristics of minimizing cut edges and relatively balanced load partitioning lay a reliable foundation for subsequent computing power evaluation; S203, performing graph neural network inference on the benchmark test dataset through the distributed environment, and calculating the ability of any edge device to infer any graph dataset for the partition according to the following formula : , where, is the number of vertices of the non-overlapping partition of the dataset partitioned to the edge device , is the edge device inferring the dataset Time of partitioning; this step aims to establish a unified metric through standardized GNN inference tasks to quantify the computational efficiency of heterogeneous devices in GNN inference, so as to achieve an accurate evaluation of the performance differences in GNN inference, where edge devices Inference dataset The time of partitioning includes: vertex sampling time, n-order (n-hop) neighbor feature aggregation time, and feature update time; S204, calculate the proportion of the graph neural network inference ability of any edge device On any graph dataset Proportion of the graph neural network inference ability : ; S205, construct a graph neural network inference ability matrix according to the proportion of the graph neural network inference ability. For example Figure 4 The graph neural network inference ability matrix constructed by 4 edge devices for 8 graph datasets in 11 ~R 41 Is the graph neural network inference ability of 4 edge devices for graph dataset 1, where the second column R 12 ~R 42 Is the graph neural network inference ability of 4 edge devices for graph dataset 2, and so on. The eighth column R 18 ~R 48 Is the graph neural network inference ability of 4 edge devices for graph dataset 8; and calculate the proportion of the graph neural network inference ability of any edge device In a distributed environment : , Among them, Is the number of graph datasets in the benchmark dataset. All edge devices In a distributed environment The set of the proportions of the graph neural network inference ability constitutes the proportion of the computing power derived by the heterogeneous computing power evaluation module HETER. This method overcomes the limitations of the traditional single-index evaluation system through the graph neural network inference performance evaluation of multiple datasets and multiple tasks, accurately reflects the true computing power of devices in the graph neural network inference scenario, and provides a reliable quantitative basis for the resource scheduling of heterogeneous devices.

[0027] In step S201 of this embodiment, different types of Among the graph datasets involved, the types of graph data include some or part of knowledge graphs, citation networks, e-commerce networks, open-source community networks, social networks, financial transaction networks, and academic collaboration networks. Specifically, in this embodiment, the benchmark test suite constructed by the heterogeneous computing power evaluation module HETER covers graph datasets with 8 different topological structures: knowledge graph (WikiCS), citation network (ogbn-arxiv), e-commerce network (Amazon), open-source community network (GitHub), social network (Flickr and FaceBookPage), financial transaction network (ElliptiBitcoinDataset), and academic collaboration network (Coauthor). These graphs cover a wide range of vertex numbers, edge densities, and feature dimensions, and can comprehensively characterize the performance of GNN inference under heterogeneous workloads. In addition, in view of the low memory capacity of edge devices, HETER ensures the availability of the framework by selecting graph datasets with less memory occupancy during the distributed inference process, as shown in Table 1 specifically.

[0028] Table 1: Information table of graph datasets in the benchmark test dataset

[0029] In Table 1, the memory occupancy is the memory occupancy when running in a distributed environment composed of 5 heterogeneous edge devices.

[0030] It can be seen that the heterogeneous computing power evaluation module HETER in this embodiment constructs a lightweight benchmark suite covering 8 topological structure graph datasets such as knowledge graphs, social networks, and financial transaction networks, breaking through the limitations of traditional evaluations based on a single metric (such as FLOPs), and realizing the accurate evaluation of the distributed GNN inference ability of edge devices. The heterogeneous computing power evaluation module HETER, based on the benchmark test and dynamic weighted quantization method, comprehensively considers the latency characteristics of the entire process of GNN inference, and adopts a cross-dataset averaging calculation model for the proportion of device comprehensive computing power to eliminate the bias of a single dataset, ensuring the universality and robustness of the evaluation results, and providing an accurate heterogeneous device ability profile for subsequent graph partitioning.

[0031] In this embodiment, the graph partitioning algorithm module GSD also includes initializing the number of vertices that a partition can accommodate according to the proportion of the graph neural network inference ability of edge devices in a distributed environment. This step aims to set an upper threshold to prevent device overload. In view of the lack of a widely recognized GNN task size quantization standard currently, and the number of partition vertices has characteristics such as being directly related to the computing requirements, being convenient for statistics, and being widely recognized in the industry, so this embodiment uses this vertex number as the index for calculating the device capacity. The calculation function expression for the number of vertices that any partition can accommodate is: , Among them, is the number of vertices that the m-th partition can accommodate, is the number of vertices in the graph dataset, is the edge device The proportion of the inference ability of the graph neural network in the distributed environment, and the m-th partition is the partition corresponding to the edge device corresponding to it, so that the capacity of edge devices in the distributed environment can be obtained .

[0032] Refer to Figure 2 Step 1 in. In step S102 of this embodiment, the vertices in the graph data are sorted in descending order of the vertex degrees and the first-order seeds and second-order seeds are screened. The high-degree vertices are preferentially selected as seeds for the following two reasons: First, in a power-law graph, the high-degree vertices and their neighbors will form large star structures. If these star structures are aggregated in the same subgraph, it will bring uneven communication in the distributed environment. Therefore, eliminating the neighbor blocks formed by the high-degree vertices and their neighbors can be an effective way to eliminate the uneven communication in the distributed environment. Second, according to the message passing mechanism of the GNN, the overhead generated within the local neighborhood of the vertex is the smallest ( is the vertex centered formed by the layer neighbors). Therefore, retaining the neighborhood blocks of the high-degree vertices during partitioning can effectively reduce the cross-partition communication. Step S102 of this embodiment for sorting the vertices in the graph data in descending order of the vertex degrees and screening the first-order seeds and second-order seeds includes: S301, sorting the vertices in the graph data in descending order of the vertex degrees; S302, selecting the vertices in the top preset ratio (which can be determined according to actual needs) from the vertex set sorted in descending order to form a seed vertex set ; S303, selecting the first highest-degree vertices from the seed vertex set as the first-order seeds, and selecting the remaining vertices as the second-order seeds, where is the total number of edge devices in the distributed environment, is the number of seed vertices in the seed vertex set

[0033] When randomly assigning partitions to the first-order seeds and their second-order neighbor blocks in step S102 of this embodiment, it includes: preferentially assigning the first-order seeds to the partition with the largest capacity according to the degree size; for the second-order neighbor blocks of the first-order seeds , the proportion of the graph neural network reasoning ability of the edge device in the distributed environment is used as the probability of assigning the second-order neighbor block to the edge device, so that the second-order neighbor block of each first-order seed is randomly assigned to the distributed environment according to the probability For example, in a distributed environment consisting of three edge devices, the proportion of graph neural network reasoning capabilities of the three edge devices in the distributed environment is calculated to be , then in this step the first-order seed Neighbor Blocks The first device randomly assigns 50% of the vertices of the neighbor block, the second device assigns 30% of the vertices of the neighbor block, and the third device assigns 20% of the vertices of the neighbor block.

[0034] In step S102 of this embodiment, when a deep expansion tree is constructed for the second-order neighbor blocks of the second-order seeds and partitions are allocated by block according to the ratio of the graph neural network reasoning capabilities of edge devices in a distributed environment, the above steps allocate the second-order seeds and their second-order neighbor blocks, and gradually build a deep expansion tree by combining the GNN message passing mechanism to maintain the integrity of message passing within the partition. Specifically: for the seeds that have been allocated as neighbors of the first-order seeds , collect its The set of vertices in the neighbor block that have not yet been assigned , constructed with The purpose of establishing a 2-layer deep tree is to use the message passing mechanism between neighbors of GNN to transfer seeds and their Placing neighbors in the same partition can maximize the integrity of message delivery, reduce the communication overhead caused by cross-partition sampling, and improve the efficiency of graph partitioning. Overall allocation to Specifically, the above steps include allocating the allocated second-order seed in the second-order neighbor block of the first-order seed to the partition where the first-order seed is located, otherwise it is allocated to the partition with the largest remaining capacity. , collect the second-level seeds The set of vertices in the second-order neighbor block that have not yet been assigned , if the set If the size of is 0, then exit the processing of the second-order neighbor block of the second-order seed and enter the partitioning phase using the streaming greedy allocation strategy for the remaining unassigned vertices; otherwise, according to the following formula, Assign partitions: , in, Indicates the allocated partition, is the mth partition, is the number of vertices that the m-th partition can accommodate. is the number of existing vertices in the m-th partition. is the number of vertices that the j-th partition can accommodate. is the j-th partition. is the number of existing vertices in the j-th partition. Among them, represents a vertex 's the number of unassigned vertices in the neighbor blocks is less than or equal to the remaining capacity of the device Therefore, the overall meaning expressed by this formula is: if the 's the number of unassigned vertices in the neighbor blocks is less than or equal to the remaining capacity of the device then the set is allocated to the device as a whole; otherwise, the set is allocated to the device with the largest remaining capacity among all devices. In addition, the above steps include an early exit mechanism where the size of the set is 0. When it is detected that the number of unassigned vertices in the seed block drops to 0, the second-order seed allocation is terminated and step 4 (allocating partitions to the remaining unassigned vertices using a streaming greedy allocation strategy) is entered. The reason for setting the early exit mechanism is: based on the power-law characteristics of the graph structure, seeds and their usually cover most vertices, and the remaining vertices are mostly low-degree edge vertices, with characteristics such as sparse connections, weak community properties, and discrete spatial distributions. Calculating for these vertices will incur computational overhead but with limited benefits, and it is cost-ineffective to continue the block-based partitioning in step 3.

[0035] The vertices that remain unassigned after the previous steps are the remaining vertices. In this embodiment, a streaming greedy allocation under capacity constraints is performed on the remaining vertices to maximize the local computational benefit while ensuring load balancing. When using the streaming greedy allocation strategy to allocate partitions to the remaining unassigned vertices in step S102 of this embodiment, for any vertex using the streaming greedy allocation strategy to allocate partitions includes: S401, collect the set of partitions that contain the already allocated first-order neighbors of the vertex Among them, is the j-th partition; S402, if the set of partitions is not empty, then jump to step S403; otherwise, jump to step S4O5; S403, for the set of partitions , calculate the number of neighbors of each partition according to the following formula: , wherein, is the number of neighbors of the j-th partition, is the vertex assigned first-order neighbor blocks, is the vertex, and this formula represents the vertex of the first-order neighbors that have been assigned to the partition ; S404, determine whether the device capacity of all partitions where the assigned first-order neighbors of the vertex are located is saturated. If it is, jump to step S405. Otherwise, allocate a partition to the vertex according to the following formula: , wherein, represents the allocated partition, is the j-th partition, is the number of vertices existing in the j-th partition, is the number of vertices that the j-th partition can accommodate. This formula represents that under the premise that the number of vertices already assigned to the partition is less than the capacity of this partition, the vertex is allocated to the partition with the most first-order neighbors; S405, allocate a partition to the vertex according to the following formula: , represents the target partition which is the device with the largest remaining capacity among all devices; and the calculation function expression of the number of vertices that any partition can accommodate is: , wherein, is the number of vertices that the m-th partition can accommodate, is the number of vertices in the graph dataset, is the edge device in the proportion of the graph neural network inference ability in the distributed environment. The m-th partition is the partition corresponding to the edge device

[0036] It can be seen that the graph partitioning algorithm module GSD in this embodiment proposes a hierarchical seed progressive partitioning mechanism, and different allocation methods are executed by selecting first-order / second-order seeds. By randomly allocating the neighbor blocks of the first-order seeds, the power-law property of the graph is effectively eliminated. For the GNN message passing mechanism, a Neighboring blocks are used to maximize the neighborhood integrity within the partition to reduce cross-partition communication. An early exit mechanism is designed to dynamically terminate the high-cost partitioning phase by monitoring the number of unassigned seed blocks in real time. Finally, a streaming greedy allocation with capacity constraints is implemented for the remaining vertices with low degrees to maximize the local computing benefit while ensuring load balancing. The graph partitioning algorithm module GSD in this embodiment supports receiving the computing power ratio parameters of heterogeneous devices to achieve unbalanced partitioning under heterogeneous computing power.

[0037] The advantages of the method (HETER-GSD) in this embodiment are as follows in three aspects: (1) The default processing method for distributed graph neural network tasks in current mainstream graph neural network processing libraries is METIS balanced partitioning, which causes serious load imbalance in edge devices and reduces the inference efficiency. The method (HETER-GSD) in this embodiment is an acceleration framework designed for heterogeneous edge graph neural network inference. The advantages are as follows: The method (HETER-GSD) in this embodiment supports unbalanced partitioning of graph neural network inference tasks under computing power evaluation, effectively promoting load balance between heterogeneous devices and improving the utilization rate of computing power resources and inference efficiency; The method (HETER-GSD) in this embodiment is a lightweight edge graph neural network inference acceleration framework that can be integrated into the mainstream graph neural network processing library PyG, suitable for running in edge environments and convenient to call. (2) Existing computing power evaluation methods mostly focus on the evaluation of single computing capabilities (such as floating-point computing capabilities, etc.), while graph neural network inference involves complex mathematical calculations and cannot be measured by a single indicator; moreover, existing evaluation methods are difficult to accurately evaluate the sampling randomness of graph neural networks and the dynamic communication overhead introduced in distributed computing. Therefore, the advantage of the heterogeneous computing power evaluation module HETER is as follows: This module is a computing power evaluation method designed for graph neural network inference in heterogeneous edge environments. By constructing a lightweight benchmark test suite consisting of 8 graph datasets with different topological structures, it evaluates the graph neural network inference capabilities of heterogeneous devices from the full process of distributed graph neural network inference, effectively making up for the deficiencies of traditional computing power evaluation methods. (3) Traditional graph partitioning methods take the edge cut rate as the core indicator without considering the message passing characteristics in graph neural network inference, and traditional graph partitioning methods are prone to retaining large star-shaped structures formed by high-degree vertices and their neighbors in a certain subgraph, resulting in unbalanced sampling between different devices. The graph partitioning algorithm module GSD, as a graph partitioning algorithm optimized for graph neural network propagation, has the following advantages: The graph partitioning algorithm module GSD eliminates the unbalanced sampling caused by large-shaped structures through a multi-layer seed progressive allocation strategy, maximizes the integrity of neighbor blocks in the subgraph, and reduces the communication overhead generated by cross-partition sampling using the message passing mechanism of graph neural networks; The block partitioning and streaming greedy partitioning adopted in the graph partitioning algorithm module GSD effectively reduce the memory occupancy during the partitioning process, improve the partitioning efficiency, and are suitable for running on edge devices; The graph partitioning algorithm module GSD supports simple calls integrated on PyG, supports receiving heterogeneous computing power parameters, and realizes unbalanced partitioning according to the strength of computing power.

[0038] In addition, this embodiment also provides a graph partitioning system for graph neural network inference in a heterogeneous edge scenario, including a microprocessor and a memory connected to each other. The microprocessor is programmed or configured to execute the graph partitioning method for graph neural network inference in the heterogeneous edge scenario. This embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the graph partitioning method for graph neural network inference in the heterogeneous edge scenario through a processor. This embodiment also provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the graph partitioning method for graph neural network inference in the heterogeneous edge scenario through a processor.

[0039] Those skilled in the art should understand that the technical solutions provided by the present invention can be in the form of a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 a block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 a block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1Steps of the functions specified in one or more boxes.

[0040] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. A graph partitioning method for graph neural network inference in heterogeneous edge scenarios, characterized in that, Including the following steps: S101, evaluating the proportion of the graph neural network inference ability of edge devices in a distributed environment through benchmark testing; S102, according to the proportion of the graph neural network inference ability of edge devices in a distributed environment, perform graph partitioning on the input graph data into partitions corresponding to edge devices, including: arranging the vertices in the graph data in descending order of vertex degree and screening first-order seeds and second-order seeds; randomly assigning partitions to the first-order seeds and their second-order neighbor blocks; constructing a depth expansion tree for the second-order neighbor blocks of the second-order seeds and allocating partitions by block according to the proportion of the graph neural network inference ability of edge devices in a distributed environment; and allocating partitions to the remaining unassigned vertices using a streaming greedy allocation strategy. For the second-order neighbor blocks of the second-order seeds, construct a depth expansion tree and allocate partitions by block in combination with the proportion of the graph neural network inference ability of edge devices in a distributed environment; for the remaining unassigned vertices, use a streaming greedy allocation strategy to allocate partitions.

2. The graph partitioning method for graph neural network inference in heterogeneous edge scenarios according to claim 1, wherein Step S101 includes: S201, Construct a benchmark test dataset, which includes different topological structures and different types of graph datasets. The different topological structures refer to partial or all of the number of vertices, the number of edges, the average degree of vertices, and the feature dimension of vertices being different; S202. For each graph dataset, use the edge cut graph partitioning method METIS to partition the graph dataset into non-overlapping partitions, where is the total number of edge devices in the distributed environment; S203, perform graph neural network inference on the benchmark test dataset in a distributed environment, and calculate the ability of any edge device to infer any graph dataset for the partition : , Among them, is the data set partitioned to the edge device the number of vertices of the non-overlapping partitions on, is the edge device the inference data set the time of the partition; S204, calculate any edge device according to the following formula In any graph data set The proportion of the inference ability of the graph neural network : ; S205. Construct a graph neural network inference ability matrix based on the proportion of the inference ability of the graph neural network, and calculate the proportion of the inference ability of the graph neural network of any edge device in a distributed environment according to the following formula : , Among them, is the number of graph datasets in the benchmark dataset.

3. The graph partitioning method for graph neural network inference in heterogeneous edge scenarios according to claim 2, wherein In step S201, among different types of graph datasets, the involved graph data types include some or all of knowledge graphs, citation networks, e-commerce networks, open-source community networks, social networks, financial transaction networks, and academic cooperation networks.

4. The graph partitioning method for graph neural network inference in heterogeneous edge scenarios according to claim 1, wherein In step S102, sorting the vertices in the graph data in descending order of the vertex degrees and screening first-order seeds and second-order seeds includes: S301, sorting the vertices in the graph data in descending order of the vertex degrees; S302, select the vertices at the front with a preset ratio from the vertex set arranged in descending order to form a seed vertex set ; S303, from the seed vertex set Select the first vertices with the highest degrees as the first-order seeds, and select the remaining vertices as the second-order seeds, where is the total number of edge devices in the distributed environment, is the set of seed vertices and is the number of seed vertices in the set.

5. The graph partitioning method for graph neural network inference in heterogeneous edge scenarios according to claim 1, wherein When randomly assigning partitions to first-order seeds and their second-order neighbor blocks in step S102, it includes: preferentially assigning first-order seeds to the partition with the largest capacity according to the degree; for the second-order neighbor blocks of the first-order seeds, taking the proportion of the graph neural network inference ability of the edge device in the distributed environment as the probability of assigning the second-order neighbor block to the edge device, so as to randomly assign the second-order neighbor blocks of each first-order seed to the partitions corresponding to the edge devices in the distributed environment according to the probability.

6. The graph partitioning method for graph neural network inference in heterogeneous edge scenarios according to claim 1, wherein In step S102, a deep expansion tree is constructed for the second-order neighbor block of the second-order seed and partitions are allocated by block according to the graph neural network reasoning capability ratio of the edge device in a distributed environment, including each allocated second-order seed in the second-order neighbor block of the first-order seed. , collect the second-level seeds The set of vertices in the second-order neighbor block that have not yet been assigned , if the set If the size is 0, then exit the processing of the second-order neighbor blocks of the second-order seed and enter the partition allocation phase using the streaming greedy allocation strategy for the remaining unallocated vertices; Otherwise, partition the set according to the following formula Allocate partitions: , Among them, represents the allocated partition, is the m-th partition, is the number of vertices that the m-th partition can accommodate, is the number of existing vertices in the m-th partition, is the number of vertices that the j-th partition can accommodate, is the j-th partition, is the number of existing vertices in the j-th partition, and the calculation function expression for the number of vertices that any partition can accommodate is: , Among them, is the number of vertices that the m-th partition can accommodate, is the number of vertices in the graph dataset, is the edge device is the proportion of the graph neural network inference ability of the edge device in the distributed environment, and the m-th partition is the partition corresponding to the edge device.

7. The graph partitioning method for graph neural network inference in heterogeneous edge scenarios according to claim 1, wherein When partitioning the remaining unallocated vertices using the streaming greedy allocation strategy in step S102, for any vertex Partitioning using the streaming greedy allocation strategy includes: S401, collect the partitions containing the vertices of the first-order neighbors that have been assigned , where is the j-th partition; S402, if the partition set is non-empty, then jump to step S403; otherwise, jump to step S405; S403, for the partition set , calculate the number of neighbors for each partition according to the following formula: , Among them, is the number of neighbors of the j-th partition, is the vertex already allocated first-order neighbor blocks, is the vertex; S404, Determine the vertex Whether the total capacity of all partition devices where the assigned first-order neighbors are located is saturated. If it holds, jump to step S405; otherwise, according to the following formula, assign a partition to the vertex as follows: , Among them, represents the allocated partition, is the j-th partition, is the number of existing vertices in the j-th partition, is the number of vertices that the j-th partition can accommodate; S405, allocate a partition to the vertex according to the following formula : ; And the calculation function expression for the number of vertices that any partition can accommodate is: , Among them, is the number of vertices that the m-th partition can accommodate, is the number of vertices in the graph dataset, is the edge device in the proportion of the graph neural network inference ability in the distributed environment. The m-th partition is the partition corresponding to the edge device corresponding.

8. A graph partitioning system for graph neural network inference in heterogeneous edge scenarios, including a microprocessor and a memory connected to each other, characterized in that, The microprocessor is programmed or configured to execute the graph partitioning method for graph neural network inference in a heterogeneous edge scenario according to any one of claims 1 to 7.

9. A computer-readable storage medium storing a computer program or instructions, characterized in that, The computer program or instruction is programmed or configured to execute the graph partitioning method for graph neural network inference in a heterogeneous edge scenario according to any one of claims 1 to 7 through a processor.

10. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or instruction is programmed or configured to execute the graph partitioning method for graph neural network inference in a heterogeneous edge scenario according to any one of claims 1 to 7 through a processor.

Citation Information

Patent Citations

  • Graph-based model division edge-end collaborative reasoning method and system

    CN117707795A

  • Distributed graph neural network training method based on heterogeneous equipment

    CN119089968A

  • Graph neural network partitioning method based on heterogeneous GPU load balancing

    CN119311414A

  • Graph neural network-oriented block-based graph data partitioning method and system

    CN120086024A

  • Techniques for learning co-engagement and semantic relationships using graph neural networks

    US20250148280A1

Cited By

  • Graph division method and system oriented to heterogeneous environment graph neural network and based on degree classification

    CN120597932A

  • Graph neural network-based graph partitioning method and system for heterogeneous environment

    CN120597932B