FPGA acceleration-based graph neural network explanation method and device
By accelerating the interpretation process of graph neural network on FPGA, calculating the k-hop neighbor node set and sub-graph adjacency matrix, the problem of high time complexity of interpretation process in the prior art is solved, and the effect of quickly generating interpretation results is achieved.
Patent Information
- Application Number
- PCT/CN2023/139066
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2025-06-19
AI Technical Summary
The prior art has high time complexity in the node classification task of graph neural networks and cannot quickly provide explanation results.
Using the graph neural network interpretation method based on FPGA acceleration, the calculation of the k-hop neighbor node set and subgraph adjacency matrix is realized on the FPGA, and the interpretation subgraph is obtained and the computing efficiency is optimized.
Through FPGA acceleration, the time complexity of the graph neural network interpretation process is reduced, the speed of generation of interpretation results is improved, and the practical application scenarios of graph neural network interpretation is expanded.
Smart Images

Figure CN2023139066_19062025_PF_FP_ABST
Abstract
Description
A graph neural network interpretation method and device based on FPGA acceleration Technical Field
[0001] The present invention relates to the field of graph neural network interpretation technology, and in particular to a graph neural network interpretation method and device based on FPGA acceleration. Background Art
[0002] Graph neural networks have demonstrated promising performance in numerous scientific fields, such as natural language processing, computer vision, biology, and social networks. Due to their superior direct inference capabilities and high efficiency, graph neural networks have been widely adopted in various graph tasks, such as node and graph classification. However, when using graph neural networks to solve these problems, they are often viewed as complex and opaque. As our requirements for the security, reliability, and performance of graph neural networks gradually increase, interpreting graph neural networks has become increasingly important.
[0003] The first explanation method specifically designed for graph neural networks was GNNExplainer, a model-agnostic, general-purpose graph neural network explainer whose primary purpose was to extract important edges and features as explanations for the model. Later, PGExplainer used parameterized learning of the explanation generation process to enable simultaneous explanation of multiple instances. These classic methods focused primarily on explaining node or edge feature dimensions, while neglecting the substructure of the graph. Subgraphx proposed a subgraph explanation method for graph neural networks, which provides a more intuitive explanation of graph neural networks and exhibits excellent explanation performance.
[0004] While the aforementioned classic algorithms perform well in graph classification tasks, few have addressed node classification and validated their effectiveness using real-world data. Furthermore, interpretation algorithms for node classification are time-consuming and inefficient when applied to real-world data, hindering rapid interpretation. FPGAs, with their high programmability, short development cycles, and efficient parallel computing, can accelerate these algorithms in real-world applications.
[0005] Summary of the Invention
[0006] The purpose of the present invention is to address the deficiencies of the existing technology and provide a graph neural network interpretation method and device based on FPGA acceleration.
[0007] The purpose of the present invention is achieved through the following technical solution: a graph neural network interpretation method based on FPGA acceleration, comprising the following steps:
[0008] (1) Obtain the target dataset and divide it into a training set and a test set;
[0009] (2) Use the training set to train the graph neural network to obtain a trained graph neural network; and use the trained graph neural network to classify each test node in the test set to obtain a test prediction label set;
[0010] (3) For each test node in the test set, the corresponding k-hop neighbor node set and subgraph adjacency matrix of each test node are obtained on the FPGA;
[0011] (4) For each test node in the test set, the corresponding explanation subgraph is obtained through the subgraph adjacency matrix corresponding to each test node.
[0012] Furthermore, the step (1) is specifically as follows: obtaining data of the node classification task scene and constructing a corresponding graph G; the graph G contains a total of N nodes, that is, the graph node set V = {x1, x2, ..., x n ,…,x N}, where x n Represents the nth node; then the graph node set V is divided into the training set and test set Among them, the training set T r There are a total of P nodes in is the training set T r Any node in the test set T e There are a total of Q nodes in For the test set T e Any node in; and get the training set T r The corresponding true label set in, For nodes The true label.
[0013] Furthermore, the step (2) specifically includes the following sub-steps:
[0014] (2.1) Training set T r Any node in The set of neighbor nodes in graph G is node Received from neighbor node set Each node Message propagation at each layer in a graph neural network in, l represents the lth layer of the graph neural network, l=1,2,…,l,…,L; Representation node In the feature representation of layer l-1, Representation node Initial feature representation of ;
[0015] (2.2) Node Messages after aggregation at each layer The calculation formula is as follows:
[0016] Among them, AGG(.) is the aggregation function;
[0017] (2.3) Graph neural network for aggregated messages and nodes Feature representation in the previous layer Perform nonlinear transformation to obtain nodes Feature representation at layer l The calculation formula is as follows:
[0018] Among them, Update(.) is the update function;
[0019] Finally, the node Final Embedding in Graph Neural Network
[0020] (2.4) For the training set T r Repeat substeps (2.1) to (2.3) for each node in to obtain the final embedding set Z r :Z r ={z1,z2,…,z p ,…,z P};
[0021] (2.5) According to the final embedding set Z r , use the fully connected layer to classify nodes, and use the cross entropy loss function to train and optimize the graph neural network to obtain the trained graph neural network:
[0022] (2.6) Use the trained graph neural network to test the test set T e Classify each node in and get the test prediction label set Y e : in, For the test set T e Any test node The predicted label is calculated as follows:
[0023] Furthermore, the step (3) specifically includes the following sub-steps:
[0024] (3.1) Obtain the adjacency matrix A and the k-order adjacency matrix A of graph G k ,in, a j,i is any element in the adjacency matrix A, j is element a j,i The first number of element a, i is j,i The second number, a j,i Represents the j-th node x in graph G j With the i-th node x i Are they connected to each other? When the jth node x j With the i-th node x i If they are connected to each other, the element value is 1: a j,i =1, otherwise the element value is 0: a j,i =0; Represents the jth node x j With the i-th node x i Are they connected to each other through at most k-1 intermediate nodes, when the jth node x j With the i-th node x i The element value is 1 when it is connected to each other through at most k-1 intermediate nodes: when When the i-th node x i is the jth node x j k-hop neighbor nodes, otherwise the element value is 0: k-order adjacency matrix A k The jth row is node x j The row, k-order adjacency matrix A k The jth column is node x j Column
[0025] And the adjacency matrix A of graph G and the k-order adjacency matrix A k Write to the DDR4 memory outside the FPGA; and set up a FIFO-cnt counter and 6 pre-processing units;
[0026] (3.2) From the test set T e Select any 6 test nodes and assign a corresponding preprocessing unit to any test node; each preprocessing unit is obtained from the k-order adjacency matrix A of graph G. k Scan the row of the corresponding test node to obtain the row data of the corresponding test node and the number of k-hop neighbor nodes of the corresponding test node;
[0027] Sort the obtained six test nodes in descending order of the number of k-hop neighbor nodes. The test node with a larger number of k-hop neighbor nodes shall be prioritized for step (3.3). If the number of k-hop neighbor nodes is the same, the test node with a smaller sequence number shall be prioritized for step (3.3).
[0028] (3.3) Set up 8 segmented traversal units, evenly divide the row data obtained by scanning the test node into 8 segments of row data and assign them to the 8 segmented traversal units in turn; each segmented traversal unit has a write initiation request and a write enable response. An arbitrator performs round-robin arbitration control on the requests initiated by the 8 segmented traversal units. Each segmented traversal unit traverses the assigned row data and finds the nodes corresponding to the second numbers of all elements with a value of 1 in graph G as the k-hop neighbor nodes of the test node, thereby obtaining the set of k-hop neighbor nodes of the test node;
[0029] And the k-order adjacency matrix A of the graph G is tested by the k-hop neighbor node set of the node k Perform adjacency matrix truncation: First, extract the corresponding row data from the k-order adjacency matrix of graph G according to the number of each k-hop neighbor node to obtain the row matrix of the test node; then extract the corresponding column data from the row matrix of the test node according to the number of each k-hop neighbor node to obtain the subgraph adjacency matrix of the test node;
[0030] (3.4) The six test nodes perform step (3.3) in sequence to obtain the corresponding k-hop neighbor node set and subgraph adjacency matrix;
[0031] (3.5) Then from the test set T e Select any 6 test nodes from the remaining test nodes and repeat steps (3.1) to (3.4) until the test set T e Each test node in the graph obtains its corresponding k-hop neighbor node set and subgraph adjacency matrix, and obtains the subgraph adjacency matrix set in, Represented as a test node The corresponding subgraph adjacency matrix.
[0032] Furthermore, the step (4) specifically includes the following sub-steps:
[0033] (4.1) When any test node in the test set The number of k-hop neighbor nodes When it is not greater than m, the test node The subgraph adjacency matrix of Perform the corresponding HN table operations;
[0034] (c1) Obtaining test nodes in FPGA hardware The result set of connectivity relationships of all permutations and combinations of nodes in the k-hop neighbor node set is obtained, and the matrix
[0035] (c2) will (H q ) 4 As Hq The iterative convergence value of (H q ) 4 The calculation is optimized, using the matrix The idempotent property is and (H q ) 4 It is divided into three sparse-dense matrix multiplications and two dense matrix multiplications. The sparse-dense matrix multiplication assigns tasks to the elements with a value of 1 in the sparse matrix and adopts a direct static mapping from matrix rows to 8 PEs to increase hardware parallel computing. There are gaps between the multiple rows assigned to each PE, and the search is carried out in a round-robin manner to find the element with a value of 1 in the middle of the row to prevent the sparse matrix from being concentrated in a few rows. Multiple rows of data to be processed are effectively connected.
[0036] (c3) According to the test node The subgraph adjacency matrix of The node numbers in the pre-calculate the eigenvalue results of all possible node combinations and splice the eigenvalue results into an eigenvalue vector according to the permutations and combinations of single nodes, double nodes, ..., m nodes
[0037] The eigenvalue of any node combination o is calculated using the following formula:
[0038] in, Indicates that all nodes in the node combination arrangement o are predicted to be The probability mean of ; Indicates that all nodes in the graph G are predicted to be The probability mean of ;
[0039] Get the HN value matrix vector Get the HN value matrix vector Center front The values are used as test nodes The set of k-hop neighbor nodes The HN value of the node is obtained The subgraph adjacency matrix of HN table, which is the test node The explanation subgraph of
[0040] (4.2) When any test node in the test set The number of k-hop neighbor nodes When it is greater than m, the test node The subgraph adjacency matrix of Implement layer-by-layer sampling based on the central node on FPGA to obtain the test node The set of sampled subgraph adjacency matrices
[0041] Then, the corresponding HN table operation is performed on each sampling subgraph adjacency matrix in the sampling subgraph adjacency matrix set to obtain the HN table corresponding to each sampling subgraph adjacency matrix; the test node The subgraph adjacency matrix of The HN value of each node in the adjacency matrix of different sampling subgraphs is the average of the HN values, and the test node is obtained. The explanation subgraph of
[0042] (4.3) According to the number of k-hop neighbor nodes of each test node in the test set, repeat step (4.1) or step (4.2) for each test node to obtain the explanation subgraph corresponding to each test node in the test set.
[0043] Furthermore, the sub-step (4.2) specifically includes the following sub-steps:
[0044] (4.2.1) When the test node The number of k-hop neighbor nodes When it is greater than m, for the test node The subgraph adjacency matrix of Test Node For the target node, for the test node The subgraph adjacency matrix of Perform shortest path traversal;
[0045] (4.2.2) On the test node The subgraph adjacency matrix of While traversing the shortest path, the sampling subgraph node set is generated synchronously to obtain the test node The set of sampled subgraph adjacency matrices
[0046] (4.2.3) Then, the corresponding HN table operation is performed on each sampling subgraph adjacency matrix in the sampling subgraph adjacency matrix set to obtain the HN table corresponding to each sampling subgraph adjacency matrix; test node The subgraph adjacency matrix of The HN value of each node in the adjacency matrix of different sampling subgraphs is the average of the HN values, and the test node is obtained. The explanation subgraph of .
[0047] Furthermore, the sub-step (4.2.1) specifically includes the following sub-steps:
[0048] (a1) First, perform traversal initialization: initialize the node on the FPGA to have all element values 0 and length List of: Used to identify whether the node has been visited, where node(c) represents the element value of the cth element in the list node; initialize flag to a queue containing only the number of the target node in the graph node set V: flag = [flag(1)] = [q′], which is used to store the neighbor node labels obtained after each BFS traversal of the node, where flag(1) is the element value of the first element in the list flag, and q′ is the target node The number in the graph node set V; initialize node_flag to the element value of the qth element to 1, the element values of the remaining elements to 0, and the length to The list of node_flag is: node_flag = [node_flag(1),…,node_flag(c),…,node_flag(7)], which is used to identify whether the node already exists in the flag, where node_flag(c) is the element value of the cth element in the list node_flag, and node_flag(q) = 1; initialize node_f to have all element values -1 and a length of List of: Used to identify the predecessor node of the node traversal process, where node_f(c) is the element value of the cth element in the list node_f; initialize bfs_cnt=0 to control the subscript address of each access queue flag; initialize node_pate to contain A list of sublists: Among them, node_pate[E] is the E-th sublist in the list node_pate, and each sublist contains elements and all element values are 0;
[0049] (a2) After the traversal is initialized, the element value of the bfs_cnt-th element in the list flag is accessed according to the value of bfs_cnt: flag(bfs_cnt); then the list node is updated: the element value of the flag(bfs_cnt)-th element in the list node is assigned to 1, and the node xflag(bfs_cnt) is used as the current traversal target node. At this time, it is determined whether the element value of the flag(bfs_cnt)-th element in the list node_f is -1: if it is -1, it means that the predecessor node of the current traversal target node is invalid, and then the subgraph adjacency matrix is read. The flag (bfs_cnt) line of , gets the row data, jumps to step (a3), and updates the limited node information; if it is not -1, it means that the predecessor node of the current traversal target node is valid, and the list node_pate is updated: the path information from the current traversal target node to the target node is stored in the list node_pate, and then the subgraph adjacency matrix is read. In the flag(bfs_cnt) line, we get the row data and jump to step (a3) to update the limited node information.
[0050] (a3) Traverse the valid elements of the read row data: the second number of the first element with a value of 1 is read from the row data as h: if the second number h is the same as the element value of any element in the list flag, the second number of the next element with a value of 1 is read from the row data; if the second number h is different from the element value of any element in the list flag, update the list flag: add the second number h to the list flag, and the updated list flag is flag=[flag(1),flag(1)]=[M,h], and update the list node_flag: assign the element value of the hth element in the list node_flag to 1, and record the node x h The predecessor node is the current traversal target node and the list node_f is updated: the element value of the hth element in the list node_f is assigned to the number M of the current traversal target node, and then the second number of the next element with an element value of 1 is obtained from the row data and the above steps are repeated until no element with an element value of 1 can be read from the row data, and the traversal ends;
[0051] After the traversal is completed, the value of bfs_cnt is updated: bfs_cnt = bfs_cnt + 1; repeat step (a2) according to the updated value of bfs_cnt until the updated value of bfs_cnt is When it jumps to the idle state, the test node is completed The subgraph adjacency matrix of Traverse the shortest path and obtain the updated list node_pate.
[0052] Furthermore, the step (4.2.2) specifically includes the following sub-steps:
[0053] (b1) First, perform sampling initialization: initialize node_out to contain A list of sublists: Among them, node_out[D] is the D-th sublist in the list node_out, each sublist contains m elements and the element values are all empty; node_y is initialized to have all element values 0 and length List of: Among them, node_y(c) is the element value of the cth element in the list node_y;
[0054] (b2) After sampling is initialized, the number of elements in the list flag is detected at any time: when the number of elements in the list flag is detected to be m, the distance test node is obtained The set of m nearest nodes: {x flag(1) ,x flag(2) ,…x flag(m)}, jump to step (b3);
[0055] (b3) Update the list node_out. The updated list node_out is node_out=[[x flag(1) ,x flag(2) ,…,x flag(m) ],…,[x flag(1) ,x flag(2) ,…,x flag(m) ]];
[0056] (b4) Check the updated value of bfs_cnt at any time after the traversal is completed: when the updated value of bfs_cnt is m+1, it means that m nodes have been traversed during the traversal of the shortest path, and jump to step (b5);
[0057] (b5) Update the list node_y: overwrite the values of each element in the list node_y with the values of each element in the list node after the traversal is completed;
[0058] (b6) After completing the test node The subgraph adjacency matrix of After traversing the shortest path, the updated list node_pate is processed: the value of the Dth element in the Dth sublist of the list node_pate is assigned to 1, and the above operation is performed on each sublist; the numbers of the elements with element values 0 in the list node_y are obtained: u1,…,u z ,…;
[0059] For each number obtained, the following processing is performed in turn: For any number u z , calculate the uth in the list node_pate z The number of elements in the row whose value is 1 is num_u z;Through combination logic (node_pate{u z}&node_y)∧node_y to get the shortest path node outside the test node Nearest Node The sampleable node set node_pate{u z}′, where node_pate{u z} represents the uth z Row data of the row, & is the AND operator, ∧ is the XOR operator; from the node The sampleable node set node_pate{u z}′ arbitrarily retain m-num_u z The element value of the element is 1, and the rest of the element values are overwritten from 1 to 0 to obtain the node The second sampleable node set node_pate{u z}″, the second sampleable node set node_pate{u z}″with node_pate{u z} Perform an OR operation to get the node The sampling node set node_pate{u z}″′; Update the list node_out: The number corresponding to each element with a value of 1 in the sampling node set sequentially covers the uth element of the list node_out. z The number of each element in the row;
[0060] Complete the processing of each obtained number and obtain the sampled list node_out;
[0061] (b7) Each row of data in the sampled list node_out is used as a test node The set of sampled neighbor nodes for the test node The subgraph adjacency matrix of Perform adjacency matrix interception operation and obtain The adjacency matrix of the sampling subgraph is obtained, that is, the test node The set of sampled subgraph adjacency matrices
[0062] The present invention also provides a graph neural network interpretation device based on FPGA acceleration, including one or more processors for implementing the above-mentioned graph neural network interpretation method based on FPGA acceleration.
[0063] The present invention also provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, it is used to implement the above-mentioned graph neural network interpretation method based on FPGA acceleration.
[0064] The beneficial effects of the present invention are:
[0065] (1) The present invention uses FPGA hardware to accelerate the interpretation process of graph neural networks, solving the problem of high time complexity in interpreting the classification results of graph neural network nodes. It can solve big data problems in practical application scenarios and expand the practical application scenarios of graph neural network interpretation.
[0066] (2) The present invention preferentially pre-calculates the eigenvalues of all permutations and combinations of node sets, thereby avoiding repeated calculations of permutations and combinations of many nodes due to the mutual adjacency between different test nodes, thereby improving the efficiency of calculations;
[0067] (3) The present invention implements the BFS algorithm after the FPGA hardware improvement. It optimizes the method of calculating the shortest path after obtaining the complete predecessor node after the algorithm is completed to obtain it based on the combinatorial logic operation of the node and the related array. This optimizes the computing and storage requirements and speeds up the generation of interpretation results.
[0068] (4) The present invention is beneficial to the optimization of matrix characteristics for multiplication operations, converting dense matrix multiplication into sparse dense matrix multiplication, using multi-PE parallel processing to optimize resource usage, and greatly improving the performance of graph neural network interpretation acceleration;
[0069] (5) The present invention designs an overall architecture based on FIFO storage and dynamic distribution of computing tasks to run the sub-steps of the above explanation method, thereby optimizing the problem of computing imbalance between nodes. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 is a flow chart of a graph neural network interpretation method based on FPGA acceleration;
[0071] FIG2 is a schematic diagram of FIGG in Example 1;
[0072] FIG3 is an operation flow chart of 8 segmented traversal units;
[0073] Figure 4 shows H q Schematic diagram of the iterative calculation process;
[0074] Figure 5 is a structural diagram of a graph neural network interpretation device based on FPGA acceleration. DETAILED DESCRIPTION
[0075] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to illustrate the present invention, rather than to represent all embodiments. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.
[0076] Example 1: As shown in FIG1 , the present invention provides a graph neural network interpretation method based on FPGA acceleration, comprising the following steps:
[0077] (1) Obtain the target dataset and divide it into a training set and a test set.
[0078] The present invention focuses on node classification tasks, including node classification task scenarios in social networks and citation networks.
[0079] The step (1) is specifically as follows: obtaining data of the node classification task scene and constructing a corresponding graph G = {V, ε}, wherein V represents the graph node set and ε represents the edge set; the graph G contains a total of N nodes, that is, V = {x1, x2, ..., x n ,…,x N}, where x n represents the nth node, n={1,2,…,n,…,N}; then the graph node set V is divided into the training set and test set Among them, the training set T r There are a total of P nodes, p=1,2,…,p,…,P, is the training set T r Any training node in the test set T e There are Q nodes in total, q=1,2,…,q,…,Q, For the test set T e Any test node in; and get the training set T r The corresponding true label set in, For training nodes The true label.
[0080] (2) Use the training set to train the graph neural network to obtain a trained graph neural network; and use the trained graph neural network to classify each test node in the test set to obtain a test prediction label set.
[0081] In this embodiment, GraphSAGE is used as an example of a graph neural network.
[0082] The step (2) specifically includes the following sub-steps:
[0083] (2.1) Training set T r Any training node The set of neighbor nodes in graph G is Training Node Received from neighbor node set Each node Message propagation at each layer in a graph neural network in, l represents the lth layer of the graph neural network, l=1,2,…,l,…,L; Representation node In the feature representation of layer l-1, Representation node The initial feature representation of .
[0084] Every node x in the graph G n Through the edges, the node's own information is propagated to each node connected to it; in this step, the node's own information is diffused to the entire receptive field, which is all neighboring nodes.
[0085] (2.2) Training Node Messages after aggregation at each layer The calculation formula is as follows:
[0086] Wherein, AGG(·) is an aggregation function, which includes summing, averaging or maximum value.
[0087] The aggregated information is transformed through a nonlinear activation function. The original information aggregation process generally involves linear operations and combinations of information. The use of nonlinear functions allows the network to obtain stronger fitting capabilities and enhance the model's expressive power.
[0088] (2.3) Graph neural network for aggregated messages and training nodes Feature representation in the previous layer Perform nonlinear transformation to obtain training nodes Feature representation at layer l The calculation formula is as follows:
[0089] Where Updata(·) is the update function;
[0090] Finally, the training node Final Embedding in Graph Neural Network
[0091] (2.4) For the training set T r Repeat substeps (2.1) to (2.3) for each training node in to obtain the final embedding set Z r :Z r ={z1,z2,…,z p ,…,z P}.
[0092] (2.5) According to Z r ={z1,z2,…,z p ,…,z P}, use the fully connected layer to classify nodes, and use the cross entropy loss function to train and optimize the graph neural network to obtain the trained graph neural network:
[0093] (2.6) Use the trained graph neural network to test the test set T e Classify each test node in and get the test prediction label set Y e : in, For the test set T e Any test node The predicted label is calculated as follows:
[0094] (3) For each test node in the test set, group and parallelly obtain the k-hop subgraph corresponding to each test node in the test set on the FPGA.
[0095] The step (3) specifically includes the following sub-steps:
[0096] (3.1) Obtain the adjacency matrix A and the k-order adjacency matrix A of graph G k ,in, a j,i is any element in the adjacency matrix A, j is element a j,i The first number of element a, i is j,i The second number, a j,i Represents the j-th node x in graph G j With the i-th node x i Are they connected to each other? When the jth node x j With the i-th node x i If they are connected to each other, the element value is 1: a j,i =1, otherwise the element value is 0: a j,i =0; Represents the jth node x j With the i-th node x i Are they connected to each other through at most k-1 intermediate nodes, when the jth node xj With the i-th node x i The element value is 1 when it is connected to each other through at most k-1 intermediate nodes: when When the i-th node x i is the jth node x j k-hop neighbor nodes, otherwise the element value is 0: k-order adjacency matrix A k The jth row is node x j The row, k-order adjacency matrix A k The jth column is node x j Column
[0097] And the adjacency matrix A of graph G and the k-order adjacency matrix A k Write to the DDR4 memory outside the FPGA; and set up a FIFO-cnt counter and 6 pre-processing units;
[0098] (3.2) From the test set T e Select any 6 test nodes and assign a corresponding preprocessing unit to any test node; each preprocessing unit is obtained from the k-order adjacency matrix A of graph G. k Scan the row of the corresponding test node to obtain the row data of the corresponding test node and the number of k-hop neighbor nodes of the corresponding test node;
[0099] Sort the obtained six test nodes in descending order of the number of k-hop neighbor nodes. The test node with a larger number of k-hop neighbor nodes shall be prioritized for step (3.3). If the number of k-hop neighbor nodes is the same, the test node with a smaller sequence number shall be prioritized for step (3.3).
[0100] (3.3) As shown in Figure 3, eight segmented traversal units are set up. The row data obtained by scanning the test node is evenly divided into eight segments of row data and sequentially allocated to the eight segmented traversal units. Each segmented traversal unit has a write initiation request and a write enable response. An arbitrator performs round-robin arbitration control on the requests initiated by the eight segmented traversal units. Each segmented traversal unit traverses the allocated row data and finds the nodes corresponding to the second numbers of all elements with a value of 1 in graph G as the k-hop neighbor nodes of the test node, thereby obtaining the set of k-hop neighbor nodes of the test node.
[0101] And the k-order adjacency matrix A of the graph G is tested by the k-hop neighbor node set of the node kPerform interception: First, intercept the corresponding row data from the k-order adjacency matrix of graph G according to the number of each k-hop neighbor node to obtain the row matrix of the test node; then intercept the corresponding column data from the row matrix of the test node according to the number of each k-hop neighbor node to obtain the subgraph adjacency matrix of the test node;
[0102] (3.4) The six test nodes perform step (3.3) in sequence to obtain the corresponding k-hop neighbor node set and subgraph adjacency matrix;
[0103] (3.5) Then from the test set T e Select any 6 test nodes from the remaining test nodes and repeat steps (3.1) to (3.4) until the test set T e Each test node in the graph obtains its corresponding k-hop neighbor node set and subgraph adjacency matrix, and obtains the subgraph adjacency matrix set in, Represented as a test node The corresponding subgraph adjacency matrix.
[0104] For example, when the graph G is shown in Figure 2, there are a total of 7 nodes in the graph G, that is, the graph node set V: V = {x1, x2, x3, x4, x5, x6, x7}. The graph node set V is divided into the training set and test set
[0105] Taking the graph G shown in FIG2 as an example, in this embodiment, k=2, the step (3) specifically includes the following sub-steps:
[0106] (3.1) Obtain the adjacency matrix A of graph G and the k=2-order adjacency matrix A k=2 .
[0107] (3.2) For the test set T e Select any 6 test nodes in the test set, and assign a corresponding preprocessing unit to any 1 test node.
[0108] The first preprocessing unit obtains the k=2 order adjacency matrix A of the graph G k=2 Scan test node The row where the test node is located Row data: And get the test node k = the number of 2-hop neighbor nodes Similarly, the second preprocessing unit obtains the test node k = the number of 2-hop neighbor nodes The third preprocessing unit obtains the test node k = the number of 2-hop neighbor nodes The fourth preprocessing unit obtains the test node k = the number of 2-hop neighbor nodes The fifth preprocessing unit obtains the test node k = the number of 2-hop neighbor nodes The sixth preprocessing unit obtains the test node k = the number of 2-hop neighbor nodes Sort in descending order, and the order of step (3.3) is:
[0109] (3.3) Set up 8 segmented traversal units to test the nodes Scanned row data The row data is evenly divided into 8 segments and distributed to 8 segment traversal units in turn: the first segment traversal unit is allocated The second segment traversal unit is allocated The third segment traversal unit is allocated The fourth segment traversal unit is allocated The 5th segment traversal unit is allocated The 6th segment traversal unit is allocated The 7th segment traversal unit is allocated The 8th segment traversal unit is allocated [] = []; each segment traversal unit has a write initiation request and a write enable response. An arbitrator performs round-robin arbitration control on the requests initiated by the 8 segment traversal units. Each segment traversal unit traverses the allocated row data and finds the node corresponding to the second number of all elements with a value of 1 in graph G as the k = 2 hop neighbor node of the test node; the test node is obtained. The set of k=2 hop neighbor nodes is: {x1,x2,x3,x4,x5,x6,x7}.
[0110] And pass the test node The k-order adjacency matrix A of the k-hop neighbor node set of graph G k To intercept: First, from the k=2 order adjacency matrix A of graph G k=2 The corresponding row data is intercepted according to the number of each k-hop neighbor node: the numbers of each k-hop neighbor node are 1, 2, 3, 4, 5, 6 and 7, that is, from the k=2 order adjacency matrix A of graph G k=2 Intercept the 1st, 2nd, 3rd, 4th, 5th, 6th, and 7th lines to get the test nodes Row data Then from the test node The row matrix According to the number of each k-hop neighbor node, the corresponding column data is intercepted to obtain the test node The subgraph adjacency matrix of
[0111] (3.4) The remaining 5 test nodes Carry out steps (3.3) in sequence to obtain the corresponding k=2-hop neighbor node set and subgraph adjacency matrix: test node The set of k=2 hop neighbor nodes is {x1,x2,x3,x4,x5,x7} and the subgraph adjacency matrix is Test Node The set of k=2 hop neighbor nodes is {x1,x2,x3,x6,x7} and the subgraph adjacency matrix is Test Node The set of k=2 hop neighbor nodes is {x2,x3,x4,x5} and the subgraph adjacency matrix is Test Node The set of k=2 hop neighbor nodes is {x2,x3,x4,x5} and the subgraph adjacency matrix is Test Node The set of k=2 hop neighbor nodes is {x1,x2,x3,x7} and the subgraph adjacency matrix is
[0112] (3.5) Due to the test set T e There are no remaining test nodes in the subgraph adjacency matrix set
[0113] (4) For each test node in the test set, the corresponding explanation subgraph is obtained through the subgraph adjacency matrix corresponding to each test node.
[0114] The step (4) specifically includes the following sub-steps:
[0115] (4.1) When any test node in the test set The number of k-hop neighbor nodes When it is not greater than m, the test node The subgraph adjacency matrix of Perform the corresponding HN table operations;
[0116] (c1) Obtaining test nodes in FPGA hardware The result set of connectivity relationships of all permutations and combinations of nodes in the k-hop neighbor node set is obtained, and the matrix
[0117] For example, for the number of k-hop neighbor nodes is the subgraph adjacency matrix of m=5 For example,
[0118] The method for generating the matrix P is as follows:
[0119] First, each element in the matrix P (m=5 node permutations) is composed of m segments, where the t-th element is composed of P{t}=[single node result|double node result|triple node result|…|m node result];
[0120] Among them, the single node result can be obtained according to the meaning P(2)=[01000|0…0|0…0|0…0|0], P(3)=[00100|0…0|0…0|0…0|0], P(4)=[00010|0…0|0…0|0…0|0], P(5)=[00001|0…0|0…0|0…0|0];
[0121] By reading the subgraph adjacency matrix The value between the corresponding node permutations and combinations is used to accelerate the double-node result. For example, P(6) corresponds to the two-node permutations and combinations. The connectivity results of For test nodes The first node in the set of k-hop neighbor nodes, For test nodes The second node in the set of k-hop neighbor nodes, if Representation node and nodes In the subgraph adjacency matrix If the two nodes are connected, then P(6) = [00000|10…0|0…0|0…0|0]; if Representation node and nodes In the subgraph adjacency matrix If the two nodes in the middle are not connected, then P(6) = [11000|00…0|0…0|0…0|0]; the connectivity relationship results of the other two-node permutations and combinations are obtained in the same way.
[0122] The connectivity results of the three-node permutation and combination are: P(16), P(17), ... P(25). For example, P(16) corresponds to the three-node permutation and combination Three-node combination Contains 3 two-node permutations: and like The number of 1s in the . node and nodes In the subgraph adjacency matrix The three nodes are connected, so P(16) = [00000|0…0|1…0|0…0|0]; if The number of 1s in is 0, then the node node and nodes In the subgraph adjacency matrix The three nodes are not connected to each other, so P(16)=[11000|0…0|0…0|0…0|0]; if Then only the node and nodes In the subgraph adjacency matrix are interconnected, so P(16)=[10000|0000100000|0…0|0…0|0]; if Then only the node and nodes In the subgraph adjacency matrix are interconnected, so P(16)=[01000|0100000000|0…0|0…0|0]; if Then only the node and nodes In the subgraph adjacency matrix are interconnected, so P(16) = [00100|1000000000|0…0|0…0|0]; the connectivity relationship results of the other three node combinations are obtained in the same way.
[0123] The connectivity relationships of the four-node permutations are: P(26), P(27), ... P(30). For example, P(26) corresponds to the four-node permutation Four-node combination Contains 4 three-node combinations: and like The number of 1s in the . node node and nodes In the subgraph adjacency matrix The four nodes are connected, so P(26) = [00000|0…0|0…0|10000|0]; if The number of 1s in is 0, then the node node node and nodes In the subgraph adjacency matrix The four nodes are not connected to each other, so P(26) = [11110|0…0|0…0|00000|0]; if Then the node node and nodes In the subgraph adjacency matrix The three nodes are connected, so P(26) = [10000|0…0|0000001000|00000|0]; if Then the node node and nodes In the subgraph adjacency matrix The three nodes are connected, so P(26) = [01000|0…0|0001000000|00000|0]; if Then the node node and nodes In the subgraph adjacency matrix The three nodes are connected, so P(26) = [00100|0…0|0100000000|00000|0]; if Then the node node and nodes In the subgraph adjacency matrix The three nodes in the middle are connected, so P(26) = [00010|0…0|1000000000|00000|0]; the connectivity relationship results of the remaining four node combinations are obtained in the same way.
[0124] In the subsequent m+1=5 node process, it is also necessary to count the number of 1s corresponding to the m node permutation and combination results to obtain the corresponding P(31). If the number of 1s is greater than or equal to 2, it indicates that the m+1 nodes are connected, and the element at that permutation position in the corresponding P(31) is 1: P(31)=[00000|0…0|0…0|0…0|1]. Otherwise, the remaining all-0 results indicate that the m nodes are not connected to each other. Otherwise, the corresponding P(31) is obtained according to the position of 1: P(31)=[11111|0…0|0…0|0…0|0].
[0125] (c2) As shown in Figure 4, (H q ) 4 As H q The iterative convergence value of (H q ) 4 The calculation is optimized, using the matrix The idempotent property is and (H q )4 It is divided into three sparse dense matrix multiplications (SPMM) and two dense matrix multiplications (DMM. The sparse dense matrix multiplication assigns tasks to the elements with a value of 1 in the sparse matrix, and adopts a direct static mapping from the matrix row to the 8 PEs to increase the parallel computing of the hardware. There is a gap between the multiple rows assigned to each PE, and the elements with a value of 1 in the middle of the row are searched in a round-robin manner to prevent the sparse matrix from gathering in a few rows. Multiple rows of data to be processed are effectively connected;
[0126] (c3) According to the test node The subgraph adjacency matrix of The node numbers in the pre-calculate the eigenvalue results of all possible node combinations and splice the eigenvalue results into an eigenvalue vector according to the permutations and combinations of single nodes, double nodes, ..., m nodes
[0127] The eigenvalue of any node combination o is calculated using the following formula:
[0128] in, Indicates that all nodes in the node combination arrangement o are predicted to be The probability mean of ; Indicates that all nodes in the graph G are predicted to be The probability mean of ;
[0129] Get the HN value matrix vector Get the HN value matrix vector Center front The values are used as test nodes The set of k-hop neighbor nodes The HN value of the node is obtained The subgraph adjacency matrix of HN table, which is the test node The explanation subgraph of .
[0130] (4.2) When any test node in the test set The number of k-hop neighbor nodes When it is greater than m, the test node The subgraph adjacency matrix of Implement layer-by-layer sampling based on the central node on FPGA to obtain the test node The set of sampled subgraph adjacency matrices
[0131] Then for the sampled subgraph adjacency matrix set Each sampling subgraph adjacency matrix performs the corresponding HN table operation to obtain the HN table corresponding to each sampling subgraph adjacency matrix; test node The subgraph adjacency matrix of The HN value of each node in the adjacency matrix of different sampling subgraphs is the average of the HN values, and the test node is obtained. The explanation subgraph of .
[0132] The sub-step (4.2) specifically includes the following sub-steps:
[0133] (4.2.1) When the test node The number of k-hop neighbor nodes When it is greater than m, for the test node The subgraph adjacency matrix of Test Node For the target node, for the test node The subgraph adjacency matrix of Perform shortest path traversal;
[0134] In step (4.2.1), for the test set T e Test nodes in For example, specifically:
[0135] (a1) In this embodiment, m=5; test node k = the number of 2-hop neighbor nodes is 7, which is greater than m;
[0136] Test Node The number in the graph node set V is q′=2; for the test node The subgraph adjacency matrix of
[0137] (a1) First, perform traversal initialization: initialize the node on the FPGA to have all element values 0 and length Initialize flag to a queue containing only the target node's number in the graph node set V: flag = [flag(1)] = [2], which is used to store the neighbor node labels obtained after each BFS traversal of the node; initialize node_flag to the element value of the qth element to 1, the element value of the remaining elements to 0, and the length to node_flag = [0, 1, 0, 0, 0, 0], this embodiment is for testing nodes That is, q=2, so the element value of the second element in the list node_flag is 1, that is, node_flag(2)=1; initialize node_f to have all element values of -1 and a length of List: node_f = [-1,-1,-1,-1,-1,-1,-1]; initialize bfs_cnt = 1 to control the subscript address of each access to the flag queue; initialize node_pate to contain A list of sublists: Among them, node_pate[E] is the E-th sublist in the list node_pate, and each sublist contains elements and all element values are 0.
[0138] (a2) After the traversal is initialized, the element value of the first element in the list flag is obtained according to bfs_cnt=1: flag(bfs_cnt=1)=2; then the list node is updated: the element value of the flag(bfs_cnt=1)=2 element in the list node is assigned to 1, and the node xflag(bfs_cnt=1)=2=x2 is used as the current traversal target node. The updated list node is node=[0,1,0,0,0,0,0]. At this time, it is judged that the element value of the flag(bfs_cnt=1)=2 element in the list node_f is -1, indicating that the predecessor node of the current traversal target node x2 is invalid. Then the subgraph adjacency matrix is read. The second row of the data is [1 1 1 0 0 0 1], and the process jumps to step (4.2.3) to update the limited node information.
[0139] (a3) Traverse the valid elements of the read row data: the second number of the element with the first element value 1 obtained from the row data is 1: because the second number 1 is different from the element value 2 in the list flag, the list flag is updated: the second number 1 is added to the list flag, after the update: flag = [2,1], and the list node_flag is updated: the element value of the first element in the list node_flag is given 1, after the update: node_flag = [1,1,0,0,0,0,0], and the predecessor node of the node x1 is recorded as the current node. Traverse the target node x2 and update the list node_f: assign the element value of one element in the list node_fl to the number M=2 of the current traversal target node. After updating: node_f=[2,-1,-1,-1,-1,-1,-1], and then read from the row data the second number of the element with the next element value of 1, which is 2: because the second number 2 is the same as the element value 2 in the list flag, read from the row data the second number of the element with the next element value of 1, which is 3: because the second number 3 is different from the element value 2 in the list flag, the list flag is f lag=[2,1] is updated, and after the update: flag=[2,1,3], and the list node_flag is updated, and after the update: node_flag=[1,1,1,0,0,0,0], and the predecessor node of node x3 is recorded as the current traversal target node x2, and the list node_f is updated, and after the update: node_f=[2,-1,2,-1,-1,-1,-1]; then the second number of the element with the next element value of 1 is read from the row data, which is 7: because the second number 7 is different from the element value 2 in the list flag, the column The table flag = [2, 1, 3] is updated to [2, 1, 3, 7]. The list node_flag is also updated to [1, 1, 1, 0, 0, 0, 1]. The predecessor of node x7 is recorded as the current traversal target node x2. The list node_f is also updated to [2, -1, 2, -1, -1, -1, 2]. At this point, no element with a value of 1 can be read from the row data, so the traversal ends. After the traversal ends, the value of bfs_cnt is updated to bfs_cnt = bfs_cnt + 1 = 1 + 1 = 2.
[0140] Repeat step (a2) according to the updated value of bfs_cnt, as follows: according to the updated bfs_cnt=2, access the element value of the second element in the list flag: flag(bfs_cnt=2)=1; then update the list node: assign the element value of the lag(bfs_cnt=2)=1 element in the list node to 1, and use the node xflag(bfs_cnt=2)=1=x1 as the traversal target node. After updating: node=[1,1,0,0,0,0,0]; at this time, it is judged that the list is obtained. The element value of the lag(bfs_cnt=2)=1 element in node_f is 2, indicating that the current traversal target node x1 is valid, and the list node_pate is updated: the path information from the current traversal target node to the target node is stored in the list node_pate, and the path information from the current traversal target node x1 to the target node x2 is node x1→node x2, that is, the element value of the second element in the first sublist in the list node_pate is stored as 1, and the updated list node_pate is node_pate=[[0 1 0 0 0 0 0],[0 0 0 0 0 0 0],[0 0 0 0 0 0 0],[0 0 0 0 0 0 0],[0 0 0 0 0 0 0],[0 0 0 0 0 0 0],[0 0 0 0 0 0 0],[0 0 0 0 0 0 0]]; then read the subgraph adjacency matrix The first row of the data is obtained, and the row data [1 1 1 0 0 1 1] is traversed. The valid elements of the read row data are traversed. After the traversal is completed, the node x6 is obtained, and the updated list flag is flag = [2, 1, 3, 7, 6], the updated list node_flag is node_flag = [1, 1, 1, 0, 0, 1, 1], and the predecessor node of node x6 is recorded as the current traversal target node x1. The updated list node_f is: node_f = [2, -1, 2, -1, -1, 1, 2], and the updated value of bfs_cnt is 3.
[0141] Repeat step (a2) according to the updated value of bfs_cnt, as follows: according to the updated bfs_cnt=3, access the element value of the third element in the list flag: flag(3)=3; then update the list node, and use node x3 as the traversal target node. After updating: node=[1,1,1,0,0,0,0]; at this time, it is judged that the element value of the third element in the list node_f is 2, indicating that the current traversal target node x3 is valid, and the list node_pate is updated: the path information from the current traversal target node to the target node is stored in the list node_pate, and the path information from the current traversal target node x3 to the target node x2 is node x3→node x2, that is, the element value of the second element in the three sublists in the list node_pate is stored as 1; then read the subgraph adjacency matrix The third row of the data is read, and the row data [1 1 1 1 1 0 1] is obtained. The valid elements of the read row data are traversed. After the traversal is completed, nodes x4 and x5 are obtained. The updated list flag is flag = [2, 1, 3, 7, 6, 4, 5], and the updated list node_flag is node_flag = [1, 1, 1, 1, 1, 1]. The predecessor node of node x4 and the predecessor node of node x5 are recorded as the current traversal target node x3. The updated list node_f is: node_f = [2, -1, 2, 3, 3, 1, 2], and the updated value of bfs_cnt is 4.
[0142] Repeat step (a2) according to the updated value of bfs_cnt, as follows: according to the updated bfs_cnt=4, access the element value of the 4th element in the list flag: flag(4)=7; then update the list node, and use node x7 as the traversal target node. After updating: node=[1,1,1,0,0,0,1]; at this time, it is judged that the element value of the 7th element in the list node_f is 2, indicating that the current traversal target node x7 is valid, and the list node_pate is updated: the path information from the current traversal target node to the target node is stored in the list node_pate, and the path information from the current traversal target node x7 to the target node x2 is node x7→node x2, that is, the element value of the 2nd element in the 7th sublist in the list node_pate is stored as 1; then read the subgraph adjacency matrix In line 7, the valid elements of the read row data are traversed; after the traversal is completed, no node is obtained in this traversal, and the list flag, list node_flag, and list node_f are not updated; the value of bfs_cnt after the update is 5;
[0143] Repeat step (a2) according to the updated value of bfs_cnt, as follows: according to the updated bfs_cnt=5, access the element value of the 5th element in the list flag: flag(5)=6; then update the list node, and use node x6 as the traversal target node. After updating: node=[1,1,1,0,0,1,1]; at this time, it is judged that the element value of the 6th element in the list node_f is 1, indicating that the current traversal target node x6 is valid, and the list node_pate is updated: the path information from the current traversal target node to the target node is stored in the list node_pate, and the path information from the current traversal target node x6 to the target node x2 is node x6→node x1→node x2, that is, the element values of the first element in the 6th sublist of the list node_pate and the second element in the 6th sublist are stored as 1; then read the subgraph adjacency matrix In line 6, the valid elements of the read row data are traversed; after the traversal is completed, no node is obtained in this traversal, and the list flag, list node_flag, and list node_f are not updated; the value of bfs_cnt after the update is 6;
[0144] Repeat step (a2) according to the updated value of bfs_cnt, as follows: according to the updated bfs_cnt=6, access the element value of the 6th element in the list flag: flag(6)=4; then update the list node, and use node x4 as the traversal target node. After updating: node=[1,1,1,1,0,1,1]; at this time, it is judged that the element value of the 4th element in the list node_f is 3, indicating that the current traversal target node x4 is valid, and the list node_pate is updated: the path information from the current traversal target node to the target node is stored in the list node_pate, and the path information from the current traversal target node x4 to the target node x2 is node x4→node x3→node x2, that is, the element value of the 3rd element in the 4th sublist and the 2nd element in the 4th sublist in the list node_pate are stored as 1; then read the subgraph adjacency matrix In the fourth line, the valid elements of the read row data are traversed; after the traversal is completed, no node is obtained in this traversal, and the list flag, list node_flag, and list node_f are not updated; the value of bfs_cnt after the update is 7;
[0145] Repeat step (a2) according to the updated value of bfs_cnt, as follows: according to the updated bfs_cnt = 7, access the element value of the 7th element in the list flag: flag (7) = 5; then update the list node, and use node x5 as the traversal target node. After updating: node = [1,1,1,1,1,1,1], at this time, it is judged that the element value of the 5th element in the list node_f is 3, indicating that the current traversal target node x5 is valid, and the list node_pate is updated: the path information from the current traversal target node to the target node is stored in the list node_pate, and the path information from the current traversal target node x5 to the target node x2 is node x5 → node x3 → node x2, that is, the element value of the 3rd element in the 5th sublist and the 2nd element in the 5th sublist in the list node_pate is stored as 1; then read the subgraph adjacency matrix In line 5, the valid elements of the read row data are traversed; after the traversal is completed, no node is obtained in this traversal, and the list flag, list node_flag, and list node_f are not updated; the value of bfs_cnt after the update is 8;
[0146] Because the value of bfs_cnt after this update is When it jumps to the idle state, the test node is completed The subgraph adjacency matrix of Traverse the shortest path and get the updated list node_pate: node_pate = [[0 1 0 0 0 0 0 0],[0 0 0 0 0 0 0 0],[0 1 0 0 0 0 0],[0 1 1 0 0 0 0],[0 1 1 0 0 0 0],[1 1 0 0 0 0 0],[0 1 0 0 0 0 0]].
[0147] (4.2.2) On the test node The subgraph adjacency matrix of While traversing the shortest path, the sampling subgraph node set is generated synchronously to obtain the test node The set of sampled subgraph adjacency matrices
[0148] The sub-step (4.2.2) specifically includes the following sub-steps:
[0149] (b1) First, perform sampling initialization: initialize node_out to contain A list of sublists, each of which contains m = 5 elements and all element values are empty; initialize node_y to have all element values 0 and a length of node_y = [0, 0, 0, 0, 0, 0], where node_y(c) is the value of the cth element in the list node_y.
[0150] (b2) After sampling is initialized, the number of elements in the list flag is detected at any time: when the number of elements in the list flag is detected to be m=5, the distance test node is obtained The set of the nearest m nodes: {x2,x1,x3,x7,x6}, jump to step (b3);
[0151] (b3) Update the list node_out. The updated list node_out is node_out = [[x2,x1,x3,x7,x6],[x2,x1,x3,x7,x6],[x2,x1,x3,x7,x6],[x2,x1,x3,x7,x6],[x2,x1,x3,x7,x6],[x2,x1,x3,x7,x6],[x2,x1,x3,x7,x6]];
[0152] (b4) Check the updated value of bfs_cnt at any time after the traversal is completed: when the updated value of bfs_cnt is m+1=5+1=6, it means that m=5 nodes have been traversed during the traversal of the shortest path, and jump to step (b5);
[0153] (b5) Update the list node_y: overwrite the element values in the list node_y with the element values in the updated list node after the traversal is completed. The updated list node_y is node_y = [1, 1, 1, 0, 0, 1, 1];
[0154] (b6) After completing the test node The subgraph adjacency matrix of After traversing the shortest path, the updated list node_pate is processed: the element value of the D-th element in the D-th sublist of the list node_pate is assigned to 1, and the above operation is performed on each sublist to obtain the processed list node_pate: node_pate=[[1 1 0 0 0 0 0],[0 1 0 0 0 0 0],[0 1 1 0 0 0 0],[0 1 1 1 0 0 0],[0 1 1 0 1 0 0],[1 1 0 0 0 1 0],[0 1 0 0 0 0 1]];
[0155] Get the numbers of the elements in the list node_y whose element value is 0: u1 = 4 and u2 = 5;
[0156] Process the number u1=4 and calculate the number of elements with value 1 in the u1=4 row in the list node_pate, which is num_u1=3. Through the combination logic (node_pate{u1=4}&node_y)∧node_y, we can get the shortest path node outside the test node and the distance from the test node. Nearest Node The sampleable node set node_pate{u1=4}′ is node_pate{u1=4}′=[1,0,0,0,0,1,1]; In the sampleable node set node_pate{u1=4}′, the element value of any m-num_u1=5-3=2 elements is retained as 1, and the rest of the element values are overwritten from 1 to 0, and the node is obtained. The second sampleable node set node_pate{u1=4}″ is node_pate{u1=4}″=[1,0,0,0,0,1,0], node_pate{u1=4}′=[1,0,0,0,0,1,1] is ORed with node_pate{u1=4} to obtain node The sampling node set node_pate{u1=4}″′: node_pate{u1=4}″′=[1,1,1,1,0,1,0]; update the list node_out: add node The number corresponding to each element with a value of 1 in the sampling node set node_pate{u1=4}″′ sequentially covers the number of each element in the u1=4th row of the list node_out. The updated list node_out is node_out=[[x2,x1,x3,x7,x6],[x2,x1,x3,x7,x6],[x2,x1,x3,x7,x6],[x1,x2,x3,x4,x6],[x2,x1,x3,x7,x6],[x2,x1,x3,x7,x6],[x2,x1,x3,x7,x6]];
[0157] Then, the node numbered u2=5 is processed, and the number of elements with element value 1 in the u2=5 row in the list node_pate is calculated to be num_u2=3. The shortest path node outside the test node is obtained through the combination logic (node_pate{u2=5}&node_y)∧node_y. Nearest Node The sampleable node set node_pate{u2=5}′ is node_pate{u2=5}′=[1,0,0,0,0,1,1]; In the sampleable node set node_pate{u2=5}′, the element value of any m-num_u2=5-3=2 elements is retained as 1, and the rest of the element values are overwritten from 1 to 0, and the node is obtained. The second sampleable node set node_pate{u2=5}″ is node_pate{u2=5}″=[1,0,0,0,0,1,0], and node_pate{u2=5}′=[1,0,0,0,0,1,1] is ORed with node_pate{u2=5} to obtain node The sampling node set node_pate{u2=5}″′: node_pate{u2=5}″′=[1,1,1,0,1,1,0];
[0158] Update the list node_out: add node ,x4,x6],[x1,x2,x3,x5,x6],[x2,x1,x3,x7,x6],[x2,x1,x3,x7,x6]], which is the list node_out after sampling.
[0159] (b7) Each row of data in the sampled list node_out is used as a test node The set of sampled neighbor nodes for the test node The subgraph adjacency matrix of Perform adjacency matrix interception operation and obtain The adjacency matrix of the sampling subgraph is obtained, that is, the test node The set of sampled subgraph adjacency matrices
[0160] (4.2.3) Then, the corresponding HN table operation is performed on each sampling subgraph adjacency matrix in the sampling subgraph adjacency matrix set to obtain the HN table corresponding to each sampling subgraph adjacency matrix; test node The subgraph adjacency matrix of The HN value of each k-hop node in is the average of the HN values in the adjacency matrix of different sampling subgraphs, and the test node is obtained. The explanation subgraph of .
[0161] (4.3) According to the number of k-hop neighbor nodes of each test node in the test set, repeat step (4.1) or step (4.2) for each test node to obtain the explanation subgraph corresponding to each test node in the test set.
[0162] Example 2: Corresponding to Example 1 of the aforementioned graph neural network interpretation method based on FPGA acceleration, the present invention also provides an embodiment of a graph neural network interpretation device based on FPGA acceleration.
[0163] Referring to Figure 5, an embodiment of the present invention provides a graph neural network interpretation device based on FPGA acceleration, including one or more processors for implementing a graph neural network interpretation method based on FPGA acceleration in the above embodiment.
[0164] An embodiment of a graph neural network interpretation device based on FPGA acceleration of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware, or by a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located to read the corresponding computer program instructions in the non-volatile memory into the memory and run them. From the hardware level, as shown in Figure 5, it is a hardware structure diagram of any device with data processing capabilities in which the graph neural network interpretation device based on FPGA acceleration of the present invention is located. In addition to the processor, memory, network interface, and non-volatile memory shown in Figure 5, the device with data processing capabilities in the embodiment of the device is usually located according to the actual function of the device with data processing capabilities. Other hardware may also be included, which will not be described in detail.
[0165] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0166] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present invention. A person of ordinary skill in the art can understand and implement the present invention without inventive work.
[0167] An embodiment of the present invention also provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements a graph neural network interpretation method based on FPGA acceleration in the above embodiment. The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), an SD card, a flash card (Flash Card), etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.
[0168] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for interpreting graph neural networks accelerated by FPGA, characterized in that, It includes the following steps: (1) Obtain the target dataset and divide the target dataset into a training set and a test set; (2) Use the training set to train the graph neural network to obtain a trained graph neural network; and use the trained graph neural network to classify each test node in the test set to obtain a set of test prediction labels; (3) For each test node in the test set, implement on the FPGA to obtain the set of k-hop neighbor nodes and the subgraph adjacency matrix corresponding to each test node; (4) For each test node in the test set, obtain the corresponding explanatory subgraph through the subgraph adjacency matrix corresponding to each test node.
2. The method for interpreting graph neural networks accelerated by FPGA according to claim 1, characterized in that, The specific steps of step (1) are as follows: Obtain the data of the node classification task scenario and construct a corresponding graph G. The graph G contains a total of N nodes, that is, the graph node set V = {x1, x2, …, x n , …, x N}, where x n represents the nth node. Then divide the graph node set V into a training set and a test set . Among them, the training set T r contains a total of P nodes, is any node in the training set T r . The test set T e contains a total of Q nodes, is any node in the test set T e . And obtain the true label set corresponding to the training set T r . Among them, is the true label of the node , and is the true label of the node.
3. The method for interpreting graph neural networks accelerated by FPGA according to claim 2, characterized in that, The step (2) specifically includes the following sub-steps: (2.1) Training set T r Any node The set of neighbor nodes in graph G is Node Received from the set of neighbor nodes each node in Messages propagated in each layer of the graph neural network Among them, $l$ represents the $l$-th layer of the graph neural network, where $l = 1, 2, \ldots, l, \ldots, L$; Indicates a node Feature representation at layer l-1 Indicates a node The initial feature representation; (2.2) Node The messages aggregated at each layer The calculation formula is as follows: where AGG(.) is an aggregation function; (2.3) Message aggregated by the graph neural network And node Feature representation in the upper layer Perform a non-linear transformation to obtain a node Feature representation in layer l The calculation formula is as follows: where Update(.) is an update function; Finally, the node Final Embedding in Graph Neural Networks (2.4) For each node in the training set T r Repeat sub-steps (2.1) - sub-step (2.3) to obtain the final embedding set Z r : Z r = {z1, z2, …, z p , …, z P}; (2.5) According to the final embedding set Z r , use a fully connected layer for node classification and use the cross-entropy loss function to train and optimize the graph neural network to obtain a trained graph neural network: (2.6) Classify each node in the test set T using the trained graph neural network to obtain the test prediction label set Y e e : Among them, For the test set T e any test node among The predicted label, the calculation formula is as follows:
4. A method for interpreting a graph neural network accelerated by FPGA according to claim 3, wherein, The step (3) specifically includes the following sub-steps: (3.1) Obtain the adjacency matrix A of graph G and the k-th order adjacency matrix A k , where a j,i is any element in the adjacency matrix A, j is the first number of element a j,i , i is the second number of element a j,i , a j,i represents whether the j-th node x j in graph G is connected to the i-th node x i . When the j-th node x j is connected to the i-th node x i , the element value is 1: a j,i = 1, otherwise the element value is 0: a j,i = 0; Denote the \(j\)-th node \(x\) j and the \(i\)-th node \(x\) i are connected to each other through at most \(k - 1\) intermediate nodes. When the \(j\)-th node \(x\) j and the \(i\)-th node \(x\) i are connected through at most \(k - 1\) intermediate nodes The element value is 1 when interconnected: When When, the i-th node x i is the k-hop neighbor node of the j-th node x j , otherwise the element value is 0: The k-th order adjacency matrix A k The j-th row in j is the row where the node x k is located, and the j-th column in the k-th order adjacency matrix A j is the column where the node x And write the adjacency matrix A of graph G and the k-th order adjacency matrix A k into the DDR4 memory outside the FPGA; and set a FIFO-cnt counter and six preprocessing units; (3.2) Select any 6 test nodes from the test set T e and respectively assign 1 corresponding preprocessing unit to any 1 test node; each preprocessing unit scans the row where the corresponding test node is located from the k-th order adjacency matrix A of the graph G k to obtain the row data of the corresponding test node and obtain the number of k-hop neighbor nodes of the corresponding test node; Sort the number of k-hop neighbor nodes of the obtained 6 test nodes in descending order. The test node with a larger number of k-hop neighbor nodes is preferentially subjected to step (3.3). When the number of k-hop neighbor nodes is the same, the test node with a smaller serial number is preferentially subjected to step (3.3); (3.3) Set 8 segmented traversal units, evenly divide the row data scanned from the test node into 8 segments of row data and assign them to 8 segmented traversal units in turn; each segmented traversal unit has a write initiation application and a write enable response, and a round-robin arbitration control is performed on the applications initiated by 8 segmented traversal units through an arbiter. Each segmented traversal unit traverses the assigned row data, and finds the node corresponding to the second number of all elements with a value of 1 in the graph G as the k-hop neighbor node of the test node, and obtains the set of k-hop neighbor nodes of the test node; And perform an adjacency matrix truncation operation on the k-th order adjacency matrix A of graph G through the set of k-hop neighbor nodes of the test node k The operation is as follows: First, intercept the corresponding row data from the k-th order adjacency matrix of graph G according to the number of each k-hop neighbor node to obtain the row matrix of the test node; then intercept the corresponding column data from the row matrix of the test node according to the number of each k-hop neighbor node to obtain the subgraph adjacency matrix of the test node; (3.4) The 6 test nodes are sequentially subjected to step (3.3) to respectively obtain the corresponding sets of k-hop neighbor nodes and subgraph adjacency matrices; (3.5) Subsequently, select any 6 test nodes from the remaining test nodes in the test set T e and repeat steps (3.1) - (3.4) until each test node in the test set T e obtains its corresponding k-hop neighbor node set and subgraph adjacency matrix, and a subgraph adjacency matrix set is obtained Among them, Denoted as a test node The corresponding subgraph adjacency matrix.
5. A method for interpreting a graph neural network accelerated by FPGA according to claim 4, wherein, The step (4) specifically includes the following sub-steps: (4.1) When any test node in the test set The number of k-hop neighbor nodes When it is not greater than m, for the test node Subgraph Adjacency Matrix Perform operations on the corresponding HN table; (c1) Obtain test nodes in FPGA hardware The set of connection relationship results of all permutations and combinations of nodes in the k-hop neighbor node set of, obtaining a matrix (c2) Take (H q ) 4 as the iterative convergence value of H q and optimize the calculation of (H q ) 4 , where the matrix The idempotent property means that And Divide (H q ) 4 into three sparse-dense matrix multiplications and two dense matrix multiplications. Among them, the sparse-dense matrix multiplication assigns tasks to the elements with a value of 1 in the sparse matrix, and adopts a direct static mapping from the matrix rows to 8 PEs to increase the parallel computing of the hardware. There are intervals between the multiple rows assigned to each PE, and the elements with a value of 1 in the middle of the rows are searched in a round-robin manner to prevent the sparse matrix from aggregating in certain rows. The multiple rows of data to be processed are effectively connected; (c3)According to the test node Sub - graph adjacency matrix The node numbers in [it] pre-calculate the eigenvalue results of all possible permutations of node combinations, and splice the eigenvalue results into an eigenvalue vector according to the permutations and combinations of single nodes, double nodes, …, m nodes The eigenvalue result of any section combination permutation o is calculated by the following formula: Among them, Indicates that all nodes in the section combination arrangement o are predicted as Probability mean value; Indicates that all nodes in graph G are predicted as The probability mean; Obtain the HN value matrix vector Obtain the HN value matrix vector Before The values are used as test nodes respectively in the set of k-hop neighbor nodes The HN value of each node to obtain the test node Sub - graph adjacency matrix The HN table is the test node The explanatory subgraph; (4.2) When any test node in the test set The number of k-hop neighbor nodes When it is greater than m, for the test node Subgraph Adjacency Matrix Implement layer-by-layer sampling based on the central node on the FPGA to obtain test nodes Sampling sub - figure adjacency matrix set Subsequently, corresponding operations on the HN table are performed for each sampled sub-graph adjacency matrix in the set of sampled sub-graph adjacency matrices to obtain the HN table corresponding to each sampled sub-graph adjacency matrix; test node Subgraph Adjacency Matrix Each node in The HN value is the average of the HN values in the adjacency matrices of different sampled subgraphs to obtain the test nodes The explanatory subgraph; (4.3) According to the number of k-hop neighbor nodes of each test node in the test set, repeat step (4.1) or step (4.2) for each test node to obtain the explanatory subgraph corresponding to each test node in the test set.
6. According to the method for interpreting a graph neural network accelerated by FPGA described in claim 4, it is characterized in that, The sub-step (4.2) specifically includes the following sub-steps: (4.2.1) When testing the node The number of k-hop neighbor nodes When it is greater than m, for the test node Subgraph Adjacency Matrix Taking the test node is the target node, for the test node Sub - graph adjacency matrix Perform a traversal of the shortest path; (4.2.2) When testing the nodes Sub - graph Adjacency Matrix While traversing the shortest path, simultaneously generate a sampled subgraph node set to obtain test nodes Set of adjacent matrices of sampling sub - diagrams (4.2.3) Subsequently, corresponding operations on the HN table are performed for each sampled subgraph adjacency matrix in the set of sampled subgraph adjacency matrices to obtain the HN table corresponding to each sampled subgraph adjacency matrix; test nodes Sub - graph adjacency matrix The HN value of each node in is the average of the HN values in the adjacency matrices of different sampled subgraphs, obtaining the test nodes The explanatory subgraph.
7. A method for interpreting a graph neural network accelerated based on FPGA according to claim 6, wherein, The sub-step (4.2.1) specifically includes the following sub-steps: (a1) First, perform traversal initialization: Initialize the node on the FPGA with all elements having a value of 0 and a length of List of: Used to identify whether the node has been visited, where node(c) represents the element value of the c-th element in the list node; initialize flag as a queue containing only the number of the target node in the graph node set V: flag = [flag(1)] = [q'], used to store the neighbor node labels obtained after each traversal of the node by BFS, where flag(1) is the element value of the first element in the list flag, and q' is the target node The number in the set V of graph nodes; initialize node_flag such that the element value of the q-th element is 1, the element values of the remaining elements are 0, and the length is List: node_flag = [node_flag(1), …, node_flag(c), …, node_flag(7)], used to identify whether the node already exists in flag, where node_flag(c) is the element value of the c-th element in the list node_flag, and node_flag(q) = 1; Initialize node_f to have all element values of -1 and a length of List of: Used to identify that the node has been traversed The predecessor node of the process, where node_f(c) is the element value of the c-th element in the list node_f; initialize bfs_cnt = 0 to control the subscript address of the queue flag accessed each time; initialize node_pate to contain List of sub - lists: Among them, node_pate[E] is the E-th sub-list in the list node_pate, and each sub-list contains respectively a elements and the element values are all 0; (a2) After the traversal initialization, access the element value of the bfs_cnt-th element in the list flag according to the value of bfs_cnt: flag(bfs_cnt); then update the list node: assign the element value of the flag(bfs_cnt)-th element in the list node to 1, and set the node x flag(bfs_cnt) as the current traversal target node. At this time, judge whether the element value of the flag(bfs_cnt)-th element in the list node_f is -1: if it is -1, it means that the predecessor node of the current traversal target node is invalid, and then read the subgraph adjacency matrix At the flag(bfs_cnt)-th row, obtain the row data, jump to step (a3), and update the information of the finite nodes; if it is not -1, it indicates that the predecessor node of the currently traversed target node is valid, and update the list node_pate: store the path information from the currently traversed target node to the target node in the list node_pate, and then read the subgraph adjacency matrix The flag(bfs_cnt)-th row of, obtain the row data, jump to step (a3), and update the finite node information; (a3)Traverse the valid elements of the read row data: The second number of the first element with a value of 1 read from the row data is h. If the second number h is the same as the value of any element in the list flag, read the second number of the next element with a value of 1 from the row data; if the second number h is different from the value of any element in the list flag, update the list flag: Add the second number h to the list flag, and the updated list flag is flag = [flag(1), flag(1)] = [M, h], and update the list node_flag: Assign the value 1 to the h-th element in the list node_flag, and record the predecessor node of node x h as the current traversal target node and update the list node_f: Assign the value of the h-th element in the list node_f to the number M of the current traversal target node, then read the second number of the next element with a value of 1 from the row data and repeat the above steps until no element with a value of 1 can be read from the row data, and end the traversal; After the traversal ends, update the value of bfs_cnt: bfs_cnt = bfs_cnt + 1; Repeat step (a2) according to the updated value of bfs_cnt until the updated value of bfs_cnt is When, jump to the idle state and complete the test node Sub - graph adjacency matrix Perform a traversal of the shortest path to obtain the updated list node_pate.
8. A method for interpreting a graph neural network accelerated based on FPGA according to claim 6, wherein, The sub-step (4.2.2) specifically includes the following sub-steps: (b1) First, perform sampling initialization: Initialize node_out to contain List of sublists: Among them, node_out[D] is the D-th sub-list in the list node_out. Each sub-list contains m elements respectively and the element values are all empty; initialize node_y to have all element values of 0 and a length of List of: where node_y(c) is the element value of the c-th element in the list node_y; (b2) After sampling initialization, the number of elements in the list flag is detected at any time: when the number of elements in the table flag is detected to be m, the distance from the test node is obtained Set of the most recent m nodes: {x flag(1) , x flag(2) , … x flag(m)}, go to step (b3); (b3) Update the list node_out, and the updated list node_out is node_out = [[x flag(1) , x flag(2) , …, x flag(m) , …, [x flag(1) , x flag(2) , …, x flag(m) ; (b4) And detect the updated value of bfs_cnt after the traversal ends at any time: When the updated value of bfs_cnt is m + 1, it means that m nodes have been traversed during the traversal of the shortest path, and jump to step (b5); (b5) Update the list node_y: Cover the respective element values in the updated list node after the traversal ends at this time with the respective element values in the list node_y; (b6) After completing the test on the test node Sub-graph adjacency matrix After traversing the shortest path, process the updated list node_pate: assign the element value of the D-th element in the D-th sub-list of the list node_pate to 1, and perform the above operation on each sub-list; obtain the numbers of the elements with element value 0 in the list node_y: u1,…,u z ,…; Perform the following processing on each obtained number in sequence: for any number u z , calculate that the number of elements with a value of 1 in the u z -th row of the list node_pate is num_u z ; obtain the node outside the shortest path node and at a distance from the test node through the combinational logic (node_pate{u z}&node_y)∧node_y Nearest node The set of sampleable nodes node_pate{u z}}′, where node_pate{u z}} represents the row data of the u z -th row in the list node_pate, & is the AND operator, and ∧ is the exclusive OR operator; from the node The set of sampleable nodes node_pate{u z}′, any m-num_u z retained elements have an element value of 1, and the remaining elements are overwritten from 1 to 0 to obtain the node The second set of sampleable nodes node_pate{u z}" and perform an OR operation on the second set of sampleable nodes node_pate{u z}" with node_pate{u z} to obtain the node The set of sampling nodes node_pate{u z}"′; Update the list node_out: The node The numbers corresponding to the elements with element value 1 in each element of the sampling node set successively cover the numbers of each element in the u-th z row of the list node_out; Complete the processing of each obtained number, and obtain the sampled list node_out; (b7) Use each row of data in the sampled list node_out as a test node The set of sampled neighbor nodes for the test node Sub - graph adjacency matrix Perform the adjacency matrix truncation operation to obtain An adjacent matrix of sampling sub - diagrams is obtained, that is, the test node Sampling subgraph adjacency matrix set 9. A graph neural network interpretation device based on FPGA acceleration, characterized in that, Include one or more processors for implementing the FPGA-accelerated graph neural network interpretation method described in any one of claims 1-8.
10. A computer-readable storage medium, on which a program is stored, characterized in that, When the program is executed by the processor, it is used to implement the FPGA-accelerated graph neural network interpretation method described in any one of claims 1-8.
Citation Information
Patent Citations
Graph neural network interpretation method and system, terminal and storage medium
CN114399025A
Semantic guidance-based semi-supervised node classification method of multilayer structure
CN115935256A
Gene regulatory network construction method and system based on graph neural network
CN116129992A
Accelerator, computer system and method
CN117223005A
Explanation of graph-based predictions using network motif analysis
US11228505B1
Cited By
Sampling-based graph neural network acceleration method and device
CN120745691A
Graph neural network node classification-oriented progressive interpretation method
CN121093101A