Graph convolutional neural network aggregation calculation acceleration device and method

By CSB encoding compression and optimized feature reuse of graph data of graph convolutional neural networks, and using parallel computing and reducing memory access conflicts, the problem of low computational efficiency in the aggregation stage of graph convolutional neural networks is solved, and efficient computing and load balancing is achieved.

CN120012824APending Publication Date: 2025-05-16HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510102515.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-15
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The computational efficiency of graph convolutional neural networks in the aggregation stage is difficult to perform efficiently due to the sparseness of graph data, irregular memory access mode and unbalanced computing load.

Method used

A graph convolutional neural network aggregation computing acceleration device is designed to improve computing efficiency by CSB encoding compression of graph data, optimize feature reuse, and adopt parallel computing and reducing memory access conflicts.

Benefits of technology

It significantly improves the computing efficiency in the aggregation stage, reduces memory access conflicts, optimizes load balancing and parallel computing, improves data reuse, and reduces power consumption and energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012824A_ABST
    Figure CN120012824A_ABST
Patent Text Reader

Abstract

The invention discloses a graph convolutional neural network aggregation calculation acceleration device and method, and relates to the field of graph neural network hardware acceleration, the device compresses a graph adjacency matrix through CSB coding, and source node features and the compressed graph adjacency matrix are stored in an off-chip storage module together; the aggregation control unit loads data from the off-chip storage module and schedules the data; the edge prefetcher controls loading of edge data according to the block offset information, and the data distribution module transmits the data to the calculation unit; the aggregation calculation unit performs calculation according to the source node feature and the boundary value, updates the target node feature and writes a result back to the storage module; the method comprises the steps of compression of graph data, data loading, edge prefetching, data distribution, feature calculation, destination node feature updating and the like so as to improve the calculation efficiency of a graph convolutional neural network aggregation stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of graph neural network hardware acceleration, and in particular to a graph convolutional neural network aggregation computing acceleration device and method. Background Art

[0002] With the emergence of big data and the rapid development of computing resources, graph data has been widely used due to its powerful expression and adaptability. Due to the particularity of graph structure, deep learning technology, which has brought profound changes in various application fields, cannot be directly applied to the field of graph data, and graph neural networks came into being. Graph neural networks enable machine learning to be applied to graph structures in non-Euclidean spaces. As one of the most widely used graph neural network models, graph convolutional neural networks (GCNs) play an indispensable role in recommendation systems, computer vision, natural language processing, real-time prediction, financial risk control and other fields, and have achieved remarkable results.

[0003] The execution process of graph convolutional neural networks is divided into two stages, namely aggregation and combination. The aggregation stage is an irregular sparse operation, while the combination stage is a regular dense operation. The main bottleneck of current graph convolutional neural network hardware acceleration is concentrated in the aggregation stage. Due to the extremely sparse characteristics of graph data, irregular memory access and computing patterns make it difficult to efficiently reuse graph data; the highly mismatched in-degree of each destination node makes it difficult to ensure load balancing between computing units in the aggregation stage. How to efficiently execute the aggregation stage is the key to graph convolutional neural network hardware acceleration.

[0004] When using general-purpose computing devices such as CPUs and GPUs to perform graph aggregation calculations, irregular memory access patterns will result in extremely low hit rates for the CPU's L2 and L3 caches. The GPU's SIMT (single-instruction, multiple-thread) execution model will also face a large number of memory access ambiguities, with different threads needing to access the same or similar memory addresses, leading to memory access conflicts and delays. Summary of the invention

[0005] The purpose of this invention is to design a hardware acceleration method to process and calculate graph data to meet the computing requirements of the graph convolutional neural network aggregation stage.

[0006] The technical solution of the present invention is: to provide a graph convolutional neural network aggregation computing acceleration device, the device includes: a front module, an off-chip storage module, an aggregation control unit, a source node feature storage module, a block offset storage module, an edge storage module, an edge prefetcher, a data distribution module, a source node feature navigation module, a destination node feature storage module, and an aggregation computing unit;

[0007] The front-end module receives the graph data and completes the combination operation, calculates all the source node features in the graph data, and compresses the graph adjacency matrix in the graph data by CSB coding. The graph adjacency matrix is ​​divided into blocks according to the source node index, and a fixed number of source node features are stored in each block. The calculated source node features are transmitted together with the compressed graph adjacency matrix to the off-chip storage module for storage;

[0008] The aggregation control unit connects the off-chip storage module, the source node feature storage module, the block offset storage module, the edge storage module, and the destination node feature storage module to coordinate the data transmission between the modules and control the calculation process;

[0009] The source node feature storage module temporarily stores the source node features corresponding to multiple blocks; the block offset storage module stores the block offset information of the graph adjacency matrix; the edge storage module stores the edge information, including the source node index, the destination node index and the edge value; the edge prefetcher controls the loading and calculation start and stop of the edge data in the edge storage module according to the block offset information in the block offset storage module, ensuring that the edge data is transmitted to the data distribution module at the right time;

[0010] The data distribution module receives the edge data provided by the edge storage module, passes the source node index to the source node feature navigation module, passes the destination node index to the destination node feature storage module, and transmits the edge value to the aggregation calculation unit;

[0011] The source node feature navigation module reads the corresponding source node feature from the source node feature storage module according to the source node index and passes it to the aggregation calculation unit;

[0012] The aggregation calculation unit connects the source node feature navigation module and the destination node feature storage module, multiplies the source node feature by the edge value and adds the result to the destination node feature, updates the destination node feature and writes it back to the destination node feature storage module;

[0013] The destination node feature storage module is connected to the data distribution module and the aggregation calculation unit, stores the updated destination node features, and provides the destination node features to the aggregation calculation unit during the calculation process.

[0014] The present invention also provides an aggregation computing acceleration method based on the above-mentioned graph convolutional neural network aggregation computing acceleration device, the method comprising:

[0015] S1. Perform CSB coding compression on the graph data of the graph convolutional neural network, divide the graph adjacency matrix into blocks according to the source node index, store a fixed number of source node features in each block, retain only the source node index, destination node index and edge value of non-zero edges, and record the block offset information;

[0016] S2, writing the source node features in step S1 and the graph adjacency matrix data compressed by CSB coding into an off-chip storage module as data input for the aggregation stage;

[0017] S3, the aggregation control unit loads the source node feature of each block from the off-chip storage module to the source node feature storage module, and loads the block offset information and edge information to the block offset storage module and the edge storage module respectively;

[0018] S4, the edge prefetcher controls the loading of edge data and the start and stop of operations according to the block offset information in the block offset storage module, and transmits the loaded edge data to the data distribution module in sequence;

[0019] S5, the data distribution module sends the source node index of each edge to the source node feature navigation module, sends the destination node index to the destination node feature storage module, and transmits the edge value to the aggregation calculation unit;

[0020] S6, the source node feature navigation module outputs the corresponding source node feature to the aggregation calculation unit according to the input source node index, and the destination node feature storage module outputs the corresponding destination node feature to the aggregation calculation unit according to the destination node index;

[0021] S7, the aggregation calculation unit multiplies the source node feature and the edge value, and adds the product result to the destination node feature, generates an updated destination node feature and writes it back to the destination node feature storage module;

[0022] S8. After completing the data calculation of all edges in sequence, the aggregation control unit updates the source node features of the next block to the source node feature navigation module, and completes the aggregation operations of all blocks in sequence according to the above steps.

[0023] In any of the above technical solutions, further, the compression step in step S1 includes:

[0024] The graph adjacency matrix is ​​divided into multiple blocks according to the source node index, and each block contains a fixed number of source node features; only non-zero elements are retained in each block, and their values ​​are stored together with the corresponding source node and destination node indexes; the number of edges in each block is recorded through block offset information, reducing storage space.

[0025] In any of the above technical solutions, further, the data operation relationship of each layer of the graph convolutional neural network algorithm is:

[0026]

[0027] in, is an adjacency matrix with self-loops, X l is the output feature of the current layer, X l-1is the output feature of the previous layer, that is, the input feature of the current layer, W l is the input weight of the current layer, δ represents the activation function ReLU, and the classic two-layer graph convolutional neural network operation is expressed as:

[0028]

[0029] Among them, the main operation is divided into two parts. The combination operation represents the process of multiplying the node features by the weights to achieve the dimension reduction of the features, namely:

[0030]

[0031] The aggregation operation is the process of multiplying the graph adjacency matrix and the feature matrix. This part is a sparse irregular operation, which can be expressed as:

[0032]

[0033] For X l The combinatorial operation part, For X l The aggregation operation part.

[0034] In any of the above technical solutions, further, the edge pre-fetching step in step S3 includes:

[0035] The edge prefetcher controls loading edge data from the edge storage module according to the block offset information in the block offset storage module, and transmits the loaded edge data to the data distribution module;

[0036] The edge prefetcher ensures that only the edge data required by the current block is loaded each time, thereby reducing the loading and latency of redundant data.

[0037] In any of the above technical solutions, further, the data allocation step in step S4 includes:

[0038] The data allocation module splits the loaded source node feature data and edge data, and transmits the source node features to the source node feature navigation module according to the source node index, and transmits the destination node features to the destination node feature storage module according to the destination node index; the edge value is passed to the aggregation calculation unit for feature calculation.

[0039] In any of the above technical solutions, further, the feature navigation step in step S5 includes: source node feature navigation and destination node feature navigation:

[0040] The source node feature navigation module reads the corresponding source node feature from the source node feature storage module according to the input source node index, and transmits it to the aggregation calculation unit;

[0041] The destination node feature storage module reads the corresponding destination node feature according to the destination node index and transmits it to the aggregation calculation unit.

[0042] In any of the above technical solutions, further, the aggregation calculation step in step S6 includes:

[0043] The aggregation calculation unit multiplies the source node feature with the edge value to generate a temporary feature; the temporary feature is accumulated to the destination node feature through an addition operation to obtain an updated destination node feature; the updated destination node feature is written back to the destination node feature storage module for subsequent calculations.

[0044] In any of the above technical solutions, further, the feature updating step in step S7 includes:

[0045] After each block calculation is completed, the aggregation calculation unit adds the product of the source node feature and the edge value to the destination node feature to update the feature of the node; the updated destination node feature is written back to the destination node feature storage module to prepare for the next round of calculation.

[0046] In any of the above technical solutions, further, the iterative calculation step in step S8 includes:

[0047] According to the description in step S6 and step S7, aggregate calculation is performed on each block in turn until the edge data of all blocks are calculated;

[0048] After all block calculations are completed, the destination node feature storage module aggregates the stored destination node features to form the final calculation result.

[0049] The beneficial effects of the present invention are:

[0050] The technical solution in the present invention mainly targets the computational bottlenecks in the aggregation stage of graph convolutional neural networks, especially the performance problems caused by the sparsity of graph data, irregular memory access, and unbalanced computing load. By compressing graph data, optimizing feature reuse, adopting parallel computing, and reducing memory access conflicts, the present invention significantly improves the computational efficiency of the aggregation stage, and has the following beneficial effects:

[0051] Improved computing efficiency: The present invention compresses the graph adjacency matrix by using the CSB (Compressed Sparse Blocks) encoding format, retains the non-zero elements in each block, greatly reduces memory usage and reduces the calculation of zero elements; this compression method effectively reduces the demand for storage bandwidth and computing bandwidth, making data reading and processing more efficient; by optimizing the storage of graph data, the access to invalid data during the calculation process is greatly reduced, thereby improving the overall computing efficiency.

[0052] Reduced memory access conflicts: Traditional graph convolutional neural networks often face high memory access conflicts during the aggregation stage due to the sparsity and irregularity of graph data, especially in multi-threaded computing. The index-driven MAC structure proposed in the present invention avoids repeated memory reading when multiple processing units (PEs) access the same destination node feature, thereby reducing the possibility of memory access conflicts. By temporarily storing destination node features and dynamically adjusting the access order, the present invention effectively avoids data conflicts, thereby reducing access delays.

[0053] Optimized load balancing and parallel computing: The present invention ensures load balancing between computing units by parallelizing the computing process. Each processing unit independently processes different edge data and performs multiplication calculations of source node features and edge values. Parallel computing not only improves the computing speed, but also avoids performance degradation due to unbalanced computing load. In addition, by designing an efficient data allocation and navigation mechanism, the transmission of source node features and edge values ​​between different modules is optimized, further enhancing the parallelism and efficiency of the calculation.

[0054] Improved data reuse rate: In the aggregation stage of graph convolutional neural networks, due to the high sparsity and irregularity of computational data, traditional methods are difficult to effectively reuse loaded data, resulting in frequent memory access and computational redundancy; the present invention greatly improves data reuse rate through the collaboration of the source node feature navigation module and the destination node feature storage module; when multiple processing units access the same source node feature, the source node feature only needs to be loaded once and is shared by multiple processing units. This strategy effectively reduces the consumption of memory bandwidth and improves computational efficiency.

[0055] Reduced power consumption and energy efficiency: By optimizing the memory access mode during the calculation process and reducing redundant calculations, the present invention reduces the waste of hardware resources, especially in a large-scale parallel computing environment, by increasing the memory reuse rate and reducing memory access conflicts, the power consumption of the system is reduced; combined with efficient hardware implementation, this method not only improves the computing efficiency, but also improves the energy efficiency, and meets the requirements of modern high-performance computing equipment for energy efficiency optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The advantages of the above and additional aspects of the present invention will become apparent and easily understood in the description of the embodiments in conjunction with the following drawings, in which:

[0057] Figure 1 It is a block diagram of a graph convolutional neural network aggregation computing device according to a graph convolutional neural network aggregation computing acceleration device and method according to an embodiment of the present invention;

[0058] Figure 2This is an example diagram of a self-loop adjacency matrix according to a graph convolutional neural network aggregation computing acceleration device and method according to an embodiment of the present invention;

[0059] Figure 3 It is a schematic diagram of the adjacency matrix compression CSB encoding format of a graph convolutional neural network aggregation computing acceleration device and method according to an embodiment of the present invention;

[0060] Figure 4 It is an example diagram of a graph convolutional neural network aggregation process of a graph convolutional neural network aggregation computing acceleration device and method according to an embodiment of the present invention;

[0061] Figure 5 It is an index-driven MAC structure diagram of a graph convolutional neural network aggregation computing acceleration device and method according to an embodiment of the present invention;

[0062] Figure 6 This is an example diagram of memory access rules of a graph convolutional neural network aggregation computing acceleration device and method according to an embodiment of the present invention;

[0063] Figure 7 It is a target node characteristic part and a distributed storage structure of a graph convolutional neural network aggregation computing acceleration device and method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0064] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.

[0065] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.

[0066] like Figure 1 As shown, this embodiment provides a graph convolutional neural network aggregation computing acceleration device and method, which aims to solve the efficiency bottleneck caused by problems such as sparse graph data and irregular access patterns when the graph convolutional neural network executes the aggregation computing stage. Specifically, by proposing a hardware acceleration method, the execution efficiency of the aggregation computing stage is improved through data compression, source node feature reuse, parallel processing and other means.

[0067] The computation process of graph convolutional neural networks is usually divided into two main stages: aggregation and combination. The combination stage is the process of multiplying node features with weight matrices, the purpose of which is to reduce the dimension of each node's features, which is usually achieved through matrix multiplication with the weight matrix. The aggregation stage performs the product operation of the graph adjacency matrix and the feature matrix, which mainly completes the aggregation of neighbor node features. The sparsity of graph data causes the aggregation stage to be an irregular sparse operation, so the computational efficiency is low. Current hardware acceleration devices (such as CPUs and GPUs) are difficult to efficiently perform such sparse computations with irregular memory access, mainly because of mismatched loads, irregular memory access patterns, and a large number of memory access conflicts. The technical solution of the present invention is to provide a new hardware acceleration method for the sparse characteristics of the aggregation computation stage, aiming to increase the reuse of feature data, improve computational efficiency, and reduce potential memory access conflicts by compressing and processing graph data.

[0068] A graph convolutional neural network aggregation computing acceleration device provided in this embodiment includes: a front-end module, an off-chip storage module, an aggregation control unit, a source node feature storage module, a block offset storage module, an edge storage module, an edge prefetcher, a data allocation module, a source node feature navigation module, a destination node feature storage module, and an aggregation computing unit.

[0069] The front-end module receives graph data and completes the combination operation to generate source node features, and transmits the source node features together with the graph adjacency matrix to the off-chip storage module.

[0070] The off-chip storage module stores the graph adjacency matrix compressed by CSB coding and the source node features output by the front module. The off-chip storage module is connected to the aggregation control unit and is responsible for providing data to the aggregation control unit.

[0071] The aggregation control unit connects the off-chip storage module, the source node feature storage module, the block offset storage module, the edge storage module, and the destination node feature storage module to coordinate data transmission and computing process control among the modules.

[0072] The source node feature storage module temporarily stores source node features corresponding to multiple blocks. The module is connected to the aggregation control unit to ensure that the source node features can be efficiently transmitted to the aggregation calculation unit during the calculation process.

[0073] The block offset storage module stores the block offset information of the graph adjacency matrix, connects to the aggregation control unit, and helps control the number of edges read by the computing unit in each block.

[0074] The edge storage module stores edge information, including source node index, destination node index and edge value. The module is connected to the aggregation control unit and is used to provide edge data to the computing unit.

[0075] The edge prefetcher controls the loading and calculation start and stop of edge data in the edge storage module according to the block offset information in the block offset storage module, ensuring that the edge data is transmitted to the data distribution module at the right time.

[0076] The data distribution module receives the edge data provided by the edge storage module, passes the source node index to the source node feature navigation module, passes the destination node index to the destination node feature storage module, and transmits the edge value to the aggregation calculation unit.

[0077] The source node feature navigation module reads the corresponding source node features from the source node feature storage module according to the source node index and passes them to the aggregation calculation unit.

[0078] The destination node feature storage module is connected to the data distribution module and the aggregation calculation unit, stores the updated destination node features, and provides the destination node features to the aggregation calculation unit during the calculation process.

[0079] The aggregation calculation unit connects the source node feature navigation module and the destination node feature storage module, multiplies the source node feature with the edge value and adds the result to the destination node feature, updates the destination node feature and writes it back to the destination node feature storage module.

[0080] The aggregation computing acceleration method based on the above device includes:

[0081] The graph data of the graph convolutional neural network is compressed by CSB coding. The graph adjacency matrix is ​​divided into blocks according to the source node index. A fixed number of source node features are stored in each block. Only the source node index, destination node index and edge value of non-zero edges are retained, and the block offset information is recorded.

[0082] In the aggregation stage of graph convolutional neural networks, the input graph data is usually represented in the form of a graph adjacency matrix. Due to the sparsity of graph data, the original adjacency matrix contains a large number of zero elements, which leads to a waste of memory and computation. Data compression processing is required. This paper proposes a CSB compression encoding method to assist the computing engine in completing efficient aggregation processing. Figure 2 As shown in the figure, the graph adjacency matrix contains 8 nodes, indicating the connection relationship between nodes, where the horizontal axis is the source node index and the vertical axis is the destination node index. The graph data is divided according to the source node index, and each block corresponds to a fixed number of source node features.

[0083] by Figure 2 The adjacency matrix in For example, let A1 be The matrix A2 is obtained by setting the elements in Block2, Block3, and Block4 to 0 and keeping the elements in Block1 unchanged. The matrix A3 is obtained by setting the elements in Block1, Block3, and Block4 to 0 and keeping the elements in Block2 unchanged. The matrix A4 is obtained by setting the elements in Block1, Block2, and Block4 to 0 and keeping the elements in Block3 unchanged. The matrix obtained by setting the elements in Block1, Block2, and Block3 to 0 and keeping the elements in Block4 unchanged is:

[0084]

[0085] The aggregation process can be expressed as:

[0086]

[0087] That is, it is divided into four operations, namely and To execute Block1, For example, during the calculation, the valid values ​​in A1 are the first two columns, and the remaining columns are all 0 elements and do not participate in the calculation, so the corresponding source node feature Only the first two rows of elements are involved in the calculation. When calculating Block1, only the first two rows of features need to be loaded, thereby achieving efficient reuse of source node features during the calculation process. The calculation process of Block2, Block3, and Block4 is consistent with Block1.

[0088] After completing Block1 operation, the target node characteristics are obtained Block2 After Block 2 is calculated, the destination node feature is obtained as follows: Block3 After Block 3 is calculated, the destination node feature is obtained as follows: Block4 After Block 4 is calculated, the destination node feature is obtained as follows: The result of the current block is added to the result of the previous block operation to achieve the reuse of the destination node features between blocks.

[0089] Figure 2 The adjacency matrix is ​​compressed in the direction of the dotted arrows in the block, and the corresponding CSB encoding format is obtained as follows: Figure 3 As shown, it is divided into two parts. One part is the block offset, which is used to parse the number of edges in the current block. Figure 2The white squares in the middle block represent 0 elements, and the colored squares represent non-zero elements, which represent the values ​​of the edges. Only non-zero values ​​are retained during the compression process. The block offset is used to parse the number of edges inside each block. The first element of the block offset is the initialization value 0. The second value minus the first value gets the number of edges in Block1. Therefore, the second value of the block offset is the first value 0 plus the number of edges in Block1 4, which equals 4. The third value of the block offset is the second value 4 plus the number of edges in Block2 5, which equals 9. The fourth value of the block offset is the third value 9 plus the number of edges in Block3 5, which equals 14. The fifth value of the block offset is the fourth value 14 plus the number of edges in Block4 6, which equals 20. The other part is the edge information, represented by (src, des, value (src,des) ), where src is the source node index, des is the destination node index, and value (src,des) Indicates the value of the edge. The compression process does not process the value of the edge. It depends on the value of the non-zero element in the initial adjacency matrix. Use value (src,des) Description. Figure 3 The source node index, destination node index, and edge value in the block represent the edge information. Each column of three squares represents an edge information. The edge information is arranged in the order of Block1, Block2, Block3, and Block4. In each block, the direction of the dotted arrow corresponds to the order of edge compression, and the source node index in the block is redistributed in the horizontal direction. The source node index corresponding to the element in the first column of the block is 1, and the source node index of the element in the second column of the block is 2. The destination node index in the block is consistent with the destination node index corresponding to the adjacency matrix. The edge value is loaded according to the edge value in the adjacency matrix, with value (src,des) Indicates the value of the edge. Taking Block1 as an example, the source node index of the first edge in the block is 1, the destination node index is 1, and the edge value is value (1,1) , the first edge is encoded as (1,1,value (1,1) ), the second edge has a source node index of 2 in the block, a destination node index of 2, and an edge value of value (2,2) , the second edge is encoded as (2,2,value (2,2) ), the third edge has a source node index of 1 in the block, a destination node index of 4, and an edge value of value (1,4) , the third edge is encoded as (1,4,value (1,4) ), the fourth edge has a source node index of 2 in the block, a destination node index of 7, and an edge value of value (2,7) , the fourth edge is encoded as (2,7,value (2,7) ), and the edges in Block2, Block3, and Block4 are compressed and encoded in the same way.

[0090] The features output by the combination operation are used as source node features and written into the off-chip storage module together with the graph adjacency matrix data compressed by CSB encoding as data input in the aggregation stage.

[0091] In the graph convolutional neural network, the output feature data of the combination phase will be used as the source node features and as the input data for the aggregation calculation. In addition, the compressed adjacency matrix also needs to be transmitted to the hardware computing device as input data. Specifically, the output result of the combination operation (i.e., the source node features) will be written to the off-chip storage module together with the compressed graph adjacency matrix, and prepared for subsequent aggregation calculations.

[0092] The data operation relationship of each layer of the graph convolutional neural network algorithm is:

[0093]

[0094] in, is an adjacency matrix with self-loops, X l is the output feature of the current layer, X l-1 is the output feature of the previous layer, that is, the input feature of the current layer, W l is the input weight of the current layer, δ represents the activation function ReLU, and the classic two-layer graph convolutional neural network operation can be expressed as:

[0095]

[0096] Among them, the main operation is divided into two parts. The combination operation represents the process of multiplying the node features by the weights to achieve the dimension reduction of the features, namely:

[0097]

[0098] The aggregation operation is the process of multiplying the graph adjacency matrix and the feature matrix. This part is a sparse irregular operation and can be expressed as:

[0099]

[0100] For X l The combinatorial operation part, For X l The aggregation operation part.

[0101] One input of the graph convolutional neural network aggregation computing device in the present invention is the output result of the combination operation, which is used as the source node feature data in the aggregation stage, and the other input is the graph adjacency matrix.

[0102] The aggregation control unit loads the source node feature of each block from the off-chip storage module to the source node feature storage module, and loads the block offset information and edge information to the block offset storage module and the edge storage module respectively.

[0103] In the hardware accelerator, the aggregation control unit is responsible for coordinating data loading and processing. First, it loads the source node feature data from the off-chip storage module, which has been calculated and stored in the combination stage. Then, the aggregation control unit loads the corresponding block offset information and edge data into the block offset storage module and edge storage module respectively, preparing for the subsequent edge data loading.

[0104] The edge prefetcher controls the loading of edge data and the start and stop of operations according to the block offset information in the block offset storage module, and transmits the loaded edge data to the data distribution module in sequence.

[0105] The main function of the edge prefetcher is to load edge data from the edge storage module one by one according to the block offset information provided by the block offset storage module. The edge prefetcher will control the start and stop of the loading process to ensure that the data of each edge can be sent to the data distribution module at the right time to prepare for subsequent calculations.

[0106] The data distribution module sends the source node index of each edge to the source node feature navigation module, sends the destination node index to the destination node feature storage module, and transmits the edge value to the aggregation calculation unit.

[0107] The data distribution module is responsible for splitting the edge data and assigning the source node index, destination node index and edge value to different modules. The source node index will be transmitted to the source node feature navigation module to extract the source node feature during calculation. The destination node index is transmitted to the destination node feature storage module to update the destination node feature. The edge value will be directly sent to the aggregation calculation unit for calculation.

[0108] The source node feature navigation module outputs the corresponding source node feature to the aggregation calculation unit according to the input source node index, and the destination node feature storage module outputs the corresponding destination node feature to the aggregation calculation unit according to the destination node index.

[0109] The source node feature navigation module is responsible for loading the corresponding source node features from the source node feature storage module according to the input source node index and passing it to the aggregation calculation unit. At the same time, the destination node feature storage module outputs the corresponding destination node features according to the input destination node index so as to participate in the calculation.

[0110] The aggregation calculation unit multiplies the source node feature with the edge value, and adds the product result to the destination node feature, generates an updated destination node feature and writes it back to the destination node feature storage module.

[0111] After receiving the source node feature and edge value, the aggregation calculation unit performs a multiplication operation: the source node feature is multiplied by the edge value, and the product result is accumulated to the destination node feature. The updated destination node feature is written back to the destination node feature storage module to provide input for subsequent calculations.

[0112] After completing the data calculation of all edges in sequence, the aggregation control unit updates the source node features of the next block to the source node feature navigation module, and completes the aggregation operations of all blocks in sequence according to the above steps.

[0113] After each block is calculated, the aggregation control unit will update the calculation results to the source node feature navigation module and start the calculation of the next block. This process will continue until the calculation of all blocks is completed, and finally the complete destination node feature is obtained.

[0114] To organize the above solutions:

[0115] First, load the data from the front-end module into the off-chip storage to complete the initialization. This includes the CSB compression code and source node features. Figure 4 for Figure 2 The flowchart of the graph data aggregation process completes the aggregation in the order of blocks, and controls the start and stop of the operation process through the edge data. When the edge data is sent to the aggregation calculation unit, the aggregation calculation unit immediately starts the data operation. The corresponding source node features in Block1 are Feature1 and Feature2, the corresponding source node features in Block2 are Feature3 and Feature4, the corresponding source node features in Block3 are Feature5 and Feature6, and the corresponding source node features in Block4 are Feature7 and Feature8. The destination node features corresponding to the current block are superimposed on the destination node features obtained by the aggregation of the previous block.

[0116] Before starting the calculation in Block1, the initial destination node features are all 0. Before the aggregation operation is started, the aggregation control unit loads the source node features corresponding to Block1, Block2, Block3, and Block4 from the off-chip storage module to the source node feature storage module, loads the block offset and edge data in the CSB compression code from the off-chip storage module to the block offset module and edge storage module respectively, and loads Fearure1 and Feature2 from the source node feature storage module to the source node feature navigation module.

[0117] The edge prefetcher reads the block offset value 4 from the block offset storage module and subtracts it from the initialization block offset 0 to obtain the number of edges in Block1, which is 4-0=0. When performing aggregation operations on Block1, the edge prefetcher fetches (1,1,value (1,1) )、(2,2,value(2,2) )、(1,4,value (1,4) )、(2,7,value (2,7) )4 edges. To execute the edge (1,1,value (1,1) ) as an example, take out the edge (1,1,value (1,1) ) is sent to the data distribution module, which sends the source node index 1 to the source node feature navigation module, sends the destination node index 1 to the destination node feature storage module, and sends the edge value (1,1) The source node feature navigation module takes out the source node feature Feature1 according to the source node feature index 1 and sends it to the PE of the aggregation operation module. The destination node feature storage module sends the destination node feature Feature1' to the PE of the aggregation operation module according to the destination node feature index 1. At this time, the internal PE of the aggregation operation module executes the edge (1,1,value (1,1) ), which is to multiply the source node feature Feature1 by the value of the upper edge. (1,1) And add the product to the destination node feature Feature1' to get the new destination node feature Feature1':

[0118] Feature 1′ out =Feature1′ in +Feature1×value (1,1) ;

[0119] After the operation is completed, the new destination node feature Feature1' is written back to the destination node feature storage module. (1,1) ) is completed, and the edge (2,2,value) is processed in the same way. (2,2) )、(1,4,value (1,4) )、(2,7,value (2,7) ) until all four edges in Block 1 are aggregated, and the edge prefetcher stops fetching data from the edge storage module.

[0120] Next, start the aggregation calculation of Block2. Before the calculation, Feature3 and Feature4 are loaded from the source node feature storage module to the source node feature navigation module. At this time, the edge prefetcher reads the block offset 9 from the block offset module, and the number of edges in Block2 is 9-4=5. The aggregation calculation unit will take out 5 edges from the edge storage module to perform the aggregation calculation of Block2. The specific calculation process is the same as Block1. When Block2 completes the calculation, Block3 and Block4 are executed in the same way until the edges in all blocks are aggregated, and the entire calculation process ends.

[0121] The aggregation unit is composed of multiple groups of PEs. The number of PEs determines the number of edges that can be processed simultaneously and can be configured according to actual application conditions and bandwidth requirements. Figure 5 The index-driven MAC shown in the figure is composed of destination node index, destination node feature, edge value, source node feature, and output is the updated destination node feature. The input edge value and source node feature are multiplied by a multiplier to obtain a temporary feature, and the temporary feature enters two adders at the same time. One temporary feature is added to the input destination node feature through adder 1 as the first input of Mux, and the other temporary feature is added to the result obtained by delaying one cycle of the output of Mux in the previous cycle through D flip-flop 2 through adder 2 as the second input of Mux. The destination node index is divided into two paths, one is directly input to the comparator, and the other is input to the comparator after delaying one cycle through D flip-flop 1, that is, the destination node index of the current cycle and the destination node index input in the previous cycle enter the comparator together, and the comparator determines whether the two input indexes are equal. If they are equal, the output signal is 1, and if they are not equal, the output signal is 0. The output of the comparator is used as the selection signal of Mux, that is, when they are equal, the result output by adder 1 is selected, and when they are not equal, the result output by adder 0 is selected. The destination node index of the previous cycle and the output destination node characteristics are retained through index MAC. If the input destination node index of the current cycle is the same as the input destination node index of the previous cycle, the output destination node characteristics are superimposed on the output destination node characteristics of the previous cycle; if the input destination node index of the current cycle is different from the input destination node index of the previous cycle, the output destination node characteristics are superimposed on the current input destination node characteristics.

[0122] By driving the MAC with an index, memory access conflicts caused by continuous access to the same destination node feature are avoided. For example, Figure 6The memory access process of continuously accessing the same destination node Ni is shown, which is divided into three stages: reading the destination node characteristics of node Ni, completing the aggregation of destination node Ni, and writing back the destination node characteristics of node Ni. The destination node characteristics of node Ni are read at time T0, and the destination node characteristics of node Ni are read at time T1 and the aggregation of destination node Ni is completed at the same time. The destination node characteristics of node Ni read at time T0 are processed at time T1; the destination node characteristics of node Ni read at time T2, the aggregation of destination node Ni is completed and the characteristics of destination node Ni aggregated at time T1 are written back at time T2; the destination node characteristics of node Nj are read at time T3, the aggregation of destination node Ni is completed and the characteristics of destination node Ni aggregated at time T2 are written back at time T3; the destination node characteristics of node Nj read at time T2 are processed at time T3 and the characteristics of destination node Ni aggregated at time T2 are written back at time T4; the destination node characteristics of node Nj read at time T3 are processed at time T4 and the destination node characteristics of node Ni aggregated at time T3 are written back at time T5. At time T2, since the processing of the Ni node has not yet been completed, simultaneous read and write requests for the node Ni will appear in the pipeline (Pipeline), resulting in a memory access conflict. If you wait for the Ni data to be updated and then read it out in the next cycle, the pipeline will be blocked, and when the destination node characteristics of Ni are continuously read at times T0, T1 and T2, the destination node characteristics that have not yet been written to the memory cannot be immediately read out at times T1 and T2. Waiting for the destination node characteristics of the node Ni to be updated before reading will also cause the pipeline to be blocked. In the present invention, PE uses an index to drive MAC to temporarily store the characteristics of the current destination node during the aggregation stage. When the next aggregation operation is performed, the destination node index is compared with the previous aggregation. When the index is consistent, the temporarily stored destination node characteristics are selected; when the index is inconsistent, the input destination node characteristics are selected. When there is a read or write request at time T2, the feature of the destination node Ni obtained by processing at time T1 is selected. For the write-back end, since there is still a read request for Ni at time T2, the aggregation calculation of node Ni has not yet ended. At time T4, the final destination node feature is written back to the memory to overwrite the data written back at T2 and T3. By driving MAC with an index, the storage end only needs to pay attention to the first read of the destination node feature corresponding to Ni at time T0 and the last write-back at time T4. At times T1, T2, and T3, the destination node feature of Ni that is temporarily stored is used by driving MAC with an index, and the memory access process at the intermediate time can be ignored, avoiding the impact of potential memory access conflicts at the intermediate time.

[0123] For parallel processing, high parallelism can be achieved by configuring the number of PEs in the aggregation computing unit. PEs are independent of each other and can process different edges in the same way at the same time. When processing multiple edges at the same time, each PE only needs to complete data interaction with its corresponding source node feature navigation port and destination node feature storage port. When initializing the source node feature navigation, the source node feature corresponding to the current Block is broadcast to the port corresponding to each group of PEs. Each port can output its source node feature according to its corresponding source node index. For the destination node feature, a distributed storage strategy is adopted, such as Figure 7 As shown, each PE has an independent destination node storage port to store its partial sum. When all blocks are aggregated, when the next stage needs to access the destination node features, the corresponding destination node features in each group of destination node feature memories are taken out, and the partial sums in each group of memories are merged through the addition tree inside the aggregation control module to obtain the final destination node features, which are sent to the next processing stage according to actual computing requirements.

[0124] In summary, the present invention proposes a graph convolutional neural network aggregation computing acceleration device and method, which includes: a front-end module, an off-chip storage module, an aggregation control unit, a source node feature storage module, a block offset storage module, an edge storage module, an edge prefetcher, a data allocation module, a source node feature navigation module, a destination node feature storage module, and an aggregation computing unit.

[0125] The front-end module receives graph data and completes the combination operation to generate source node features, and transmits the source node features together with the graph adjacency matrix to the off-chip storage module.

[0126] The off-chip storage module stores the graph adjacency matrix compressed by CSB coding and the source node features output by the front module. The off-chip storage module is connected to the aggregation control unit and is responsible for providing data to the aggregation control unit.

[0127] The aggregation control unit connects the off-chip storage module, the source node feature storage module, the block offset storage module, the edge storage module, and the destination node feature storage module to coordinate data transmission and computing process control among the modules.

[0128] The source node feature storage module temporarily stores source node features corresponding to multiple blocks. The module is connected to the aggregation control unit to ensure that the source node features can be efficiently transmitted to the aggregation calculation unit during the calculation process.

[0129] The block offset storage module stores the block offset information of the graph adjacency matrix, connects to the aggregation control unit, and helps control the number of edges read by the computing unit in each block.

[0130] The edge storage module stores edge information, including source node index, destination node index and edge value. The module is connected to the aggregation control unit and is used to provide edge data to the computing unit.

[0131] The edge prefetcher controls the loading and calculation start and stop of edge data in the edge storage module according to the block offset information in the block offset storage module, ensuring that the edge data is transmitted to the data distribution module at the right time.

[0132] The data distribution module receives the edge data provided by the edge storage module, passes the source node index to the source node feature navigation module, passes the destination node index to the destination node feature storage module, and transmits the edge value to the aggregation calculation unit.

[0133] The source node feature navigation module reads the corresponding source node features from the source node feature storage module according to the source node index and passes them to the aggregation calculation unit.

[0134] The destination node feature storage module is connected to the data distribution module and the aggregation calculation unit, stores the updated destination node features, and provides the destination node features to the aggregation calculation unit during the calculation process.

[0135] The aggregation calculation unit connects the source node feature navigation module and the destination node feature storage module, multiplies the source node feature with the edge value and adds the result to the destination node feature, updates the destination node feature and writes it back to the destination node feature storage module.

[0136] The method includes:

[0137] S1. Perform CSB encoding compression on the graph data of the graph convolutional neural network, divide the graph adjacency matrix into blocks according to the source node index, store a fixed number of source node features in each block, retain only the source node index, destination node index and edge value of non-zero edges, and record the block offset information.

[0138] S2. The features output by the combined operation are used as source node features and written into the off-chip storage module together with the graph adjacency matrix data compressed by CSB encoding as data input for the aggregation stage.

[0139] S3. The aggregation control unit loads the source node features of each block from the off-chip storage module to the source node feature storage module, and loads the block offset information and edge information to the block offset storage module and the edge storage module respectively.

[0140] S4. The edge prefetcher controls the loading of edge data one by one and the start and stop of operations according to the block offset information in the block offset storage module, and transmits the loaded edge data to the data distribution module in sequence.

[0141] S5. The data distribution module sends the source node index of each edge to the source node feature navigation module, sends the destination node index to the destination node feature storage module, and transmits the edge value to the aggregation calculation unit.

[0142] S6. The source node feature navigation module outputs the corresponding source node feature to the aggregation calculation unit according to the input source node index, and the destination node feature storage module outputs the corresponding destination node feature to the aggregation calculation unit according to the destination node index.

[0143] S7. The aggregation calculation unit multiplies the source node feature with the edge value, and adds the product result to the destination node feature, generates an updated destination node feature and writes it back to the destination node feature storage module.

[0144] S8. After completing the data calculation of all edges in sequence, the aggregation control unit updates the source node features of the next block to the source node feature navigation module, and completes the aggregation operations of all blocks in sequence according to the above steps.

[0145] Although the present invention is disclosed in detail with reference to the accompanying drawings, it should be understood that these descriptions are merely exemplary and are not intended to limit the application of the present invention. The scope of protection of the present invention is defined by the appended claims and may include various modifications, alterations and equivalents made to the invention without departing from the scope and spirit of the present invention.

Claims

1. A graph convolutional neural network aggregation computing acceleration device, characterized in that: The device comprises: a front-end module, an off-chip storage module, an aggregation control unit, a source node feature storage module, a block offset storage module, an edge storage module, an edge prefetcher, a data distribution module, a source node feature navigation module, a destination node feature storage module, and an aggregation calculation unit; The front-end module receives the graph data and completes the combination operation, calculates all the source node features in the graph data, and compresses the graph adjacency matrix in the graph data by CSB coding. The graph adjacency matrix is ​​divided into blocks according to the source node index, and a fixed number of source node features are stored in each block. The calculated source node features are transmitted together with the compressed graph adjacency matrix to the off-chip storage module for storage; The aggregation control unit connects the off-chip storage module, the source node feature storage module, the block offset storage module, the edge storage module, and the destination node feature storage module to coordinate the data transmission between the modules and control the calculation process; The source node feature storage module temporarily stores the source node features corresponding to multiple blocks; the block offset storage module stores the block offset information of the graph adjacency matrix; the edge storage module stores the edge information, including the source node index, the destination node index and the edge value; the edge prefetcher controls the loading and calculation start and stop of the edge data in the edge storage module according to the block offset information in the block offset storage module, ensuring that the edge data is transmitted to the data distribution module at the right time; The data distribution module receives the edge data provided by the edge storage module, passes the source node index to the source node feature navigation module, passes the destination node index to the destination node feature storage module, and transmits the edge value to the aggregation calculation unit; The source node feature navigation module reads the corresponding source node feature from the source node feature storage module according to the source node index and passes it to the aggregation calculation unit; The aggregation calculation unit connects the source node feature navigation module and the destination node feature storage module, multiplies the source node feature by the edge value and adds the result to the destination node feature, updates the destination node feature and writes it back to the destination node feature storage module; The destination node feature storage module is connected to the data distribution module and the aggregation calculation unit, and is used to store the updated destination node features and provide the destination node features to the aggregation calculation unit during the calculation process.

2. A method for accelerating aggregate computing based on the graph convolutional neural network aggregate computing acceleration device of claim 1, characterized in that: The method comprises: S1. Perform CSB coding compression on the graph data of the graph convolutional neural network, divide the graph adjacency matrix into blocks according to the source node index, store a fixed number of source node features in each block, retain only the source node index, destination node index and edge value of non-zero edges, and record the block offset information; S2, writing the source node features in step S1 and the graph adjacency matrix data compressed by CSB coding into an off-chip storage module as data input for the aggregation stage; S3, the aggregation control unit loads the source node feature of each block from the off-chip storage module to the source node feature storage module, and loads the block offset information and edge information to the block offset storage module and the edge storage module respectively; S4, the edge prefetcher controls the loading of edge data and the start and stop of operations according to the block offset information in the block offset storage module, and transmits the loaded edge data to the data distribution module in sequence; S5, the data distribution module sends the source node index of each edge to the source node feature navigation module, sends the destination node index to the destination node feature storage module, and transmits the edge value to the aggregation calculation unit; S6, the source node feature navigation module outputs the corresponding source node feature to the aggregation calculation unit according to the input source node index, and the destination node feature storage module outputs the corresponding destination node feature to the aggregation calculation unit according to the destination node index; S7, the aggregation calculation unit multiplies the source node feature and the edge value, and adds the product result to the destination node feature, generates an updated destination node feature and writes it back to the destination node feature storage module; S8. After completing the data calculation of all edges in sequence, the aggregation control unit updates the source node features of the next block to the source node feature navigation module, and completes the aggregation operations of all blocks in sequence according to the above steps.

3. The graph convolutional neural network aggregation calculation acceleration method according to claim 2, characterized in that: The compression step in step S1 includes: The graph adjacency matrix is ​​divided into multiple blocks according to the source node index, and each block contains a fixed number of source node features; only non-zero elements are retained in each block, and their values ​​are stored together with the corresponding source node and destination node indexes; the number of edges in each block is recorded through block offset information, reducing storage space.

4. The graph convolutional neural network aggregation calculation acceleration method according to claim 2, characterized in that: The data operation relationship of each layer of the graph convolutional neural network algorithm is: in, is an adjacency matrix with self-loops, X l is the output feature of the current layer, X l-1 is the output feature of the previous layer, that is, the input feature of the current layer, W l is the input weight of the current layer, δ represents the activation function ReLU, and the classic two-layer graph convolutional neural network operation is expressed as: Among them, the main operation is divided into two parts. The combination operation represents the process of multiplying the node features by the weights to achieve the dimension reduction of the features, namely: The aggregation operation is the process of multiplying the graph adjacency matrix and the feature matrix. This part is a sparse irregular operation, which can be expressed as: For X l The combinatorial operation part, For X l The aggregation operation part.

5. The graph convolutional neural network aggregation calculation acceleration method according to claim 2, characterized in that: The edge pre-fetching step in step S3 includes: The edge prefetcher controls loading edge data from the edge storage module according to the block offset information in the block offset storage module, and transmits the loaded edge data to the data distribution module; The edge prefetcher ensures that only the edge data required by the current block is loaded each time, thereby reducing the loading and latency of redundant data.

6. The graph convolutional neural network aggregation calculation acceleration method according to claim 2, characterized in that: The data allocation step in step S4 includes: The data allocation module splits the loaded source node feature data and edge data, and transmits the source node features to the source node feature navigation module according to the source node index, and transmits the destination node features to the destination node feature storage module according to the destination node index; the edge value is passed to the aggregation calculation unit for feature calculation.

7. The graph convolutional neural network aggregation calculation acceleration method according to claim 2, characterized in that: The feature navigation step in step S5 includes: source node feature navigation and destination node feature navigation: The source node feature navigation module reads the corresponding source node feature from the source node feature storage module according to the input source node index, and transmits it to the aggregation calculation unit; The destination node feature storage module reads the corresponding destination node feature according to the destination node index and transmits it to the aggregation calculation unit.

8. The graph convolutional neural network aggregation calculation acceleration method according to claim 2, characterized in that: The aggregation calculation step in step S6 includes: The aggregation calculation unit multiplies the source node feature with the edge value to generate a temporary feature; the temporary feature is accumulated to the destination node feature through an addition operation to obtain an updated destination node feature; the updated destination node feature is written back to the destination node feature storage module for subsequent calculations.

9. The graph convolutional neural network aggregation calculation acceleration method according to claim 2, characterized in that: The feature updating step in step S7 includes: After each block calculation is completed, the aggregation calculation unit adds the product of the source node feature and the edge value to the destination node feature to update the feature of the node; the updated destination node feature is written back to the destination node feature storage module to prepare for the next round of calculation.

10. The graph convolutional neural network aggregation calculation acceleration method according to claim 2, characterized in that: The iterative calculation step in step S8 includes: According to the description in step S6 and step S7, aggregate calculation is performed on each block in turn until the edge data of all blocks are calculated; After all block calculations are completed, the destination node feature storage module aggregates the stored destination node features to form the final calculation result.