A software-hardware collaborative acceleration method for graph convolutional networks based on in-memory computing
Through quantitative fixed-point and graph clustering mapping technology, graph data is converted into fixed-point numbers with low bit width and divided into multiple sub-graphs, solving the problems of low computational parallelism and low hardware resource utilization in the prior art, and the acceleration of GCN calculations of graph convolution networks and the support of 32-bit floating-point numbers are realized.
Patent Information
- Application Number
- CN202210267969.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-03-17
AI Technical Summary
The existing in-memory computing-based architecture used for GCN calculations has problems such as low computational parallelism, low hardware resource utilization and inability to support 32-bit floating-point calculations.
A method of co-acceleration of graph convolution network software and hardware based on in-memory computing is designed. The graph data is converted into fixed-point numbers with low bit width through quantitative fixed-points, and the graph data is divided into multiple sub-graphs using a mapping algorithm based on graph clustering, and mapped onto RRAM cross-arrays. At the same time, edge deletion technology is used to improve the computational parallelism and hardware resource utilization.
It realizes the acceleration of GCN calculation of graph convolution network, improves the parallelism of the calculation, the throughput of the architecture and the resource utilization of the hardware, and can support the GCN model of 32-bit floating point number.
Smart Images

Figure CN114707648B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of in-memory computing, and in particular to a method and device for collaborative software and hardware acceleration of a graph convolutional network based on in-memory computing. Background Art
[0002] Graph structure is a commonly used data structure, which consists of many nodes and edges connecting the nodes. Traffic networks, social networks, molecular structures and other structures commonly used in daily life and research can be abstractly represented by graph structures. In recent years, the graph neural network model (GNN) has demonstrated powerful learning capabilities in graph-based tasks, such as recommendation systems and document classification. The graph convolutional network model (GCN) has become the most commonly used GNN model due to its low computational cost and excellent accuracy performance. Figure 1 As shown in the figure, each node in the graph task will have a feature vector containing the information of the node, and the purpose of GCN is to infer the required conclusions based on this information and the connection relationship between nodes, such as inferring the classification of documents through the authors and citation relationships of documents. A typical graph convolutional network model contains several network layers, each of which contains two stages, namely neighbor node feature aggregation and feature vector combination. In the neighbor node feature aggregation stage, each node will obtain the feature vectors of its neighboring points, and make these vectors act on an aggregation function (such as summation, averaging function) together. This step can usually be abstracted into sparse matrix multiplication using a sparse adjacency matrix and a dense feature matrix. In the vector combination stage, the results obtained in the previous stage will be sent to a multilayer perceptron (MLP), and the results will be used as the node feature matrix of the next layer of the network. If written in the form of a matrix, the graph convolutional network can be expressed as follows:
[0003] H l+1 =σ(AH l W)
[0004] Where σ represents a nonlinear function, H l+1 ,H 1 represents the node feature matrix of the l+1th and lth layers, W represents the weight matrix of the MLP, these matrices are generally dense; and A represents the adjacency matrix of the graph, which is generally irregular and sparse.
[0005] The difference between GCN and traditional neural networks (such as convolutional neural networks (CNN)) lies in the irregularity of the graph structure. In the feature aggregation stage, GCN needs to collect the feature vectors of all (or part) of the node's neighbor nodes, and these feature vectors are not placed continuously in the memory. And because the scale of real graphs is generally large, the burden of memory access during GCN calculations will be further increased. In the traditional von Neumann architecture, the computing performance of the architecture will be greatly limited due to the irregular memory access and large amount of data handling in the GCN calculation process. Taking the calculation of GCN with a graphics processing unit (GPU) as an example, some scholars pointed out that when the dimension of the feature vector increases, the amount of data that the GPU needs to load from the global memory for each calculation also increases accordingly, and the global throughput is gradually limited by the global bandwidth. As Figure 2 As shown in the figure, when the feature vector dimension increases to 32, the global throughput is limited by the global bandwidth and no longer increases with the feature dimension. This reflects the huge burden of data transfer when using the von Neumann architecture to calculate GCN. These data transfers cause the overall GCN computing performance to be limited by the bandwidth bottleneck.
[0006] In 2008, scientists developed Resistive Random Access Memory (RRAM), a device with variable resistance. As long as a specific voltage is applied to both ends of the device, its resistance can be changed, and the resistance can remain unchanged after power failure. Therefore, the resistance value of RRAM can be used to represent the stored data. The researchers arranged the RRAM devices in rows and columns to form a cross array structure, using the conductivity value of RRAM to represent the value of the matrix element, and using the voltage on the word line as the input vector to realize the matrix-vector multiplication in the memory. According to Ohm's law, the relationship between the output current on the bit line and the word line voltage and the RRAM conductivity is:
[0007]
[0008] Written in matrix form:
[0009] I=CV
[0010] In this way, matrix-vector multiplication with O(1) time complexity can be implemented in memory, and the weight data can always be kept in the memory, eliminating the need to move a large amount of weight data. Since the main computational process of a neural network can be represented in the form of matrix-vector multiplication, existing research has designed a CNN accelerator based on RRAM in-memory computing based on this method, and successfully achieved an extremely high acceleration ratio. The implementation scheme that is most similar to this scheme is PRIME, which expands the convolution kernel weights of CNN and maps them to an RRAM cross array, and uses the feature map of each layer as the input vector to calculate the matrix-vector multiplication to calculate the convolution of each layer. The data mapping method and hardware architecture design proposed in this scheme have made important contributions to solving the in-memory computing problem of CNN. If this implementation scheme is used to implement the calculation of GCN, because the size of the adjacency matrix A is extremely large (the number of matrix elements generally exceeds 10 8 ) and is very sparse (the number of non-zero elements generally does not exceed 0.1%). Mapping A to RRAM requires an extremely large array, so the feature matrix H and the weight matrix W should be mapped to the RRAM cross array, and each row of A should be input separately, and then matrix-vector multiplication with H and W in turn. Finally, when all rows of A are input and the calculation is completed, the calculation result will be used to update the H matrix, marking the end of a layer of GCN calculation.
[0011] Currently, there are still many problems with the in-memory computing architecture for GCN computing, as listed below:
[0012] 1. The computational parallelism between nodes in in-memory computing is low, and feature aggregation of different nodes cannot be calculated in parallel. The existing in-memory computing architecture uses a cross array of RRAM to store matrices, and converts vectors into voltages loaded on word lines to implement matrix-vector multiplication. In order to calculate the neighbor feature aggregation step in GCN, the matrix multiplication between the adjacency matrix and the feature matrix needs to be divided into many matrix-vector multiplications, and these matrix-vector multiplications can only be executed sequentially, resulting in a high computational delay cost. Taking a typical graph dataset, the ARXIV dataset, as an example, in the feature aggregation stage, the original sparse matrix-matrix multiplication needs to be decomposed into 169,343 matrix-vector multiplications executed sequentially, resulting in a high computational delay;
[0013] 2. The existing in-memory computing architecture has very low hardware resource utilization when calculating GCN, and does not fully utilize the high computational parallelism of the RRAM crossbar array in-memory computing. In the existing in-memory computing architecture, the input vector of the matrix-vector multiplication will be converted into a voltage vector and loaded onto the word line of the RRAM crossbar array to activate the corresponding row for calculation. However, when calculating sparse matrix multiplication, the sparse input vector will only activate a small part of the array rows, and the remaining array rows will be idle. For example, on the typical graph dataset PUBMED, the feature aggregation operation of GCN will only activate 4.5 rows out of 19717 rows on average each time, resulting in a serious problem of low array utilization;
[0014] 3. The existing RRAM-based in-memory computing architecture mainly supports fixed-point computing and uses a lower data bit width to further reduce the computing cost. The mainstream high-accuracy GCN models all use 32-bit floating-point numbers. Deploying a 32-bit floating-point GCN model on an in-memory computing architecture suitable for low-data-bit-width fixed-point computing is still a problem that needs to be solved. Summary of the invention
[0015] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.
[0016] To this end, the purpose of the present invention is to design a GCN deployment algorithm for multi-core in-memory computing hardware at the deployment algorithm level, to solve the problem of how to map irregular graph data to a hardware array of a certain scale and at the same time improve the computing parallelism. At the hardware level, it solves the problem of how to design in-memory computing units and data flows, and proposes a graph convolutional network hardware and software collaborative acceleration method based on in-memory computing.
[0017] Another object of the present invention is to propose a graph convolutional network software and hardware collaborative acceleration device based on in-memory computing.
[0018] To achieve the above objectives, the present invention proposes a method for co-accelerating graph convolutional networks based on in-memory computing, comprising the following steps:
[0019] Software: In the in-memory calculation of graph data, the high-bit-width floating-point number of the graph data is converted into a low-bit-width fixed-point number by quantizing the fixed-point method; and the original graph of the graph data is divided into multiple subgraphs to obtain clustering results by using a mapping algorithm based on graph clustering, and the node features of the multiple subgraphs are mapped to the RRAM crossbar array for calculating feature aggregation; and, based on the clustering results, the edges connecting different subgraphs in the clustering results are deleted by using an edge deletion method, so as to deploy the graph data to the hardware through a hardware-oriented deployment algorithm;
[0020] Hardware: configure a computing module and a control module; wherein the computing module includes: an aggregation core array, a vector combination array and an intermediate cache, wherein the aggregation core array is used to calculate the feature aggregation of the graph data, the vector combination array is used to calculate the vector combination of the graph data, and the intermediate cache is used for data exchange between the aggregation core array and the vector combination array; the control module includes: an instruction queue, a graph data decoder and a neighbor cache, which are respectively used for temporary storage of instructions of the graph data, format conversion and temporary storage of adjacent vectors;
[0021] Through the collaborative design of the software and the hardware, the acceleration of graph convolutional network (GCN) calculation is achieved.
[0022] According to the in-memory computing-based graph convolution network software-hardware collaborative acceleration method of the embodiment of the present invention, in terms of software, the high-bitwidth floating-point numbers of the graph data are converted into low-bitwidth fixed-point numbers by quantizing the fixed-point method, the original graph of the graph data is divided into multiple subgraphs to obtain clustering results, and the node features of the multiple subgraphs are mapped to the RRAM cross array for feature aggregation, and then the edges connecting different subgraphs in the clustering results are deleted by edge deletion to deploy the graph data on the hardware; in terms of hardware, the computing module and the control module are configured, wherein the computing module includes: an aggregation core array, a vector combination array and an intermediate cache, and the control module includes: an instruction queue, a graph data decoder and a neighbor cache; through the collaborative design of software and hardware, the acceleration of the graph convolution network GCN calculation is achieved. The present invention improves the parallelism of the calculation, the throughput of the architecture and the resource utilization of the hardware.
[0023] In addition, the graph convolutional network software and hardware collaborative acceleration method based on in-memory computing according to the above embodiment of the present invention also includes:
[0024] Furthermore, in the in-memory calculation of the graph data, converting the high-bit-width floating-point number of the graph data into a low-bit-width fixed-point number by quantizing the fixed-point method includes:
[0025] In the in-memory calculation based on the RRAM, a quantized fixed-point method for graph convolutional networks is used. For the feature data and network weights of each layer of the graph convolutional network, the total quantization bit width of the data is preset. total Finally, according to the data to be quantized, the number of bits Width allocated to the integer and decimal parts is calculated by the following formula int With Width dec :
[0026]
[0027] Width dec =Width total-Width int
[0028] Furthermore, the mapping algorithm based on graph clustering is used to divide the original graph of the graph data into multiple subgraphs to obtain clustering results, and the node features of the multiple subgraphs are mapped to the RRAM crossbar array for calculating feature aggregation, including:
[0029] Use the Graclus algorithm to cluster the original graph of the graph data into multiple subgraphs, and maximize the connections within the subgraphs and minimize the connections between subgraphs to obtain the clustering results;
[0030] According to the clustering result, the feature matrix is segmented, and the feature vectors corresponding to the nodes of different subgraphs are mapped to different resistive random access memory cross arrays RRAM; wherein each cross array calculates the feature aggregation of the nodes in the corresponding subgraph.
[0031] Furthermore, the method further includes: predefining an edge deletion rate threshold T, and making the number of adjacent edges deleted from each node in the original graph smaller than the product of the original degree of the node and T in the edge deletion algorithm; wherein the value range of T is 0 <T≤1。
[0032] Furthermore, the edge deletion algorithm is:
[0033] Input the nodes of the original graph and information of each of the subgraphs;
[0034] For each node in the original graph, iteratively select a subgraph that participates in the feature aggregation of this node;
[0035] In each iteration, a subgraph that covers the most uncovered neighbors of this node is selected to participate in the feature aggregation of this node. When the proportion of covered neighbors of a node is greater than or equal to T, the iteration terminates.
[0036] Furthermore, the aggregate core array includes a plurality of RRAM-based aggregate cores, and the plurality of RRAM-based aggregate cores are interconnected in the form of an on-chip network to perform matrix-vector multiplication calculations in parallel.
[0037] Furthermore, the aggregation core array is divided into a working area and an update area in a ping-pong working form. When the working area performs feature aggregation operations, the update area writes the calculated node features into the RRAM cross array.
[0038] Furthermore, the method also includes: using the data forwarding unit to receive data from adjacent units or the corresponding RRAM cross array, forwarding it to other adjacent units after merging processing, and organizing it in the form of an on-chip network to be controlled by the control module.
[0039] Furthermore, the vector combination array includes a waiting queue and multiple MLP cores, which are used to parallelly calculate the vector combination stage of GCN; wherein the MLP core uses the RRAM cross array to calculate the matrix-vector multiplication and uses the lookup table to calculate the nonlinear activation function. When calculating the vector combination stage of GCN, the intermediate data stored in the intermediate cache will first be written into the waiting queue of the vector combination array.
[0040] To achieve the above object, the present invention proposes a graph convolutional network software and hardware collaborative acceleration device based on in-memory computing, comprising:
[0041] A software design module, for converting a high-bit-width floating-point number of the graph data into a low-bit-width fixed-point number by quantizing the fixed-point method in the in-memory calculation of the graph data; and using a mapping algorithm based on graph clustering to divide the original graph of the graph data into multiple subgraphs to obtain clustering results, and mapping the node features of the multiple subgraphs to an RRAM crossbar array for calculating feature aggregation; and, based on the clustering results, using an edge deletion method, deleting the edges connecting different subgraphs in the clustering results, so as to deploy the graph data on hardware through a hardware-oriented deployment algorithm;
[0042] A hardware design module, used to configure a computing module and a control module; wherein the computing module includes: an aggregation core array, a vector combination array and an intermediate cache, wherein the aggregation core array is used to calculate feature aggregation of the graph data, the vector combination array is used to calculate vector combination of the graph data, and the intermediate cache is used for data exchange between the aggregation core array and the vector combination array; the control module includes: an instruction queue, a graph data decoder and a neighbor cache, which are respectively used for temporary storage of instructions of the graph data, format conversion and temporary storage of adjacent vectors;
[0043] The collaborative acceleration module is used to accelerate the calculation of the graph convolutional network GCN through the collaborative design of the software and the hardware.
[0044] According to the in-memory computing-based graph convolution network software-hardware collaborative acceleration device of the embodiment of the present invention, in terms of software, the high-bit-width floating-point numbers of the graph data are converted into low-bit-width fixed-point numbers by quantizing the fixed-point method, the original graph of the graph data is divided into multiple subgraphs to obtain clustering results, and the node features of the multiple subgraphs are mapped to the RRAM cross array for feature aggregation, and then the edges connecting different subgraphs in the clustering results are deleted by edge deletion to deploy the graph data on the hardware; in terms of hardware, the computing module and the control module are configured, wherein the computing module includes: an aggregation core array, a vector combination array and an intermediate cache, and the control module includes: an instruction queue, a graph data decoder and a neighbor cache; through the collaborative design of software and hardware, the acceleration of the graph convolution network GCN calculation is achieved. The present invention improves the parallelism of the calculation, the throughput of the architecture and the resource utilization of the hardware.
[0045] Beneficial effects of the present invention:
[0046] (1) At the software level, the data is deployed on a multi-core computing architecture through the node clustering-based GCN mapping algorithm, which converts large-scale sparse graphs into small-scale dense subgraphs that can be processed independently in parallel, thereby increasing the computational parallelism at the subgraph level. Based on this algorithm, the present invention proposes an edge deletion technology, which reduces the communication cost in the aggregation stage and also reduces the data dependency between subgraphs, thereby further improving the subgraph level parallelism. In addition, the present invention also proposes a graph data quantization method to reduce the amount of hardware computation and storage.
[0047] (2) At the hardware level, the present invention designs a multi-core GCN computing architecture based on RRAM in-memory computing, which supports parallel computing at the subgraph level. At the same time, the present invention also designs the data flow of the intra-layer pipeline, further improving the throughput of the architecture.
[0048] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0050] Figure 1 A schematic diagram of the calculation process of GCN in the prior art;
[0051] Figure 2 A schematic diagram showing the relationship between feature vector dimension and throughput when computing GCN on a GPU in the prior art;
[0052] Figure 3A schematic diagram of a framework of a method for co-accelerating graph convolutional networks using software and hardware based on in-memory computing according to an embodiment of the present invention;
[0053] Figure 4 A schematic diagram of a flow chart of a method for co-accelerating graph convolutional networks using software and hardware based on in-memory computing according to an embodiment of the present invention;
[0054] Figure 5 A schematic diagram of GCN mapping based on graph clustering according to an embodiment of the present invention;
[0055] Figure 6 The pseudo code of the algorithm of the edge deletion technology according to the embodiment of the present invention;
[0056] Figure 7 A schematic diagram of a GCN computing architecture based on in-memory computing according to an embodiment of the present invention;
[0057] Figure 8 Schematic diagram of delay and energy consumption experimental results according to an embodiment of the present invention;
[0058] Fig. 9 A schematic diagram of a trade-off between accuracy and latency according to an embodiment of the present invention;
[0059] Fig.10 It is a structural schematic diagram of a graph convolutional network software and hardware collaborative acceleration device based on in-memory computing according to an embodiment of the present invention. DETAILED DESCRIPTION
[0060] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0061] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0062] The following describes the graph convolutional network software and hardware collaborative acceleration method and device based on in-memory computing proposed in accordance with an embodiment of the present invention with reference to the accompanying drawings. First, the graph convolutional network software and hardware collaborative acceleration method based on in-memory computing proposed in accordance with an embodiment of the present invention will be described with reference to the accompanying drawings.
[0063] The overall framework of the embodiment of the present invention is as follows Figure 3As shown. At the beginning of the entire computing flow, a hardware-oriented deployment algorithm is designed to perform preliminary processing on the graph data and determine the data mapping method, so as to deploy the data on the hardware. First, the quantization and fixed-point methods of GCN are designed to convert high-bitwidth floating-point numbers into low-bitwidth fixed-point numbers to reduce the cost of hardware deployment and calculation. Later, in order to better map the graph structure to the hardware, a mapping algorithm based on graph clustering is proposed to divide the large graph into multiple subgraphs, and map the node features of the subgraphs to the RRAM cross array used to calculate feature aggregation. After the mapping is completed, an edge deletion technology is designed to further make the irregular graph data hardware-friendly.
[0064] The embodiment of the present invention also designs a hardware architecture based on RRAM in-memory computing. The main computing modules of the architecture include an aggregate core array and a vector combination array.
[0065] Figure 4 It is a flowchart of a graph convolutional network software and hardware collaborative acceleration method based on in-memory computing according to an embodiment of the present invention.
[0066] like Figure 4 As shown in the figure, the graph convolution network software and hardware collaborative acceleration method based on in-memory computing includes the following steps:
[0067] Step S1, software aspect: in the in-memory calculation of graph data, the high-bit-width floating-point numbers of the graph data are converted into low-bit-width fixed-point numbers by quantizing the fixed-point method; and a mapping algorithm based on graph clustering is used to divide the original graph of the graph data into multiple subgraphs to obtain clustering results, and the node features of the multiple subgraphs are mapped to the RRAM cross array for calculating feature aggregation; and, based on the clustering results, the edge deletion method is used to delete the edges connecting different subgraphs in the clustering results, so as to deploy the graph data to the hardware through a hardware-oriented deployment algorithm.
[0068] Specifically, the software is divided into the following three parts:
[0069] (1) Design of quantized fixed-point solutions. In RRAM-based in-memory computing, if fixed-point calculations are used, the addition operation on the exponent bit will be omitted compared to the use of floating-point calculations, the computational cost will be lower, and the hardware architecture will be simpler. In the existing in-memory computing architecture for CNN, most work uses fixed-point numbers for calculations. Similarly, the data bit width will also affect the performance of the in-memory computing architecture. The original graph data bit width in GCN is generally 32 bits. If 32-bit data is used directly for in-memory calculations, a lot of storage space will be consumed and a higher-precision digital-to-analog and analog-to-digital conversion interface will be required. Generally speaking, the smaller the data bit width, the smaller the time, energy consumption and area required for calculation, but the greater the loss of accuracy.
[0070] The embodiment of the present invention designs a quantization fixed-point method for GCN. For each layer of feature data and network weights, the user can specify the total quantization bit width of the data. total After the given, the algorithm calculates the number of bits Width allocated to the integer and decimal parts according to the data to be quantized by the following formula int With Width dec :
[0071]
[0072] Width dec =Width total -Width int
[0073] As shown in the above formula, the algorithm will first determine the number of bits allocated to the integer part based on the maximum and minimum values of the data. If the number of bits required for the integer part is less than the quantization bit width specified by the user, the remaining bits are all allocated to the decimal part. When the bit widths of the integer and decimal parts are determined, the program will uniformly quantize the integer and decimal parts of the data according to the bit width.
[0074] (2) Mapping algorithm based on graph clustering. In order to realize sparse matrix multiplication calculation based on in-memory calculation, the in-memory calculation architecture needs to store one matrix on the RRAM crossbar array and split the other matrix into vectors, and complete the matrix multiplication by calculating multiple matrix-vector multiplications. Because the number of nodes in the graph is often much larger than the dimension of the feature vector (for example, the ARXIV dataset has 169,343 nodes, while the feature vector dimension is only 128), and the graph structure is generally very sparse (for example, only 0.028% of the data in the adjacency matrix in the PUBMED dataset is non-zero), it is chosen here to store the feature matrix in the RRAM crossbar array to reduce storage costs and improve resource utilization. Specifically, one embodiment of the present invention maps the dense feature matrix H to the RRAM crossbar array and converts the sparse adjacency matrix A into a series of continuous input vectors. However, due to the large number of input vectors and the ultra-sparseness of the graph structure, the existing in-memory computing architecture requires multiple iterations to complete the calculation, and each iteration only activates a small number of rows of the RRAM crossbar array to participate in the calculation, resulting in low utilization of the RRAM crossbar array hardware resources.
[0075] To solve the above two problems, the embodiment of the present invention proposes a mapping method based on graph clustering. This method uses a graph clustering algorithm to divide a sparse large graph into many smaller and denser subgraphs, and maps the feature vectors of nodes to the corresponding RRAM crossbars according to the clustering results. These crossbars can perform calculations simultaneously, thus reducing the number of iterations required in the aggregation process. The mapping process is shown in Figure 5 . First, the present invention uses the Graclus algorithm to cluster the graph into smaller subgraphs, maximizing the connections within the subgraphs and minimizing the connections between subgraphs. Then, according to the clustering results, the embodiment of the present invention divides the feature matrix and maps the feature vectors corresponding to the nodes of different subgraphs to different RRAM crossbars. Each crossbar can calculate the feature aggregation of the points within the corresponding subgraph, and different crossbars can work in parallel.
[0076] (3) Edge deletion technique. In addition to providing support for information transmission from the architecture, the embodiment of the present invention also proposes an edge deletion technique to reduce the data communication cost and the computational dependence in feature aggregation by deleting the edges between subgraphs. After the graph clustering mapping, there are still a small number of cross-subgraph edges in the graph, which connect the nodes belonging to different subgraphs. During the execution of the algorithm, due to the existence of these edges, the computing units responsible for different subgraphs need to perform calculations together and merge the calculation results, increasing the additional communication cost and data dependence. Therefore, the embodiment of the present invention reduces the communication cost and data dependence by deleting these edges. However, deleting too many edges will lead to a decrease in the accuracy of the GCN. To control the number of deleted edges, the embodiment of the present invention defines an edge deletion rate threshold T, and ensures that the number of adjacent edges deleted for each point in the algorithm is less than the product of the original degree of the point and T. The value range of T is 0 < T ≤ 1, and the larger T is, the higher the acceptable edge deletion rate is.
[0077] Figure 6 The pseudocode of Figure 6 shows the flow of the algorithm. The input of the algorithm is the nodes of the graph and the information of each subgraph. For each node in the whole graph, the algorithm iteratively selects the subgraphs participating in the feature aggregation of this node. In each iteration, the algorithm will select a subgraph that covers the most uncovered neighbors of this node ( Figure 6 line 4), and let it participate in the feature aggregation of this node ( line 6). When the covered ratio of a node's neighbors is greater than or equal to T, this iteration terminates. In this way, the algorithm obtains which subgraphs are needed for the feature aggregation of each node, that is, determines which edges are retained and which edges are deleted.
[0078] Step S2, hardware aspect: configure the computing module and the control module; wherein the computing module includes: an aggregation core array, a vector combination array and an intermediate cache, the aggregation core array is used to calculate the feature aggregation of the graph data, the vector combination array is used to calculate the vector combination of the graph data, and the intermediate cache is used for data exchange between the aggregation core array and the vector combination array; the control module includes: an instruction queue, a graph data decoder and a neighbor cache, which are respectively used for temporary storage of graph data instructions, format conversion and temporary storage of adjacent vectors.
[0079] Specifically, at the hardware level, such as Figure 7 As shown, an embodiment of the present invention designs a hardware architecture based on RRAM in-memory computing to realize in-memory GCN computing with high parallelism. The entire architecture mainly includes a control module and a computing module. The computing module includes an aggregation core array for calculating feature aggregation, a vector combination array for calculating vector combination, and an intermediate cache for data exchange between the two arrays. The control module includes an instruction queue, a graph data decoder and a neighbor cache, which are respectively used for temporary storage of instructions, conversion of graph data from CSR format to adjacent vector format, and temporary storage of adjacent vectors.
[0080] Polymer Core Array ( Figure 7 The (A) in Figure 1 contains multiple RRAM-based aggregate cores that are interconnected in the form of an on-chip network ( Figure 7 (B) in the figure), and the matrix-vector multiplication can be calculated in parallel. In the feature aggregation step, the feature matrix is statically stored in multiple RRAM cross arrays, and the vectors obtained by segmenting the adjacency matrix are dynamically loaded from the neighbor cache to the aggregation core. In each iteration of feature aggregation, multiple vectors are read out from the cache and then sent to each aggregation core through the data forwarding unit. In order to make full use of the computational parallelism brought by the mapping algorithm above, these aggregation cores can work independently and simultaneously. Because there may be connections between subgraphs, the embodiment of the present invention adds summers, comparators and other computing units to the data forwarding unit of the on-chip network to support the merging of result data. Similarly, these data forwarding units can also work in parallel and independently. Finally, the merged result will be sent to the intermediate cache, waiting for subsequent vector combination calculations.
[0081] At the end of each GCN network layer, the results of the vector combination stage will be used to update the feature matrix, that is, written back to the aggregation core array. In order to disperse the dense feature update operations at the end of each layer of calculation into the iterative process, the embodiment of the present invention divides the aggregation core array into two parts, namely the working area and the update area, and the two work in a ping-pong manner. When the working area performs feature aggregation operations, the update area can simultaneously write the calculated node features into the RRAM cross array. When the calculation of the next layer begins, the roles of the two will be exchanged, that is, the original update area will be responsible for the calculation and will be changed to the working area to perform feature aggregation calculations, while the original working area will be responsible for the update and will be changed to the update area to store the calculation results.
[0082] The internal details of the aggregation core are as follows Figure 7 As shown in (A) in the figure. In order to support different core sizes, each core includes multiple RRAM cross arrays and their supporting peripheral circuit components. In order to support different aggregation functions, the embodiment of the present invention places comparators and adders next to the RRAM cross array to merge the results output by different RRAM cross arrays. On the one hand, if the aggregation function is a mean function or a sum function, a digital-to-analog converter (DAC) is used to convert the input vector into a voltage, and after matrix-vector multiplication with the RRAM cross array, the result is read out through an analog-to-digital converter (ADC). On the other hand, if the aggregation function is a maximum value function, the embodiment of the present invention only turns on one row of the RRAM cross array at a time, and uses a comparator to compare row by row to finally select the maximum value.
[0083] The design of the data forwarding unit is as follows: Figure 7 As shown in (B) in the figure, the data forwarding unit can receive data from adjacent units or the corresponding RRAM crossbar array, and forward it to other adjacent units after merging. They are organized in the form of an on-chip network and are controlled by the control module.
[0084] like Figure 7As shown in (C) in the figure, the vector combination array contains a waiting queue and multiple MLP cores, which can calculate the vector combination stage in GCN in parallel. The MLP core uses the RRAM crossbar array to calculate the matrix-vector multiplication and uses the lookup table to calculate the nonlinear activation function. When calculating the vector combination stage of GCN, the intermediate data stored in the intermediate cache will first be written to the waiting queue of the vector combination array. Unlike the aggregate core, for each GCN network layer, each MLP core of the vector combination array stores all the MLP weights of this layer. Therefore, the vectors stored in the waiting queue can be sent to any idle MLP core of the layer for calculation, so that the vector combination array can well inherit the node-level computing parallelism brought by the aggregate core array. In each MLP core, the weights of different MLP layers are stored on different RRAM crossbar arrays, so different MLP layers can work in a pipelined manner. When the first input vector completes the calculation of the first layer, the second input vector can be loaded into the RRAM crossbar array of the first layer to prepare for the matrix-vector multiplication of the first layer. The embodiment of the present invention designs a multi-layer calculation pipeline in the MLP core, further improving the throughput of the vector combination array. Finally, the calculation result of the last layer of the MLP will obtain the final result through the lookup table of the nonlinear activation function and be sent to the intermediate cache to wait for being written back to the update area.
[0085] Benefiting from the aggregation core array working in a ping-pong manner, an embodiment of the present invention can divide an iterative process of GCN calculation into three stages: feature aggregation, vector combination, and feature update. In the same iteration, these three stages will be performed in the current working area of the aggregation core, the vector combination core, and the current update area of the aggregation core. Therefore, when performing GCN calculations, different GCN layers perform calculations in sequence, and within the layer, the three stages of feature aggregation, vector combination, and feature update can work in the form of a pipeline. In this pipeline-style data flow, the feature update operations that were originally concentrated at the end of each layer are dispersed throughout the entire calculation process and are hidden by the delays of feature aggregation and vector combination operations, thereby improving the throughput of the architecture.
[0086] Step S3, accelerate the calculation of graph convolutional network GCN through collaborative design of software and hardware.
[0087] It can be understood that the present invention first solves the problem of how to map irregular graph data to a hardware array of a certain scale and improve the computational parallelism by designing a GCN deployment algorithm for multi-core in-memory computing hardware; at the hardware level, the present invention solves the problem of how to design in-memory computing units and data flows. Through the above hardware and software collaborative design, the present invention solves the problems of low computational parallelism and low resource utilization, and ultimately achieves the acceleration of GCN computing.
[0088] The embodiments of the present invention are further described below by means of the accompanying drawings:
[0089] Specifically, the experimental effects of the present invention are as follows:
[0090] The experiment uses the GCN model to evaluate the performance of the embodiments of the present invention on four typical data sets: CORA, CITESEER, PUBMED, and ARXIV. In the in-memory computing baseline architecture and the architecture simulation of the present invention, the data of RRAM, ADC, and DAC are respectively derived from other literature. The digital circuit part is simulated at a frequency of 300MHz using Synopsys Design Compiler and TSMC65nm process. The experiment uses CACTI to simulate the cache, and the cache size is set to 64KB. At the same time, the experiment also uses the DGL framework to evaluate GCN on a 40-core Intel Xeon E5-2630 v4 2.20GHz CPU and Nvidia RTX 2080Ti GPU as CPU and GPU baselines.
[0091] (1) Delay and energy consumption simulation
[0092] The experiment compared the latency and energy consumption of the present invention with those of CPU, GPU, and PRIME on four data sets. The results are as follows: Figure 8 As shown. Figure 8 As can be seen from (a) in the figure, PRIME only achieved a 44-86 times speed increase on the GCN task compared to the CPU, which is much lower than its speed increase on the CNN task (1596-11802 times). This is because the graph data is extremely sparse and the edges are irregularly connected. By using the hardware-software co-design described above, the embodiments of the present invention achieved 87-664 times, 45-120 times, and 2.79-12.49 times speed increases compared to the CPU, GPU, and PRIME, respectively. Figure 8 (b) shows the energy efficiency improvement result of the present invention. Compared with CPU, GPU, and PRIME, the embodiments of the present invention achieve energy efficiency improvements of 2834 to 7108 times, 172 to 606 times, and 2.42 to 5.33 times, respectively.
[0093] (2) Accuracy Verification
[0094] The experiment mainly considers the impact of two factors on the accuracy of GCN, namely edge deletion technology and quantization error. The embodiment of the present invention uses edge deletion technology to delete edges that span two subgraphs, and sets an edge deletion threshold T to limit the proportion of edge deletion, thereby controlling the loss of accuracy. Therefore, there is a trade-off between delay and accuracy: a larger threshold T means that more edges are deleted, the delay will be reduced and the loss of accuracy will increase, and vice versa. There are similar results for the quantization error. When the number of quantization bits is higher, the calculation accuracy will also be higher, but the calculation delay will also increase accordingly. Therefore, the experiment took the above two factors into consideration and tested the changes in accuracy and delay under different T and different numbers of quantization bits. The results are as follows: Fig. 9 shown. Fig. 9 The T in represents the edge deletion rate threshold. When T becomes larger, more edges will be deleted. The gray line represents the T value corresponding to the minimum delay when the accuracy loss is less than 1%. It can be seen that in all data sets, as the delay decreases, the accuracy also shows a downward trend. If the acceptable accuracy loss is limited to less than 1%, the gray vertical line in the figure indicates the optimal T value for delay under the limited accuracy loss. For CORA and CITESEER, 8-bit quantization and T=1 can be used. For PUBMED, ARXIV and REDDIT, 12-bit data is required, and for ARXIV, T needs to be selected not to exceed 0.2, and for REDDIT, T needs to be selected not to exceed 0.5.
[0095] Through the above steps, in terms of software, the high-bit-width floating-point numbers of the graph data are converted into low-bit-width fixed-point numbers by quantizing fixed-point methods, the original graph of the graph data is divided into multiple subgraphs to obtain clustering results, and the node features of the multiple subgraphs are mapped to the RRAM cross array for feature aggregation, and then the edges connecting different subgraphs in the clustering results are deleted by edge deletion to deploy the graph data on hardware; in terms of hardware, the computing module and the control module are configured, wherein the computing module includes: an aggregation core array, a vector combination array and an intermediate cache, and the control module includes: an instruction queue, a graph data decoder and a neighbor cache; through the collaborative design of software and hardware, the acceleration of the graph convolution network GCN calculation is achieved. The present invention improves the parallelism of the calculation, the throughput of the architecture and the resource utilization of the hardware.
[0096] In order to implement the above embodiment, Fig.10 As shown, this embodiment also provides a graph convolutional network software and hardware collaborative acceleration device 10 based on in-memory computing, and the device 10 includes: a software design module 100, a hardware design module 200, and a collaborative acceleration module 300.
[0097] The software design module 100 is used to convert the high-bit-width floating-point numbers of the graph data into low-bit-width fixed-point numbers by quantizing the fixed-point method in the in-memory calculation of the graph data; and use a mapping algorithm based on graph clustering to divide the original graph of the graph data into multiple subgraphs to obtain clustering results, and map the node features of the multiple subgraphs to the RRAM crossbar array for calculating feature aggregation; and, based on the clustering results, use an edge deletion method to delete the edges connecting different subgraphs in the clustering results, so as to deploy the graph data to the hardware through a hardware-oriented deployment algorithm;
[0098] The hardware design module 200 is used to configure the computing module and the control module; wherein the computing module includes: an aggregation core array, a vector combination array and an intermediate cache, wherein the aggregation core array is used to calculate the feature aggregation of the graph data, the vector combination array is used to calculate the vector combination of the graph data, and the intermediate cache is used for data exchange between the aggregation core array and the vector combination array; the control module includes: an instruction queue, a graph data decoder and a neighbor cache, which are respectively used for temporary storage of graph data instructions, format conversion and temporary storage of adjacent vectors;
[0099] The collaborative acceleration module 300 is used to accelerate the calculation of the graph convolutional network GCN through collaborative design of software and hardware.
[0100] According to the in-memory computing-based graph convolution network software-hardware collaborative acceleration device of the embodiment of the present invention, in terms of software, the high-bit-width floating-point numbers of the graph data are converted into low-bit-width fixed-point numbers by quantizing the fixed-point method, the original graph of the graph data is divided into multiple subgraphs to obtain clustering results, and the node features of the multiple subgraphs are mapped to the RRAM cross array for feature aggregation, and then the edges connecting different subgraphs in the clustering results are deleted by edge deletion to deploy the graph data on the hardware; in terms of hardware, the computing module and the control module are configured, wherein the computing module includes: an aggregation core array, a vector combination array and an intermediate cache, and the control module includes: an instruction queue, a graph data decoder and a neighbor cache; through the collaborative design of software and hardware, the acceleration of the graph convolution network GCN calculation is achieved. The present invention improves the parallelism of the calculation, the throughput of the architecture and the resource utilization of the hardware.
[0101] It should be noted that the aforementioned explanation of the embodiment of the graph convolutional network software and hardware collaborative acceleration method based on in-memory computing is also applicable to the graph convolutional network software and hardware collaborative acceleration device based on in-memory computing in this embodiment, and will not be repeated here.
[0102] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0103] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0104] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. A software-hardware collaborative acceleration method for graph convolutional networks based on in-memory computing. It is characterized in that The following steps are involved: Software: In the in-memory calculation of graph data, the high-bit-width floating-point numbers of the graph data are converted into low-bit-width fixed-point numbers by quantizing the fixed-point method; and the original graph of the graph data is divided into multiple subgraphs to obtain clustering results by using a mapping algorithm based on graph clustering, and the node features of the multiple subgraphs are mapped to the RRAM crossbar array for calculating feature aggregation; And, based on the clustering result, using an edge deletion method, deleting the edges connecting different subgraphs in the clustering result, so as to deploy the graph data on the hardware through a hardware-oriented deployment algorithm; Hardware: configure a computing module and a control module; wherein the computing module includes: an aggregation core array, a vector combination array and an intermediate cache, wherein the aggregation core array is used to calculate the feature aggregation of the graph data, the vector combination array is used to calculate the vector combination of the graph data, and the intermediate cache is used for data exchange between the aggregation core array and the vector combination array; the control module includes: an instruction queue, a graph data decoder and a neighbor cache, which are respectively used for temporary storage of instructions of the graph data, format conversion and temporary storage of adjacent vectors; Through the collaborative design of the software and the hardware, the acceleration of graph convolutional network (GCN) calculation is achieved.
2. The method according to claim 1, It is characterized in that The step of converting a high-bit-width floating-point number of the graph data into a low-bit-width fixed-point number by quantizing the fixed-point number in the in-memory calculation of the graph data includes: In the in-memory calculation based on the RRAM, a quantized fixed-point method for graph convolutional networks is used. For the feature data and network weights of each layer of the graph convolutional network, the total quantization bit width of the data is preset. total Finally, according to the data to be quantized, the number of bits Width allocated to the integer and decimal parts is calculated by the following formula int With Width dec : Width dec =Width total -Width int 。 3. The method according to claim 1, It is characterized in that The method uses a mapping algorithm based on graph clustering to divide the original graph of the graph data into multiple subgraphs to obtain clustering results, and maps the node features of the multiple subgraphs to an RRAM crossbar array for calculating feature aggregation, including: Use the Graclus algorithm to cluster the original graph of the graph data into multiple subgraphs, and maximize the connections within the subgraphs and minimize the connections between subgraphs to obtain the clustering results; According to the clustering result, the feature matrix is segmented, and the feature vectors corresponding to the nodes of different subgraphs are mapped to different resistive random access memory cross arrays RRAM; wherein each cross array calculates the feature aggregation of the nodes in the corresponding subgraph.
4. The method according to claim 3, It is characterized in that The method further includes: predefining an edge deletion rate threshold T, and making the number of adjacent edges deleted from each node in the original graph smaller than the product of the original degree of the node and T in the edge deletion algorithm; wherein the value range of T is 0 <T≤1。 5. The method according to claim 4, It is characterized in that The edge deletion algorithm is: Input the nodes of the original graph and information of each of the subgraphs; For each node in the original graph, iteratively select a subgraph that participates in the feature aggregation of this node; In each iteration, a subgraph that covers the most uncovered neighbors of this node is selected to participate in the feature aggregation of this node. When the proportion of covered neighbors of a node is greater than or equal to T, the iteration terminates.
6. The method according to claim 1, It is characterized in that The cluster core array includes a plurality of RRAM-based cluster cores, which are interconnected in the form of an on-chip network to perform matrix-vector multiplication calculations in parallel.
7. The method according to claim 1, It is characterized in that The aggregation core array is divided into a working area and an update area in a ping-pong working form. When the working area performs feature aggregation operations, the update area writes the calculated node features into the RRAM cross array.
8. The method according to claim 1, It is characterized in that The method further includes: using a data forwarding unit to receive data from an adjacent unit or a corresponding RRAM cross array, forwarding data to other adjacent units after merging processing, and organizing in the form of an on-chip network to be controlled by the control module.
9. The method according to claim 1, It is characterized in that The vector combination array includes a waiting queue and multiple MLP cores, which are used to parallelly calculate the vector combination stage of GCN; wherein the MLP core uses the RRAM cross array to calculate the matrix-vector multiplication and uses the lookup table to calculate the nonlinear activation function. When calculating the vector combination stage of GCN, the intermediate data stored in the intermediate cache will first be written into the waiting queue of the vector combination array.
10. A graph convolutional network software and hardware collaborative acceleration device based on in-memory computing, It is characterized in that include: A software design module, for converting a high-bit-width floating-point number of the graph data into a low-bit-width fixed-point number by quantizing the fixed-point method in the in-memory calculation of the graph data; and using a mapping algorithm based on graph clustering to divide the original graph of the graph data into multiple subgraphs to obtain clustering results, and mapping the node features of the multiple subgraphs to an RRAM crossbar array for calculating feature aggregation; And, based on the clustering result, using an edge deletion method, deleting the edges connecting different subgraphs in the clustering result, so as to deploy the graph data on the hardware through a hardware-oriented deployment algorithm; A hardware design module, used to configure a computing module and a control module; wherein the computing module includes: an aggregation core array, a vector combination array and an intermediate cache, wherein the aggregation core array is used to calculate feature aggregation of the graph data, the vector combination array is used to calculate vector combination of the graph data, and the intermediate cache is used for data exchange between the aggregation core array and the vector combination array; the control module includes: an instruction queue, a graph data decoder and a neighbor cache, which are respectively used for temporary storage of instructions of the graph data, format conversion and temporary storage of adjacent vectors; The collaborative acceleration module is used to accelerate the calculation of the graph convolutional network GCN through the collaborative design of the software and the hardware.
Citation Information
Patent Citations
Methods and Apparatus for Incremental Frequent Subgraph Mining on Dynamic Graphs
US20180032587A1
Hardware Accelerator for Convolutional Neural Networks and Method of Operation Thereof
US20180341495A1