Heterogeneous graph neural network reasoning accelerator and reasoning method based on ReRAM
By designing a heterogeneous graph neural network inference accelerator based on ReRAM, using metapath information to build a composite metapath link and perform instance matching, combined with load balancing methods, the problem of difficulty in accelerating heterogeneous graph neural networks in the existing technology is solved, and efficient heterogeneous graph neural network acceleration and memory optimization are achieved.
Patent Information
- Application Number
- CN202510436703.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-25
AI Technical Summary
The existing ReRAM graph neural network accelerators are mainly used to accelerate isomorphic neural networks, but it is difficult to directly apply to complex heteromorphic neural networks. The existing heteromorphic neural network accelerators have challenges in computing and storage performance, especially the low parallelism, which leads to large time and energy overhead.
A heterogeneous graph neural network inference accelerator based on ReRAM is designed, including off-chip memory, chain structure processor, memory controller and resistive random access memory chip. By obtaining metapathic information, a composite metapath chain is constructed, combined with the original graph information for instance matching, and a load balancing method is used to dynamically allocate cross-arrays between engines to achieve load balancing and efficient aggregation.
By fully exploring the parallel computing potential of heterogeneous graph neural networks, the inference speed and energy efficiency are significantly improved, and memory usage is reduced, and the efficient acceleration of heterogeneous graph neural networks is achieved, which is several times higher than that of traditional methods.
Smart Images

Figure CN120373456A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence hardware acceleration, and particularly relates to a heterogeneous graph neural network inference accelerator and an inference method based on ReRAM. Background Art
[0002] With the development of artificial intelligence technology, deep learning has brought about great changes in all aspects of modern social development. For example, the wide application of graph neural networks in social network prediction, recommendation systems, knowledge graphs, etc., has demonstrated their powerful capabilities in processing graph data. However, traditional graph neural networks can only process homogeneous graphs, that is, graph data has only one type of node and edge, while real-world graph data is usually heterogeneous, that is, graph data has multiple types of nodes and edges. Therefore, researchers have proposed heterogeneous graph neural networks, which use meta-paths (i.e., sequences of node types) to define semantic information, and obtain the embedding features of nodes by aggregating the structural and semantic information in the heterogeneous graph in stages, so as to better capture the complex structures and relationships in the heterogeneous graph. Although this type of model has achieved good performance in processing heterogeneous graph data, its complex computational mode also poses higher requirements for computational and storage performance.
[0003] Resistive Random Access Memory (ReRAM) is one of the most promising Processing In Memory (PIM) technologies. It allows memory cells to perform matrix-vector multiplication operations in place in the form of analog circuits, and is often used to accelerate deep neural networks that require a large amount of matrix calculations. However, existing ReRAM graph neural network accelerators only focus on accelerating homogeneous graph neural networks and are difficult to be directly applied to more complex heterogeneous graph neural networks. Summary of the Invention
[0004] In order to solve the above technical problems, the object of the present invention is to provide a heterogeneous graph neural network inference accelerator and an inference method based on ReRAM, which can estimate the load through a sampling-based method, so as to dynamically allocate the number of cross arrays between engines and achieve load balancing.
[0005] The first technical solution adopted by the present invention is: a heterogeneous graph neural network inference accelerator based on ReRAM, including off-chip memory, a chain structure processor, a memory controller, and a resistive random access memory chip. The output end of the off-chip memory is connected to the input end of the memory controller, the output end of the memory controller is connected to the input end of the resistive random access memory chip, the output end of the resistive random access memory chip is connected to the input end of the chain structure processor, the output end of the chain structure processor is connected to the input end of the off-chip memory, and the chain structure processor and the memory controller are connected to each other, where:
[0006] The chain structure processor and the memory controller are used to obtain the original graph and meta-path information for path construction and matching processing, and generate a composite meta-path instance chain;
[0007] The off-chip memory is used to obtain the composite meta-path instance chain for guided aggregation;
[0008] The resistive random access memory chip is used to aggregate the composite meta-path instance chain and output node embedding features.
[0009] Further, the resistive random access memory chip includes an external input / output interface and a plurality of processing arrays. The external input / output interface is electrically connected to the plurality of processing arrays, where:
[0010] The external input / output interface is used for data transmission between the plurality of processing arrays, the chain structure processor, and the memory controller;
[0011] The plurality of processing arrays are used to aggregate the composite meta-path instance chain and generate node embedding features.
[0012] Further, the processing array includes an on-chip buffer module, an intra-meta-path instance aggregation engine, an intra-meta-path aggregation engine, an inter-meta-path aggregation engine, and a first computing unit. The first output end of the on-chip buffer module is connected to the input end of the intra-meta-path instance aggregation engine, the second output end of the on-chip buffer module is connected to the input end of the intra-meta-path aggregation engine, the third output end of the on-chip buffer module is connected to the input end of the inter-meta-path aggregation engine, the intra-meta-path instance aggregation engine, the intra-meta-path aggregation engine, and the inter-meta-path aggregation engine are connected to each other, and the intra-meta-path instance aggregation engine, the intra-meta-path aggregation engine, and the inter-meta-path aggregation engine are all connected to the first computing unit, where:
[0013] The on-chip buffer module is used to obtain the composite meta-path instance chain;
[0014] The in - meta - path instance aggregation engine is used to perform in - meta - path instance aggregation on the composite meta - path instance chain, generating intermediate features within the meta - path instance;
[0015] The intra - meta - path aggregation engine is used to perform intra - meta - path aggregation on the intermediate features within the meta - path instance, obtaining intermediate features within the meta - path;
[0016] The inter - meta - path aggregation engine is used to perform inter - meta - path aggregation on the intermediate features within the meta - path, obtaining node embedding features;
[0017] The first computing unit is used to implement the computation of the aggregation process.
[0018] Furthermore, the in - meta - path instance aggregation engine includes a prefetcher, a scheduler, an input buffer module, a cross - array buffer module, a second computing unit, a functional unit, and an output buffer module. The prefetcher is connected to the scheduler, the scheduler is connected to the input buffer module, the input buffer module is electrically connected to the cross - array buffer module, the output ends of the input buffer module and the cross - array buffer module are both connected to the input end of the second computing unit, the output end of the second computing unit is connected to the input end of the functional unit, and the output end of the functional unit is connected to the input end of the output buffer module, where:
[0019] The prefetcher and the scheduler are used to read the composite meta - path instance chain for pre - fetching and scheduling of node features;
[0020] The input buffer module is used to cache the weights during node feature aggregation;
[0021] The cross - array buffer module is used to cache the node features to be mapped onto the cross - array;
[0022] The second computing unit and the functional unit are used to perform matrix - vector multiplication calculations;
[0023] The output buffer module is used to cache the results of the matrix - vector multiplication calculation and output the intermediate features within the meta - path instance.
[0024] The second technical solution adopted by the present invention is: An inference method for a heterogeneous graph neural network inference accelerator based on ReRAM, including the following steps:
[0025] Obtain meta - path information to construct a composite meta - path chain, obtaining a composite meta - path chain;
[0026] Combine the original graph information to perform meta - path instance matching on the composite meta - path chain, generating a composite meta - path instance chain;
[0027] Based on the load balancing method, aggregate the composite meta-path instance chain and output the node embedding features.
[0028] Furthermore, the step of obtaining meta-path information to construct a composite meta-path chain, resulting in a composite meta-path chain, specifically includes:
[0029] Obtain meta-path information according to a predefined set of meta-paths;
[0030] Perform a first classification on the meta-path information, grouping the meta-paths with the same head node type into the same group and those with different head node types into different groups;
[0031] Based on the result of the first classification, perform a second classification on the meta-path information, grouping the meta-paths with the same second node type into the same group and those with different second node types into different groups;
[0032] Until all node classifications are traversed, merge the same prefixes of the meta-paths within the same group to form a composite meta-path chain, and add a marker to the end node of the meta-path;
[0033] Aggregate the meta-paths within the same group serially on the ReRAM and those between different groups in parallel to obtain a composite meta-path chain.
[0034] Furthermore, the composite meta-path chain includes a value domain, a type domain, a meta-path marker, a shared pointer, and a next node pointer. Among them, the value domain is used to store the node number, the type domain is used to store the node type, the meta-path marker is used to indicate whether the current node is the end point of the meta-path instance, the shared pointer is used to indicate the replaceable node of the current node, and the next node pointer is used to indicate the next node of the meta-path instance chain.
[0035] Furthermore, the step of combining the original graph information to perform meta-path instance matching on the composite meta-path chain to generate a composite meta-path instance chain specifically includes:
[0036] Obtain the original graph information, and select a node with a matching type in the original graph information as the starting point of the instance chain according to the starting node of the composite meta-path chain;
[0037] According to the type of the successor node of the starting point of the composite meta-path chain, find several instance nodes with matching types among the neighbors of the starting point of the instance chain as the successor nodes of the starting point of the instance chain;
[0038] Match the structure in the original graph information according to the type of the successor node of the composite meta-path chain until the composite meta-path chain is completely matched to generate a composite meta-path instance chain.
[0039] Further, for the step of aggregating the composite meta-path instance chain and outputting node embedding features in the load balancing method, it specifically includes:
[0040] Determine the scheduling order of the composite meta-path instance chain by means of depth-first traversal, generate an instruction for aggregating meta-path instances, and determine the node features and node weights of the composite meta-path instance chain according to the instruction;
[0041] Write the node features into the in-meta-path instance aggregation engine, write the node weights into the input buffer of the engine, perform a matrix-vector multiplication operation once to complete the aggregation of node features of a single meta-path instance chain, and output the intermediate features within the meta-path instance;
[0042] Write the intermediate features within the meta-path instance into the in-meta-path aggregation engine, perform a matrix-vector multiplication operation once to complete the aggregation of node features of a single meta-path, and output the intermediate features within the meta-path;
[0043] Write the intermediate features within the meta-path into the inter-meta-path aggregation engine, perform a matrix-vector multiplication operation once to aggregate the intermediate features of each meta-path, and obtain the node embedding features.
[0044] Further, the expression of the load balancing method is specifically as follows:
[0045]
[0046] In the above formula, C r represents the row size of the crossbar array, w and r respectively represent the number of cycles required for each write and read operation, I i represents the length of meta-path i, I i represents the number of instances of meta-path i, M represents the number of meta-paths, W1 represents the in-meta-path instance aggregation engine, W2 represents the in-meta-path aggregation engine, and W3 represents the inter-meta-path aggregation engine.
[0047] The beneficial effects of the method and system of the present invention are as follows: By obtaining meta-path information to construct a composite meta-path chain, the present invention balances the parallelism between meta-paths and the reuse of node features, and then combines the original graph information to perform meta-path instance matching on the composite meta-path chain to generate a composite meta-path instance chain, reducing memory occupancy. Finally, based on the load balancing method, the composite meta-path instance chain is aggregated to output node embedding features, and the load is estimated by a sampling-based method, so as to dynamically allocate the number of crossbar arrays between engines to achieve load balancing. Description of the Drawings
[0048] Figure 1 is a schematic structural diagram of an ReRAM-based heterogeneous graph neural network inference accelerator of the present invention;
[0049] Figure 2 It is a schematic diagram of the steps of an inference method for a heterogeneous graph neural network inference accelerator based on ReRAM according to the present invention;
[0050] Figure 3 It is a schematic diagram of the process of heterogeneous graph neural network inference provided by a specific embodiment of the present invention;
[0051] Figure 4 It is a schematic diagram of the process of hardware inference provided by a specific embodiment of the present invention;
[0052] Figure 5 It is a schematic diagram of the construction of a composite meta-path chain and the matching of meta-path instances provided by a specific embodiment of the present invention;
[0053] Figure 6 It is a schematic diagram of the storage structure provided by a specific embodiment of the present invention;
[0054] Figure 7 It is a schematic diagram of the speedup ratio provided by a specific embodiment of the present invention;
[0055] Figure 8 It is a schematic diagram of the memory overhead reduction rate provided by a specific embodiment of the present invention;
[0056] Figure 9 It is a schematic diagram of the comparison of optimized scheduling provided by a specific embodiment of the present invention;
[0057] Figure 10 It is a schematic diagram of the comparison of load balancing provided by a specific embodiment of the present invention. Detailed implementation manners
[0058] The following further describes the present invention in detail with reference to the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of description and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0059] Traditional graph neural network accelerators can usually only accelerate homogeneous graph neural networks, that is, the graph data is limited to one type of node and edge, and it is difficult and inefficient to directly apply them to heterogeneous graph neural networks. There are a large number of multi-stage node feature aggregation processes in heterogeneous graph neural networks. The large amount of computation and data movement in this process directly leads to large time and energy overheads, which is not conducive to its wide application.
[0060] In addition, although the current heterogeneous graph neural network accelerators adopt the Near Memory Processing (NMP) technology to improve parallelism and reuse of intermediate results, there is still a large amount of data movement, the control logic is complex, the parallelism is not high, and the acceleration potential of the heterogeneous graph neural network has not been fully exploited.
[0061] The embodiments of the present invention aim to fully exploit the parallel computing potential of the heterogeneous graph neural network, utilize the powerful in-situ parallel computing ability of ReRAM, greatly improve the inference speed and energy efficiency of the heterogeneous graph neural network, and promote its wider application and deployment.
[0062] Refer to Figure 1 , the present invention provides a ReRAM-based heterogeneous graph neural network inference accelerator, including off-chip memory, a chain structure processor, a memory controller, and a resistive random access memory chip. The output end of the off-chip memory is connected to the input end of the memory controller, the output end of the memory controller is connected to the input end of the resistive random access memory chip, the output end of the resistive random access memory chip is connected to the input end of the chain structure processor, the output end of the chain structure processor is connected to the input end of the off-chip memory, and the chain structure processor and the memory controller are connected to each other, wherein:
[0063] The chain structure processor and the memory controller are used to obtain the original graph and meta-path information for path construction and matching processing, and generate a composite meta-path instance chain;
[0064] Specifically, the memory controller is mainly used to coordinate the data transmission between the off-chip memory and the ReRAM chip. On the one hand, it ensures that the data requested by the prefetcher can flow into the on-chip buffer of the ReRAM chip in advance, greatly reducing the memory access latency. These requested data include the original graph data, node initial features, composite meta-path instance chains, etc. On the other hand, it also ensures that the results of HGNN inference can be batch written back to the off-chip memory. The chain structure processor is mainly used to generate a composite meta-path chain and implement instance matching of the meta-path.
[0065] The off-chip memory is used to obtain the composite meta-path instance chain for guided aggregation;
[0066] Specifically, the off-chip memory can be any large-capacity random access memory providing a memory interface, such as DRAM, HBM, etc. The off-chip memory is mainly used to store the input and output data of the heterogeneous graph neural network. The input data includes meta-paths, graph structure information, meta-paths, composite meta-path instance chains, and node features. The output data mainly includes the inference results of HGNN, that is, the embedded features of nodes.
[0067] The resistive random access memory chip is used to aggregate the composite meta-path instance chain and output node embedding features.
[0068] Furthermore, the resistive random access memory chip includes an external input / output interface and a plurality of processing arrays. The external input / output interface is electrically connected to the plurality of processing arrays. Among them, the external input / output interface is used for data transmission between the plurality of processing arrays, the chain structure processor, and the memory controller; the plurality of processing arrays are used to aggregate the composite meta-path instance chain and generate node embedding features.
[0069] Specifically, a processing array (Tile) is an independent processing unit that can be used to process the aggregation process of a single node. It includes an intra-meta-path instance aggregation engine (MIAE), an intra-meta-path aggregation engine (IAAE), and an inter-meta-path aggregation engine (IEAE). Different engines work in a pipeline manner. Different processing arrays work in parallel, so they can support the parallel processing at the HGNN node level.
[0070] Furthermore, the processing array includes an on-chip buffer module, an intra-meta-path instance aggregation engine, an intra-meta-path aggregation engine, an inter-meta-path aggregation engine, and a first computing unit. The first output terminal of the on-chip buffer module is connected to the input terminal of the intra-meta-path instance aggregation engine. The second output terminal of the on-chip buffer module is connected to the input terminal of the intra-meta-path aggregation engine. The third output terminal of the on-chip buffer module is connected to the input terminal of the inter-meta-path aggregation engine. The intra-meta-path instance aggregation engine, the intra-meta-path aggregation engine, and the inter-meta-path aggregation engine are connected to each other. The intra-meta-path instance aggregation engine, the intra-meta-path aggregation engine, and the inter-meta-path aggregation engine are all connected to the first computing unit. Among them, the on-chip buffer module is used to obtain the composite meta-path instance chain; the intra-meta-path instance aggregation engine is used to perform intra-meta-path instance aggregation on the composite meta-path instance chain and generate intermediate features within the meta-path instance; the intra-meta-path aggregation engine is used to perform intra-meta-path aggregation on the intermediate features within the meta-path instance and obtain intermediate features within the meta-path; the inter-meta-path aggregation engine is used to perform inter-meta-path aggregation on the intermediate features within the meta-path and obtain node embedding features; the first computing unit is used to implement the calculation of the aggregation process.
[0071] Specifically, the aggregation engine includes three categories, namely the Meta-Path Instance Aggregation Engine (MIAE), the Intra-Meta-Path Aggregation Engine (IAAE), and the Inter-Meta-Path Aggregation Engine (IEAE). Their structures are similar and they work in a pipeline manner. Taking MIAE as an example, it includes a Scheduler, a Prefetcher, on-chip buffers, and a Compute Unit (CU). The Scheduler is used for the prefetch scheduling of node features. It first reads the composite meta-path instance chain on off-chip memory, then performs a depth-first traversal of the composite meta-path instance chain. Each time it encounters the end point of a meta-path instance, a meta-path instance is generated. According to the order of generation of meta-path instances, a processing sequence of locality-aware meta-path instances can be obtained, and then the node features are scheduled according to this sequence to control the aggregation process within the entire engine. The Prefetcher is used on the one hand to prefetch the composite meta-path instance chain from off-chip memory for the Scheduler to generate a scheduling sequence, and on the other hand, it is also used to prefetch node features into the on-chip buffers according to the scheduling sequence to reduce the memory access time.
[0072] The on-chip buffers include three categories: The first category is the Input Buffer, which is used to cache the weights during node feature aggregation. These weights serve as input vectors during matrix-vector multiplication calculations. The second category is the Crossbar Buffer, which is used to cache the node features to be mapped onto the crossbar. These node features form the matrix during matrix-vector multiplication calculations and are written row by row onto the crossbar. The third category is the Output Buffer, which is used to cache the results of matrix-vector multiplication calculations, that is, the aggregation results of multiple features. These results are directly written into the buffers of the next-level engine or directly written back to off-chip memory depending on the engine type.
[0073] The Compute Unit (CU) includes multiple Crossbars at the core and peripheral circuits. The Crossbar is used for matrix-vector multiplication and is the core component for HGNN node feature aggregation. The peripheral circuits include an Analog-to-Digital Converter (ADC), a Digital-to-Analog Converter (DAC), a Sample and Hold circuit (S&H), Special Function Units (SFU), etc., which are used to assist the ReRAM crossbar in completing matrix-vector multiplication.
[0074] In this embodiment, as Figure 4As shown, the chain structure processor and the memory controller process the original graph and meta-paths to generate a chain of composite meta-path instances; the generated chain of composite meta-path instances is written into off-chip memory to guide the subsequent aggregation process; within the meta-path instance aggregation engine, the composite meta-path instance chain and the corresponding initial node features in the off-chip memory are read to perform in-meta-path instance aggregation, generating intermediate features within the meta-path instance, which are written into the in-meta-path aggregation engine; based on the input intermediate features, the in-meta-path aggregation engine further completes the in-meta-path aggregation to obtain intermediate features within the meta-path, and these intermediate features are written into the inter-meta-path aggregation engine; based on the input intermediate features, the inter-meta-path aggregation engine finally completes the inter-meta-path aggregation to obtain the final node embedding features; the inferred node embedding features are written back into the off-chip memory.
[0075] Among them, the in-meta-path instance aggregation engine includes a prefetcher, a scheduler, an input buffer module, a crossbar buffer module, a second computing unit, a functional unit, and an output buffer module. The prefetcher is connected to the scheduler, the scheduler is connected to the input buffer module, the input buffer module is electrically connected to the crossbar buffer module, the output ends of the input buffer module and the crossbar buffer module are both connected to the input end of the second computing unit, the output end of the second computing unit is connected to the input end of the functional unit, and the output end of the functional unit is connected to the input end of the output buffer module. The prefetcher and the scheduler are used to read the chain of composite meta-path instances for prefetch scheduling of node features; the input buffer module is used to cache the weights during node feature aggregation; the crossbar buffer module is used to cache the node features to be mapped onto the crossbar; the second computing unit and the functional unit are used to perform matrix-vector multiplication calculations; the output buffer module is used to cache the results of matrix-vector multiplication calculations and output the intermediate features within the meta-path instance.
[0076] In summary, as Figure 1 shown, the embodiments of the present invention mainly include off-chip memory, a chain structure processor, a memory controller, a processing array (Tile), an aggregation engine, on-chip buffer, and a computing unit (Compute Unit, CU). As Figure 1 shown in (a) of Figure 1 shown in (b) of Figure 1As shown in (c), the aggregation engine internally includes a prefetcher and a scheduler. In addition, the aggregation engine estimates the load through a sampling-based method, so as to dynamically allocate the number of crossbars between the engines to achieve load balancing.
[0077] Referring to Figure 2 , an inference method for a ReRAM-based heterogeneous graph neural network inference accelerator, includes:
[0078] S100. Obtain meta-path information to construct a composite meta-path chain, and obtain a composite meta-path chain;
[0079] Specifically, obtain meta-path information according to a predefined set of meta-paths; perform a first classification on the meta-path information, and divide the meta-paths with the same head node type into the same group, and divide the meta-paths with different head node types into different groups; based on the first classification result, perform a second classification on the meta-path information, and divide the meta-paths with the same second node type into the same group, and divide the meta-paths with different second node types into different groups; until all node classifications are traversed, merge the same prefixes of the meta-paths in the same group to form a composite meta-path chain, and add a mark to the end node of the meta-path; the meta-paths in the same group are serially aggregated on the ReRAM, and the meta-paths between different groups are parallelly aggregated to obtain a composite meta-path chain.
[0080] In this embodiment, the steps of constructing a composite meta-path chain include, as Figure 5 shown in (a), given a set of meaningful meta-paths defined by domain experts, divide the meta-paths with the same head node type into the same group, and divide the meta-paths with different head node types into different groups. Within each group, continue to group based on the second node type, and this process will be repeated n times. Here, n is called the repetition factor, which is a hyperparameter for balancing hardware parallel capabilities and node reuse. Merge the same prefixes of the meta-paths in the same group to form a composite meta-path chain, and add a mark to the end node of the meta-path. The meta-paths in the same group are serially aggregated on the ReRAM, and the meta-paths between different groups are parallelly aggregated.
[0081] S200. Combine the original graph information to perform meta-path instance matching on the composite meta-path chain to generate a composite meta-path instance chain;
[0082] Specifically, obtain the original graph information. According to the starting node of the composite meta-path chain, select a node with a matching type in the original graph information as the starting point of the instance chain; according to the successor node type of the starting point of the composite meta-path chain, find several instance nodes with matching types among the neighbors of the starting point of the instance chain as the successor nodes of the starting point of the instance chain; according to the successor node type of the composite meta-path chain, perform matching in the structure of the original graph information until the composite meta-path chain is completely matched, and generate a composite meta-path instance chain.
[0083] Among them, the composite meta-path chain includes a value domain, a type domain, a meta-path flag, a shared pointer, and a next node pointer. Among them, the value domain is used to store the node number, the type domain is used to store the node type, the meta-path flag is used to indicate whether the current node is the end point of the meta-path instance, the shared pointer is used to indicate the replaceable node of the current node, and the next node pointer is used to indicate the next node of the meta-path instance chain.
[0084] In this embodiment, as Figure 5 shown in (b) of, according to the starting node of the composite meta-path chain, select a node with a matching type in the graph as the starting point of the instance. Subsequently, according to the successor node type of the starting point of the composite meta-path chain, find several instance nodes with matching types among the neighbors of the starting point of the instance chain as the successor nodes of the starting point of the instance chain. Continue to perform matching in the graph structure according to the successor node type of the composite meta-path chain until the composite meta-path chain is completely matched. At this time, a composite meta-path instance chain will be obtained on the graph. Store the obtained composite meta-path instance chain in the off-chip memory. This composite instance chain can generate a scheduling order through subsequent processing.
[0085] Further as Figure 6 shown, where (a) represents a traditional instance storage example, Figure 6 and (b) in is the specific composition of the chain structure of the embodiment of the present invention. Among them, a single node includes: a value domain (4Bytes), used to store the node number. A type domain (1Byte), used to store the node type. A meta-path flag (1bit), used to indicate whether the current node is the end point of the meta-path instance, 0 indicates that it is not the end point of the meta-path instance, and 1 indicates that it is the end point of the meta-path instance. A shared pointer (4Bytes), used to indicate the replaceable node of the current node. A next node pointer (4Bytes), used to indicate the next node of the meta-path instance chain. Although this method occupies a little more space than the traditional node unit, it will reduce a large number of repeatedly stored nodes, thereby macroscopically reducing the space occupied by the meta-path instance.
[0086] S300. Based on the load balancing method, perform aggregation processing on the composite meta-path instance chain and output the node embedding features.
[0087] Specifically, the scheduling order of the composite metapath instance chain is determined by means of depth-first traversal to generate an instruction for aggregating metapath instances. The node features and node weights of the composite metapath instance chain are determined according to the instruction. The node features are written into the in-instance aggregation engine of the metapath instance, and the node weights are written into the input buffer of the engine. A matrix-vector multiplication operation is performed once to complete the aggregation of the node features of a single metapath instance chain, and the intermediate features within the metapath instance are output. The intermediate features within the metapath instance are written into the intra-metapath aggregation engine, and a matrix-vector multiplication operation is performed once to complete the aggregation of the node features of a single metapath, and the intermediate features within the metapath are output. The intermediate features within the metapath are written into the inter-metapath aggregation engine, and a matrix-vector multiplication operation is performed once to aggregate the intermediate features of each metapath to obtain the node embedding features.
[0088] In this embodiment, the steps of structure aggregation (corresponding to Figure 4 sequence numbers 3 and 4 in
[0089] include. First, the prefetcher fetches a composite metapath instance chain from off-chip memory. Then, the scheduler determines the scheduling order according to this composite metapath instance chain by means of depth-first traversal. Starting from the starting node, every time a node with a metapath marked as "1" is traversed, an instruction for aggregating metapath instances is generated. This instruction controls the prefetcher to prefetch the node features and their weights of this metapath instance chain from off-chip memory. The further prefetched node features are written row by row into the Metapath Intra-instance Aggregation Engine (MIAE), and the corresponding weight vectors are written into the input buffer of the engine. When all the node features on a single metapath instance chain are written onto the ReRAM crossbar array of the MIAE, the ReRAM crossbar array performs a matrix-vector multiplication operation once to complete the aggregation of the node features of a single metapath instance chain. The intermediate features after the aggregation of a single metapath instance chain are written into the next-level engine, that is, the Intra-metapath Aggregation Engine (IAAE). When all the metapath instances in a composite metapath instance chain have completed aggregation, the ReRAM crossbar array of the IAAE performs a matrix-vector multiplication operation once to complete the aggregation of the node features of a single metapath. This intermediate feature is also written into the next-level engine, that is, the Inter-metapath Aggregation Engine (IEAE). Finally, through the processing of the MIAE and IAAE, the structure aggregation of a single node is completed. In addition, the structure aggregations on different composite metapath instance chains are performed in parallel. Figure 4The step of item 5) includes that after the structure aggregation is completed, the intermediate features of each meta-path have been written on the ReRAM crossbar array of IEAE. Then, IEAE will perform a matrix-vector multiplication operation to aggregate the intermediate features of each meta-path and obtain the final embedded features. Through the processing of IEAE, the final embedded features of a single node are obtained. These features will then flow into the output buffer and finally be written back to off-chip memory. Thus, the inference acceleration of the heterogeneous graph neural network is completed.
[0090] Finally, it should also be noted that the expression of the load balancing method is specifically as follows:
[0091]
[0092] In the above formula, C r represents the row size of the crossbar array, w and r respectively represent the number of cycles required for each write and read operation, L i represents the length of meta-path i, I i represents the number of instances of meta-path i, M represents the number of meta-paths, W1 represents the aggregation engine within the meta-path instance, W2 represents the aggregation engine within the meta-path, and W3 represents the aggregation engine between meta-paths. According to the estimated load, the number of crossbar arrays is proportionally allocated to each engine.
[0093] In summary, as Figure 3 shown, taking a heterogeneous graph with three types of nodes as an example. Feature projection projects node features of different types onto the same dimension as the initial features of the nodes. According to the obtained meta-path instance chain by matching and the initial features of the nodes, structure aggregation (corresponding to Figure 3 items 2 and 3) is performed to obtain the intermediate features of a single meta-path. According to the intermediate features of each meta-path, semantic aggregation (corresponding to Figure 3 item 4) is performed to obtain the final embedded features of the nodes.
[0094] Finally, the embodiments of the present invention are explained and illustrated in combination with an actual engineering case. First, ReRAM is configured. The simulator is configured such that 1 accelerator contains 32 processing arrays (Tiles), each processing array contains 24 computing units (CUs), each computing unit contains 32 crossbars, the specification of the crossbar is 128×128, and the number of bits of each cell is 2. The read and write latencies of the ReRAM cells are 29.31 ns and 50.88 ns respectively. The off-chip memory is configured as an HBM memory with a capacity of 128 GB and a bandwidth of 900 GB / s. The repetition factor n is set to 1. Furthermore, 4 state-of-the-art inference schemes for heterogeneous graph neural networks are selected, namely Intel Xeon Gold 5117 CPU, Nvidia Tesla V100 GPU, REFLIP, and MetaNMP. 3 representative heterogeneous graph neural networks are selected, namely MAGNN, HAN, and SHGNN. 5 representative datasets are DBLP (DP), IMDB (IB), LastFM (LF), OGB-MAG (OM), and OAG (OG).
[0095] Furthermore, data preprocessing is carried out. First, the meta-paths of each dataset are grouped, and a composite meta-path chain is constructed for each group. Then, according to the composite meta-path chain, meta-path instance matching is performed on the original graph data to construct a composite meta-path instance chain. Finally, the composite meta-path instance chain is written into the off-chip memory for use by the heterogeneous graph neural network during subsequent inference. Since it is stored in a compressed chain structure, compared with the traditional storage method, the repeated node storage is reduced, so certain effects are achieved on each meta-path instance of each dataset. The memory overhead reduction rate of the embodiments of the present invention has reduced the memory occupancy by 47.86%.
[0096] Finally, according to the pre-stored meta-path instance chain, the scheduler starts a depth-first traversal from the starting node. Each time it encounters a node with the end marker of the meta-path instance being 1, it generates a meta-path instance. The order in which the meta-path instances are generated is the scheduling order of the locality-aware meta-path instances. This locality-aware scheduling order reduces the loading of node features by an average of 2.81 times and achieves a 2.17-fold performance improvement compared to the scheduling order without the guidance of the chain structure. Subsequently, the scheduler controls the prefetcher to prefetch the initial features of the nodes and the corresponding aggregation weights from off-chip memory. The aggregation engine within the meta-path instance starts to perform the aggregation of the meta-path instance. By using the aggregation weight vector as the input to the ReRAM crossbar array, the node initial features are mapped row by row onto the ReRAM crossbar array. The ReRAM crossbar array performs a matrix-vector multiplication operation in the form of an analog circuit to obtain the intermediate features of the meta-path instance. These intermediate features are written into the crossbar buffer of the next-level aggregation engine. Then, the next-level aggregation engine, i.e., the intra-meta-path aggregation engine, continues to prefetch the corresponding aggregation weights from off-chip memory as the input to the ReRAM crossbar array at this level. The intermediate features in the crossbar buffer are written onto the ReRAM crossbar array at this level. After the engine at this level performs a matrix-vector multiplication operation, it can obtain the intermediate features of a single meta-path. These intermediate features are written into the buffer of the last-level engine, i.e., the buffer of the inter-meta-path aggregation engine. The inter-meta-path aggregation engine continues to prefetch the weights for semantic aggregation from off-chip memory and writes the intermediate features in the crossbar cache row by row into the crossbar array within the engine. By performing a matrix-vector multiplication operation, the inter-meta-path aggregation engine aggregates the semantic information of each meta-path to obtain the final node embedding features. Finally, these node embedding features are written into the output cache and further batch-written back to off-chip memory. Since this accelerator makes full use of the in-memory computing technology to significantly reduce data movement, exploits the parallelism at the meta-path instance level and meta-path level in the heterogeneous graph neural network, enhances the reuse of nodes through locality-aware scheduling, and achieves load balancing using a sampling-based method, it can achieve good acceleration effects. The method of the embodiment of the present invention accelerates by an average of 1320 times, 128.29 times, 15.35 times, and 6.01 times compared to CPU, GPU, REFLIP, and MetaNMP respectively.
[0097] Therefore, by making full use of the ability of ReRAM to perform matrix multiplication in situ, the control logic of the embodiment of the present invention is simpler, the data movement is reduced, and the acceleration ratio and energy efficiency are greatly improved. Existing ReRAM accelerators can only accelerate homogeneous graph neural networks. Compared with them, our ReRAM accelerator fully exploits the parallelism at the meta-path instance level and meta-path level unique to heterogeneous graph neural networks, so the acceleration efficiency has been significantly improved, as Figure 7As shown below. After that, we designed a chained storage structure for meta-path instances, which reduced the storage of a large number of duplicate nodes compared with the traditional method, and thus significantly reduced the memory occupancy, as shown in Figure 8 As shown below. In addition, through locality-aware scheduling, we enhanced the reuse of node features and reduced the writing of node features in the ReRAM crossbar array, as shown in Figure 9 As shown below, further improving the efficiency. Finally, we also reasonably designed the pipelines of the three engines and optimized the utilization rate of the ReRAM crossbar array through load balancing, as shown in Figure 10 As shown below, thereby significantly improving the performance of the accelerator.
[0098] The content in the above method embodiments is applicable to the present system embodiment. The functions specifically implemented by the present system embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0099] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A ReRAM-based heterogeneous graph neural network inference accelerator, characterized in that It includes off-chip memory, a chain structure processor, a memory controller, and a resistive random access memory chip. The output end of the off-chip memory is connected to the input end of the memory controller. The output end of the memory controller is connected to the input end of the resistive random access memory chip. The output end of the resistive random access memory chip is connected to the input end of the chain structure processor. The output end of the chain structure processor is connected to the input end of the off-chip memory. The chain structure processor and the memory controller are interconnected. Among them: The chain structure processor and the memory controller are used to obtain the original graph and meta-path information for path construction and matching processing, and generate a composite meta-path instance chain; The off-chip memory is used to obtain the composite meta-path instance chain for guiding aggregation; The resistive random access memory chip is used to perform aggregation processing on the composite meta-path instance chain and output node embedding features.
2. The inference accelerator for heterogeneous graph neural network based on ReRAM according to claim 1, wherein The resistive random access memory chip includes an external input / output interface and a plurality of processing arrays. The external input / output interface is electrically connected to the plurality of processing arrays. Among them: The external input / output interface is used to perform data transmission between the plurality of processing arrays, the chain structure processor, and the memory controller; The plurality of processing arrays are used to perform aggregation processing on the composite meta-path instance chain and generate node embedding features.
3. The inference accelerator for heterogeneous graph neural network based on ReRAM according to claim 2, wherein The processing array includes an on-chip buffer module, an intra-meta-path instance aggregation engine, an intra-meta-path aggregation engine, an inter-meta-path aggregation engine, and a first computing unit. The first output end of the on-chip buffer module is connected to the input end of the intra-meta-path instance aggregation engine. The second output end of the on-chip buffer module is connected to the input end of the intra-meta-path aggregation engine. The third output end of the on-chip buffer module is connected to the input end of the inter-meta-path aggregation engine. The intra-meta-path instance aggregation engine, the intra-meta-path aggregation engine, and the inter-meta-path aggregation engine are interconnected. The intra-meta-path instance aggregation engine, the intra-meta-path aggregation engine, and the inter-meta-path aggregation engine are all interconnected with the first computing unit. Among them: The on-chip buffer module is used to obtain the composite meta-path instance chain; The intra-meta-path instance aggregation engine is used to perform intra-meta-path instance aggregation on the composite meta-path instance chain and generate intermediate features within the meta-path instance; The intra-meta-path aggregation engine is used to perform intra-meta-path aggregation on the intermediate features within the meta-path instance to obtain intermediate features within the meta-path; The inter-meta-path aggregation engine is used to perform inter-meta-path aggregation on the intermediate features within the meta-path to obtain node embedding features; The first computing unit is used to implement the calculation of the aggregation process.
4. The inference accelerator of a heterogeneous graph neural network based on ReRAM according to claim 3, wherein The aggregation engine within the meta-path instance includes a prefetcher, a scheduler, an input buffer module, a crossbar buffer module, a second computing unit, a functional unit, and an output buffer module. The prefetcher is interconnected with the scheduler, the scheduler is interconnected with the input buffer module, the input buffer module is electrically connected to the crossbar buffer module, the output ends of the input buffer module and the crossbar buffer module are both connected to the input end of the second computing unit, the output end of the second computing unit is connected to the input end of the functional unit, and the output end of the functional unit is connected to the input end of the output buffer module, where: The prefetcher and the scheduler are used to read the composite meta-path instance chain for prefetch scheduling of node features; The input buffer module is used to cache the weights during node feature aggregation; The crossbar buffer module is used to cache the node features to be mapped onto the crossbar; The second computing unit and the functional unit are used to perform matrix-vector multiplication calculations; The output buffer module is used to cache the results of matrix-vector multiplication calculations and output the intermediate features within the meta-path instance.
5. An inference method for an inference accelerator of a heterogeneous graph neural network based on ReRAM, characterized in that, It includes the following steps: Obtain meta-path information to construct a composite meta-path chain, resulting in a composite meta-path chain; Combine the original graph information to perform meta-path instance matching on the composite meta-path chain to generate a composite meta-path instance chain; Based on the load balancing method, perform aggregation processing on the composite meta-path instance chain and output node embedding features.
6. The inference method of a heterogeneous graph neural network inference accelerator based on ReRAM according to claim 5, characterized in that, The step of obtaining meta-path information to construct a composite meta-path chain, resulting in a composite meta-path chain specifically includes: Obtain meta-path information according to a predefined set of meta-paths; Perform the first classification on the meta-path information, grouping the meta-paths with the same head node type of the meta-path information into the same group, and those with different head node types into different groups; Based on the result of the first classification, perform the second classification on the meta-path information, grouping the meta-paths with the same second node type of the meta-path information into the same group, and those with different second node types into different groups; Until all node classifications are traversed, merge the same prefixes of the meta-paths within the same group to form a composite meta-path chain, and add a marker to the end node of the meta-path; The meta-paths within the same group are aggregated serially on the ReRAM, and the meta-paths between different groups are aggregated in parallel to obtain a composite meta-path chain.
7. The inference method of a heterogeneous graph neural network inference accelerator based on ReRAM according to claim 6, wherein The composite meta-path chain includes a value domain, a type domain, a meta-path marker, a shared pointer, and a next node pointer. Among them, the value domain is used to store the node number, the type domain is used to store the node type, the meta-path marker is used to indicate whether the current node is the end point of the meta-path instance, the shared pointer is used to indicate the replaceable node of the current node, and the next node pointer is used to indicate the next node of the meta-path instance chain.
8. The inference method of a heterogeneous graph neural network inference accelerator based on ReRAM according to claim 7, characterized in that, The step of combining the original graph information to perform meta-path instance matching on the composite meta-path chain to generate a composite meta-path instance chain specifically includes: Obtain the original graph information. According to the starting node of the composite meta-path chain, select a node with a matching type in the original graph information as the starting point of the instance chain; According to the type of the successor node of the starting point of the composite meta-path chain, find several instance nodes with matching types among the neighbors of the starting point of the instance chain as the successor nodes of the starting point of the instance chain; According to the type of the successor node of the composite meta-path chain, perform matching on the structure in the original graph information until the composite meta-path chain is completely matched to generate a composite meta-path instance chain.
9. The inference method of a heterogeneous graph neural network inference accelerator based on ReRAM according to claim 8, characterized in that, The step of aggregating the composite meta-path instance chain based on the load balancing method and outputting the node embedding features specifically includes: Determine the scheduling order of the composite meta-path instance chain in a depth-first traversal manner, generate an instruction for aggregating the meta-path instances, and determine the node features and node weights of the composite meta-path instance chain according to the instruction; Write the node features into the in-meta-path instance aggregation engine, write the node weights into the input buffer of the engine, and perform a matrix-vector multiplication operation to complete the aggregation of the node features of a single meta-path instance chain and output the intermediate features within the meta-path instance; Write the intermediate features within the meta-path instance into the in-meta-path aggregation engine, perform a matrix-vector multiplication operation to complete the aggregation of the node features of a single meta-path, and output the intermediate features within the meta-path; Write the intermediate features within the meta-path into the between-meta-path aggregation engine, perform a matrix-vector multiplication operation to aggregate the intermediate features of each meta-path, and obtain the node embedding features.
10. The inference method of a heterogeneous graph neural network inference accelerator based on ReRAM according to claim 9, wherein, The expression of the load balancing method is specifically as follows: In the above formula, C r represents the row size of the crossbar array, w and r respectively represent the number of cycles required for each write and read operation, L i represents the length of the meta-path i, I i represents the number of instances of the meta-path i, M represents the number of meta-paths, W1 represents the aggregation engine within the meta-path instance, W2 represents the aggregation engine within the meta-path, and W3 represents the aggregation engine between the meta-paths.