GCN reasoning acceleration method based on multiple data streams and high-bandwidth memory
The GCN inference acceleration method uses multi-data flow engines and high-bandwidth memory to optimize GCN operations, addressing data set variability and improving performance and energy efficiency.
Patent Information
- Application Number
- CN202510382015.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-15
AI Technical Summary
The existing technology is difficult to effectively deal with the different sizes and sparseness of graph structure data sets, resulting in inefficient computing in the GCN inference process and serious memory access latency and bandwidth bottlenecks.
The GCN inference acceleration method based on multi-data streams and high-bandwidth memory is designed, and the multi-data stream computing engine, adaptive selector, multi-channel high-bandwidth memory mapping strategy, task pipeline optimization and hybrid fixed-point quantization strategy are used to optimize the computing and memory access of the GCN model.
With almost no loss of accuracy, the GCN inference speed is significantly improved, the calculation amount and parameter amount are reduced, and the GCN accelerator with high flexibility and low energy consumption is achieved.
Smart Images

Figure CN120317366A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of graph convolutional networks, and in particular, to a GCN inference acceleration method based on multi-data streams and high-bandwidth memory. Background Art
[0002] In the real world, graph-structured data has been widely used in fields such as social networks, transportation networks, molecular networks, and e-commerce networks. Graph convolutional neural networks (GCNs) have become an effective method for processing these non-Euclidean graph data. However, graph datasets with different scales and sparsities, as well as the dependence of the data flow pattern of GCN calculations on the graph structure, make accelerating GCN inference increasingly valuable and challenging.
[0003] In order to effectively address the characteristics of graph-structured datasets in the real world with different scales and sparsities, and to solve the above challenges faced in the process of GCN acceleration. It is necessary to design a GCN accelerator with diverse computing modes and multi-data stream modes to flexibly handle the complex diversity of input graph-structured data. In addition, due to the frequent memory accesses in GCN calculations, the GCN accelerator also needs to make full use of high-bandwidth memory to alleviate the memory latency and bandwidth bottlenecks caused by a large amount of data transfer and frequent memory access behaviors. Summary of the Invention
[0004] To achieve the above object, this application provides the following technical solutions:
[0005] According to the first aspect of the present invention, the present invention claims a GCN inference acceleration method based on multi-data streams and high-bandwidth memory, including:
[0006] Configuring a computing engine for the GCN model that supports multi-data streams, aggregation priority, and combination priority order;
[0007] An adaptive selector for a multi-data stream computing engine based on a decision tree to select the best computing engine according to the characteristics of various datasets and the characteristics of the GCN model;
[0008] Configuring a multi-channel high-bandwidth memory HBM and designing a mapping strategy for the multi-channels of the multi-channel high-bandwidth memory to data and computing units;
[0009] Performing task pipeline optimization when performing feature aggregation operations and combination operations in each layer of the GCN model;
[0010] Adopting a mixed fixed-point quantization strategy for hierarchical classification of the GCN model, and obtaining performance evaluation results of the GCN inference acceleration method through performance evaluation experiments on various datasets.
[0011] Further, the configuration supports a computing engine for multi-data streams, aggregation priority, and combination priority order, and further includes:
[0012] The GCN model includes an aggregation stage and a combination stage;
[0013] When the data set satisfies the condition that the GCN reduces the feature dimension after executing the combination stage, the combination stage is executed first, and then the aggregation stage is executed;
[0014] The activation functions used in the aggregation stage and the combination stage satisfy the idempotent property;
[0015] In the GCN model, the core operations of the aggregation and combination stages are matrix multiplications;
[0016] The data flow patterns of the matrix multiplications include four types: inner product, outer product, row-aware product, and column-aware product;
[0017] Based on the characteristics and advantages of the data flows of the inner product, outer product, row-aware product, and column-aware product, general matrix multiplication and sparse-dense matrix multiplication computing engine cores are designed for the GCN model respectively.
[0018] Further, the adaptive selector of the multi-data stream computing engine based on the decision tree selects the best computing engine according to the characteristics of various data sets and the characteristics of the GCN model, and further includes:
[0019] The selector of the multi-data stream computing engine based on the decision tree uses the number of vertices N, the number of non-zero elements of matrix A, the number of non-zero elements of matrix X, the feature matrix F1, the weight matrix W, and the node degree D as the characteristics of the data set and uses them as the input data of the decision tree;
[0020] A decision tree model is trained using data sets with various different characteristics. In actual application scenarios, the decision tree model is used to select the most suitable data stream computing engine.
[0021] Further, the configuration is a multi-channel high-bandwidth memory HBM, and a mapping strategy for the multi-channel of the multi-channel high-bandwidth memory to the data and computing units is designed, and further includes:
[0022] A partition mapping mechanism is adopted to group the graph data rows in the convolution operation of the GCN model based on the degree of vertices to achieve reasonable partitioning of the graph;
[0023] By mapping the partition of the graph data to the multi-channels of the HBM, and mapping the multi-channels of the HBM to the data stream computing engine cores;
[0024] The partitioned graph data is stored in the on-chip memory, and the synergistic effect of the pipeline and the mapping mechanism is utilized;
[0025] Use the compilation guidance function of HLS to implement the mapping of multiple channels of HBM.
[0026] Furthermore, when performing feature aggregation operations and combination operations in each layer of the GCN model, the execution of task pipeline optimization further includes:
[0027] Based on the pipeline task scheduling optimization strategy of the greedy algorithm, optimize the pipeline scheduling between the feature aggregation operation and the combination operation of GCN;
[0028] Based on the strategy of fine-grained data partitioning, and introducing the deep pipeline and data flow parallel processing methods, adopt a fine-grained pipeline design on each computing node of the CPU-FPGA to achieve the pipeline parallelism of the data transmission and calculation parts in the GCN model.
[0029] Furthermore, when using the hybrid fixed-point quantization strategy for hierarchical classification of the GCN model and obtaining the performance evaluation results of the GCN inference acceleration method through performance evaluation experiments on multiple data sets, it further includes:
[0030] Convert the original floating-point matrix multiplication into a low-precision integer multiplication operation;
[0031] Adopt a hybrid fixed-point quantization strategy based on hierarchical and classification. For the feature matrix of the first layer of the GCN model, use a high bit width, and for the feature matrix of the second layer, use a low bit width;
[0032] Perform performance tests on various data flow engines based on different data sets on the Vanilla GCN model.
[0033] Furthermore, when performing performance tests on various data flow engines based on different data sets on the Vanilla GCN model, it further includes:
[0034] Measure and compare the execution times of various data flow engines of the GCN model on different data sets, and obtain the efficiency of the data flow engines when processing different data sets;
[0035] The execution times of different execution orders of the aggregation and combination stages vary on different data sets;
[0036] When the aggregation stage precedes the combination stage in execution, if the dimension of the result matrix after performing the aggregation operation decreases more, then the execution order with the aggregation stage preceding the combination stage will reduce the execution time;
[0037] When the combination stage precedes the aggregation stage in execution, if the dimension of the result matrix after performing the combination operation decreases more, then the execution order with the combination stage preceding the aggregation stage will result in less execution time.
[0038] This application relates to the technical field of graph convolutional networks, and particularly to a GCN inference acceleration method based on multi-data streams and high-bandwidth memory. By combining a multi-data stream computing engine and an effective aggregation combination stage sequence, it explores the PCs mapping strategy of multi-channel HBM, a hierarchical classification hybrid fixed-point quantization strategy, and designs an adaptive selector for the multi-data stream computing engine based on a decision tree. According to the characteristics of the dataset and the characteristics of the GCN model, it selects the most suitable data stream computing engine to adapt to graph structure datasets of different scales and sparsities in the real world, improving the inference speed of GCN. The GCN inference acceleration method proposed by the present invention reduces the computational amount and the number of parameters of the GCN model with almost no loss of accuracy, and can obtain an efficient GCN accelerator, which has the advantages of high flexibility and low energy consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a flowchart of a GCN inference acceleration method based on multi-data streams and high-bandwidth memory claimed in the embodiments of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0041] The terms "first", "second", and "third" in this application are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined. All directional indications (such as up, down, left, right, front, back...) in the embodiments of this application are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0042] References to "embodiments" in this specification mean that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0043] According to the first embodiment of the present invention, the present invention claims a GCN inference acceleration method based on multi-data streams and high-bandwidth memory. Referring to Figure 1 , it includes:
[0044] Configure a computing engine for the GCN model that supports multi-data streams, aggregation priority, and combination priority order;
[0045] An adaptive selector for the multi-data stream computing engine based on a decision tree selects the best computing engine according to the characteristics of various data sets and the characteristics of the GCN model;
[0046] Configure multi-channel high-bandwidth memory HBM and design a mapping strategy for the multi-channels of the multi-channel high-bandwidth memory to data and computing units;
[0047] Perform task pipeline optimization when performing feature aggregation operations and combination operations in each layer of the GCN model;
[0048] Adopt a hybrid fixed-point quantization strategy for hierarchical classification of the GCN model, and obtain the performance evaluation results of the GCN inference acceleration method through performance evaluation experiments on various data sets.
[0049] Among them, in this embodiment, a basic Vanilla GCN model is constructed using the PyTorch-Geometric (PyG) framework, and GCN networks with different numbers of layers from 1 to 10 are implemented. In the CPU environment, we tested the performance of these GCN models with different numbers of layers on multiple data sets (including Cora, Citeseer, Pubmed, Nell).
[0050] In most cases, when the number of layers of the GCN model is set to 2, its accuracy reaches the peak. However, as the number of layers of the model further increases, the accuracy shows a downward trend. This phenomenon can be explained by the over-smoothing problem, that is, as the number of GCN layers increases, the receptive field of the nodes continuously expands, resulting in the neighbor information aggregated by each node gradually becoming the same, thus causing over-smoothing.
[0051] Based on the above experimental results and analysis, the GCN inference accelerator designed in this embodiment adopts a two-layer GCN model as the benchmark architecture. This choice aims to reduce the performance loss caused by the over-smoothing problem while maintaining a high accuracy through a moderate network depth.
[0052] In GCN, the core idea of graph convolution is to generate a comprehensive feature representation of a node by aggregating the features of all its neighbor nodes and the node's own features. The GCN model realizes the step-by-step extraction and integration of high-level comprehensive information of each node by stacking multiple graph convolution layers.
[0053] The overall architecture of the multi-data-stream computing engine and the GCN model accelerator based on HBM mapping designed in this embodiment. The overall architecture of this GCN accelerator is based on a CPU-FPGA hybrid architecture. Among them, the CPU host side is responsible for preprocessing graph data, partitioning graph data, and scheduling and mapping graph tasks. The FPGA side focuses on executing the core computing tasks of GCN and deploys the multi-data-stream computing engine core of the GCN model.
[0054] This multi-data-stream computing engine core includes computing cores based on inner product, outer product, row-aware product, and column-aware product to adapt to GCN for processing various graph data sets with different feature types. At the same time, in order to further solve the bandwidth bottleneck, this design makes full use of the multi-channel high bandwidth of FPGA based on HBM. By efficiently mapping graph tasks to multi-channel HBM, the bandwidth and efficiency of memory access are effectively improved.
[0055] The CPU host side and the FPGA side are connected through a PCIe interface to ensure efficient data exchange and communication between the two ends. The application running on the CPU host side is implemented based on OpenCL programming and is used to manage the execution and scheduling of the engine core on the FPGA to achieve efficient cooperation between the CPU and the FPGA.
[0056] Furthermore, the configuration supports a computing engine with multi-data-stream, aggregation priority, and combination priority order, and further includes:
[0057] The GCN model includes an aggregation stage and a combination stage;
[0058] When the data set satisfies the condition of reducing the feature dimension after GCN executes the combination stage, the combination stage is executed first, and then the aggregation stage;
[0059] The activation functions used in the aggregation stage and the combination stage satisfy the idempotent property;
[0060] In the GCN model, the core operations in the aggregation and combination stages are matrix multiplications;
[0061] The data flow patterns of the matrix multiplication include four types: inner product, outer product, row-aware product, and column-aware product;
[0062] Based on the characteristics and advantages of the data flows of inner product, outer product, row-aware product, and column-aware product, a general matrix multiplication and a sparse-dense matrix multiplication computational engine core are respectively designed for the GCN model.
[0063] Among them, in this embodiment, the overall structure of the GCN accelerator engine core is composed of components such as the SpDeGEMM core, GEMM core, buffer, and scheduler. Among them, the SpDeGEMM core serves as an aggregation processing unit and is usually used to be responsible for aggregation operation calculations. The GEMM core serves as a combination processing unit and is used to be responsible for combination operation calculations. In addition, in order to adapt to different graph data sets, the ratio of the number of units of the SpDeGEMM core and the GEMM core is determined according to the characteristics of the graph data set being processed. The scheduler is responsible for controlling task allocation and the transmission of intermediate results between the SpDeGEMM core and the GEMM core. At the same time, in order to optimize data transmission efficiency, corresponding buffers are set between the SpDeGEMM core and the GEMM core.
[0064] The SpDeGEMM core and the GEMM core support multiple data flow patterns, including inner product, outer product, row-aware product, and column-aware product. The characteristics, advantages, and disadvantages of these data flow patterns have been described in the second paragraph. Combining the combination order of aggregation priority and combination optimization, a total of eight different data flow patterns can be formed, namely AC-IP, AC-OP, AC-RP, AC-CP, CA-IP, CA-OP, CA-RP, and CA-CP. The selection of these patterns depends on the characteristics of the data set and the characteristics of the GCN model.
[0065] In the aggregation-combination order method, when the SpDeGEMM core completes a part of the task, it will send the task ID and partial results to the scheduler together. Once all the SpDeGEMM cores have completed the relevant tasks, the scheduler will assign the corresponding combination tasks to the GEMM core until all the combination tasks are completed. On the contrary, in the combination-aggregation order method, the scheduler performs the operations in the reverse order, first processing the combination tasks and then the aggregation tasks. This flexible task scheduling strategy enables the GCN accelerator to adapt to different graph data characteristics and achieve efficient computing performance.
[0066] Furthermore, the adaptive selector of the multi-data flow computational engine based on the decision tree selects the best computational engine according to the characteristics of various data sets and the characteristics of the GCN model, and further includes:
[0067] A selector for a multi-data stream computing engine based on a decision tree uses the number of vertices N, the number of non-zero elements of matrix A, the number of non-zero elements of matrix X, the feature matrix F1, the weight matrix W, and the node degree D as the features of the data set, and serves as the input data of the decision tree.
[0068] Use data sets with various different features to train the decision tree model. In an actual application scenario, use the decision tree model to select the most suitable data stream computing engine.
[0069] Among them, in this embodiment, Algorithm 1 represents a selector for a multi-data stream computing engine based on a decision tree. This algorithm analyzes the characteristics of the input graph data and uses the decision logic of the decision tree to quickly and accurately determine the most suitable data stream computing engine for the current data set.
[0070]
[0071]
[0072]
[0073] Furthermore, configuring the multi-channel high-bandwidth memory HBM and designing the mapping strategy of the multi-channel of the multi-channel high-bandwidth memory with the data and computing units further includes:
[0074] Adopt a partition mapping mechanism to group the graph data rows in the convolution operation of the GCN model based on the degree of vertices to achieve reasonable partitioning of the graph.
[0075] Through the mapping of the partitions of the graph data with the multi-channels of the HBM and the mapping of the multi-channels of the HBM with the data stream computing engine cores;
[0076] Store the partitioned graph data in the on-chip memory and utilize the synergistic effect of the pipeline and the mapping mechanism;
[0077] Adopt the compilation guidance function of HLS to implement the mapping of the HBM multi-channels.
[0078] Among them, in this embodiment, the compilation guidance function of HLS is used to implement the mapping of the HBM multi-channels. For example, allocate multiple PCs of the HBM to a computing unit, and use the following compilation guidance of HLS:
[0079] #pragma HLS INTERFACE m_axi port=cu0_pc01 bundle=axi8
[0080] #pragma HLS INTERFACE m_axi port=cu0_pc02 bundle=axi8.
[0081] Assign one PC of HBM to multiple computing units. The compilation guidance using HLS is as follows:
[0082] #pragma HLS INTERFACE m_axi port=pc01_cu0 bundle=axi0
[0083] #pragma HLS INTERFACE m_axi port=pc01_cu1 bundle=axi0
[0084] Effectively implement the mapping between computing units and multiple channels of HBM through the method of HLS-based compilation guidance, so as to give full play to the effectiveness of the high bandwidth of HBM. Through precise compilation guidance, we can flexibly adjust the mapping and access patterns of data in HBM according to specific computing requirements and workloads, thereby optimizing the mapping relationship between computing units and multiple channels of HBM. This mapping strategy not only improves the memory bandwidth and data transfer efficiency, but also ensures the accuracy and consistency of data access.
[0085] Furthermore, when performing feature aggregation operations and combination operations in each layer of the GCN model, task pipeline optimization is also included:
[0086] Based on the pipeline task scheduling optimization strategy of the greedy algorithm, optimize the pipeline scheduling between the feature aggregation operation and the combination operation of GCN;
[0087] Based on the strategy of fine-grained data partitioning, and introducing deep pipeline and data flow parallel processing methods, on each computing node of CPU-FPGA, adopt a fine-grained pipeline design to achieve pipeline parallelism of the data transmission and computing parts in the GCN model.
[0088] Among them, in this embodiment, based on the pipeline task scheduling optimization strategy of the greedy algorithm, this strategy is used to optimize the pipeline scheduling between the feature aggregation operation and the combination operation of GCN, and achieve pipeline processing between stages with a finer granularity, so as to give play to the pipeline parallel advantage of the heterogeneous platform FPGA and further improve the performance of the GCN accelerator.
[0089] Aiming at the characteristics of GCN in the aggregation and combination stages, this embodiment adopts a strategy based on fine-grained data partitioning, and introduces deep pipeline and data flow parallel processing methods. On each computing node of CPU-FPGA, we use a fine-grained pipeline design to achieve pipeline parallelism of the data transmission and computing parts in GCN, so as to hide the communication time and computing time from each other and reduce the overall execution time of GCN.
[0090] Therefore, in this embodiment, through a deep pipeline design, the bandwidth delay problem of GCN during data transmission is further hidden. This measure greatly optimizes the utilization efficiency of limited hardware resources and improves the performance of the GCN accelerator as a whole.
[0091] Furthermore, the hybrid fixed-point quantization strategy using hierarchical classification of the GCN model obtains the performance evaluation results of the GCN inference acceleration method through performance evaluation experiments on multiple datasets, and further includes:
[0092] Convert the original floating-point matrix multiplication into low-precision integer multiplication operations;
[0093] Adopt a hybrid fixed-point quantization strategy based on hierarchical and classification. For the feature matrix of the first layer of the GCN model, use a high bitwidth, and for the feature matrix of the second layer, use a low bitwidth;
[0094] Perform performance tests on various data flow engines based on different datasets on the Vanilla GCN model.
[0095] Among them, in this embodiment, the performance comparison (throughput) of multi-channel HBM will be obtained:
[0096] In various dataset scenarios, there are differences in the bandwidth benefits obtained by the GCN accelerator using HBM and effective mapped PCs. The evaluation results show that when processing large-scale datasets such as Yelp, Amazon, Ogbn, etc., higher bandwidth benefits can be achieved by using HBM and reasonable PC mapping, further improving the performance of the GCN accelerator;
[0097] Performance evaluation of hybrid fixed-point quantization: By using appropriate hybrid fixed-point quantization, the inference speed of GCN can be effectively improved, and at the same time, the accuracy loss is almost negligible, and the accuracy loss is controlled within 1%.
[0098] In the GCN model, the hybrid fixed-point quantization method using hierarchical classification is because the feature matrices of the GCN model are passed layer by layer during the inference process, that is, the output feature matrix of the previous layer of GCN will be used as the input of the feature matrix of the next layer to participate in the inference operation. Therefore, the quantization error of the previous layer of GCN will accumulate into the next layer, resulting in a gradual increase in the error. Through comprehensive theoretical analysis and experimental results, we found that the GCN model adopting the hybrid fixed-point quantization strategy of int16|int8 can achieve a fast inference speed while maintaining a high accuracy, thus achieving a good balance between speed and accuracy.
[0099] Performance Evaluation of Multiple Compute Units (CUs): There are significant differences in the scalability of various datasets under multiple CUs. For small-scale datasets such as Cora and Citeseer, only a relatively small number of CUs are needed to meet the load requirements of the entire working dataset. In this case, as the number of CUs further increases, the performance improvement is not significant. However, for larger-scale datasets such as Yelp, Amazon, and Ogbn, multiple CUs can produce nearly linear acceleration effects. This is because in the processing of large-scale datasets, multiple CUs can more effectively process data in parallel, thus significantly improving the overall performance.
[0100] Performance Comparison of CPU-based Accelerators: In terms of the execution time on each dataset, the accelerator DAHBM-GCN achieved significant performance improvement compared to the PyG-GCN on CPU, with an average speedup of 52.5 - 129.3 times. At the same time, compared to DGL-GCN, the average speedup can also reach 4.9 - 7.9 times.
[0101] Comprehensive Performance Comparison and Evaluation based on FPGA: In terms of the execution time on each dataset, the accelerator DAHBM-GCN can achieve an average speedup of 1.21 - 2.21× compared to AWB-GCN, 1.25 - 1.98× compared to HyGCN, and 1.65 - 2.68× compared to the HLS-GCN accelerator. This experimental result shows that our DAHBM-GCN accelerator has good performance acceleration on the FPGA platform.
[0102] Energy Consumption Comparison: Among various datasets, DAHBM-GCN achieved energy savings of 9.8 - 36.23×, 1.36 - 3.34×, 1.49 - 5.45×, and 1.68 - 6.64× compared to PyG-GCN, AWB-GCN, HyGCN, and HLS-GCN respectively. Because the proposed accelerator has less computational volume and better memory storage access, it can achieve better utilization of computing resources.
[0103] Furthermore, the performance testing of various data flow engines based on different datasets on the Vanilla GCN model also includes:
[0104] Measuring and comparing the execution time of various data flow engines of the GCN model on different datasets to obtain the efficiency of the data flow engines when processing different datasets;
[0105] There are differences in the execution time of different execution orders of the aggregation and combination stages on different datasets;
[0106] When the aggregation stage is executed prior to the combination stage, if the dimension of the result matrix after performing the aggregation operation is reduced more, then the execution order with the aggregation stage prior to the combination stage will reduce the execution time;
[0107] When the combination stage is executed prior to the aggregation stage, if the dimension of the result matrix after performing the combination operation is reduced more, then the execution order with the combination stage prior to the aggregation stage will result in less execution time.
[0108] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0109] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. The above is only the implementation manner of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, is equally included in the patent protection scope of the present application.
[0110] The specific implementation manners of the invention have been described in detail above, but they are only examples, and the present application is not limited to the specific implementation manners described above. For those skilled in the art, any equivalent modification or substitution to the invention is also within the scope of the present application. Therefore, equivalent transformations, modifications, improvements, etc. made without departing from the spirit and principles of the present application should all be covered within the scope of the present application.
Claims
1. A GCN inference acceleration method based on multi-data streams and high-bandwidth memory, characterized in that, Including: A computing engine configured for the GCN model that supports multi-data streams, aggregation-first, and combination-first orders; An adaptive selector for the multi-data stream computing engine based on a decision tree, which selects the best computing engine according to the characteristics of various data sets and the characteristics of the GCN model; Configuring a multi-channel high-bandwidth memory HBM and designing a mapping strategy for the multi-channels of the multi-channel high-bandwidth memory to data and computing units; Performing task pipeline optimization when performing feature aggregation operations and combination operations in each layer of the GCN model; Adopting a hybrid fixed-point quantization strategy for hierarchical classification of the GCN model, and obtaining the performance evaluation results of the GCN inference acceleration method through performance evaluation experiments on various data sets.
2. The GCN inference acceleration method based on multi-data streams and high-bandwidth memory according to claim 1, wherein The computing engine configured to support multi-data streams, aggregation-first, and combination-first orders further includes: The GCN model includes an aggregation stage and a combination stage; When the data set satisfies that the GCN reduces the feature dimension after executing the combination stage, first execute the combination stage, and then execute the aggregation stage; The activation functions used in the aggregation stage and the combination stage satisfy the idempotent property; In the GCN model, the core operations of the aggregation and combination stages are matrix multiplications; The data flow patterns of the matrix multiplications include four types: inner product, outer product, row-aware product, and column-aware product; Based on the characteristics and advantages of the data streams of the inner product, outer product, row-aware product, and column-aware product, designing general matrix multiplication and sparse-dense matrix multiplication computing engine kernels for the GCN model respectively.
3. A GCN inference acceleration method based on multi-data streams and high-bandwidth memory according to claim 1, characterized in that The adaptive selector for the multi-data stream computing engine based on a decision tree, which selects the best computing engine according to the characteristics of various data sets and the characteristics of the GCN model, further includes: A selector for the multi-data stream computing engine based on a decision tree, using the number of vertices N, the number of non-zero elements of matrix A, the number of non-zero elements of matrix X, the feature matrix F1, the weight matrix W, and the node degree D as the characteristics of the data set and as the input data of the decision tree; Using data sets with various different characteristics to train the decision tree model, and in actual application scenarios, using the decision tree model to select the most suitable data stream computing engine.
4. A GCN inference acceleration method based on multi-data streams and high-bandwidth memory according to claim 1, characterized in that, The configuration of the multi-channel high-bandwidth memory HBM and the design of the mapping strategy for the multi-channels of the multi-channel high-bandwidth memory to data and computing units further includes: Adopting a partition mapping mechanism, grouping the graph data rows in the convolution operation of the GCN model based on the degree of vertices to achieve reasonable partitioning of the graph; Through the mapping of the partition of the graph data to the multi-channels of the HBM, and the mapping of the multi-channels of the HBM to the data stream computing engine kernels; Storing the partitioned graph data in on-chip memory and utilizing the synergistic effect of the pipeline and the mapping mechanism; Using the compilation guidance function of HLS to achieve the mapping of the multi-channels of the HBM.
5. A GCN inference acceleration method based on multi-data streams and high-bandwidth memory according to claim 1, characterized in that The execution of task pipeline optimization when performing feature aggregation operations and combination operations in each layer of the GCN model further includes: A pipeline task scheduling optimization strategy based on a greedy algorithm to optimize the pipeline scheduling between the feature aggregation operation and the combination operation of the GCN. Based on the strategy of fine-grained data partitioning, and introducing deep pipelining and data stream parallel processing methods, on each computing node of the CPU-FPGA, a fine-grained pipeline design is adopted to achieve pipeline parallelism in the data transmission and calculation parts of the GCN model.
6. The GCN inference acceleration method based on multi-data streams and high-bandwidth memory according to claim 1, wherein The hybrid fixed-point quantization strategy using hierarchical classification of the GCN model, and obtaining the performance evaluation results of the GCN inference acceleration method through performance evaluation experiments on multiple datasets, further includes: Converting the original floating-point matrix multiplication into low-precision integer multiplication operations; Adopting a hybrid fixed-point quantization strategy based on hierarchical and classification, using a high bit-width for the feature matrix of the first layer of the GCN model and a low bit-width for the feature matrix of the second layer; Performing performance tests on various data stream engines based on different datasets on the Vanilla GCN model.
7. A GCN inference acceleration method based on multi-data streams and high-bandwidth memory according to claim 1, characterized in that The performance tests on various data stream engines based on different datasets on the Vanilla GCN model further include: Measuring and comparing the execution times of various data stream engines of the GCN model on different datasets to obtain the efficiency of the data stream engines when processing different datasets; There are differences in the execution times of different execution orders of the aggregation and combination stages on different datasets; When the aggregation stage precedes the combination stage in execution, if the dimension of the result matrix after the aggregation operation decreases more, the execution order with the aggregation stage prior to the combination stage will reduce the execution time; When the combination stage precedes the aggregation stage in execution, if the dimension of the result matrix after the combination operation decreases more, the execution order with the combination stage prior to the aggregation stage will result in less execution time.