Graph neural network processor simulation framework-oriented acceleration method and system

By adopting GPU parallel computing and optimization algorithms in the simulation framework, the performance bottlenecks and insufficient simulation accuracy in simulating dedicated GNN processors in the existing technology are solved, efficient simulation and evaluation are achieved, and the overall efficiency of GNN accelerator design is improved.

CN120124673APending Publication Date: 2025-06-10INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510266543.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art has performance bottlenecks when simulating dedicated GNN processors, with limited computing power, which is difficult to meet the needs of large-scale graph datasets, and the simulation accuracy is insufficient, making it impossible to comprehensively evaluate computing and resource management performance.

Method used

Using a GPU-based simulation framework, the computing tasks of a dedicated GNN processor are broken down into multiple nodes, each node corresponds to a computing unit of the GPU. Through the rearrangement optimization of graph neural network operators, kernel fusion strategy and GNAS programming model, the allocation and execution of computing tasks are optimized.

Benefits of technology

It significantly improves the simulation efficiency of dedicated GNN processors, reduces redundant computing and memory transmission, improves simulation accuracy and efficiency, and provides more accurate performance predictions for dedicated GNN accelerator design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124673A_ABST
    Figure CN120124673A_ABST
Patent Text Reader

Abstract

The invention provides a graph neural network processor simulation framework-oriented acceleration method, which comprises the following steps of: decomposing a calculation task of a special GNN processor into a plurality of nodes by constructing a GPU (Graphics Processing Unit)-based simulation framework, and enabling each node to correspond to a calculation unit of the GPU; operators of the graph neural network are rearranged and optimized, so that vertex feature transformation operators are executed preferentially in the simulation framework, and edge data distribution operators are executed later; adopting a kernel fusion strategy to merge a plurality of GPU kernels into a single kernel; on the basis of a GNAS programming model, a workload is divided into thread blocks with edges as centers, each thread block processes calculation tasks of multiple edges in parallel, and results are accumulated to a target vertex through reduction and atomicAdd function operation. The invention further provides an acceleration system for the graph neural network processor simulation framework, a storage medium and electronic equipment. Therefore, redundant calculation and memory transmission can be reduced, and the simulation efficiency of the special GNN processor is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of graph neural networks, and particularly to an acceleration method, system, storage medium and electronic device for a graph neural network processor simulation framework. Background Art

[0002] Graph neural networks (GNNs) have achieved extensive applications in recent years in fields such as social networks, recommendation systems, and bioinformatics. However, due to the sparsity and randomness of graph structures, traditional computing architectures (such as general-purpose CPUs and GPUs) often face performance bottlenecks when performing GNN computations. Especially in the case of large-scale graph datasets and high-dimensional feature vectors, existing execution platforms (such as CPUs and GPUs) are difficult to meet their requirements for computing performance, memory bandwidth, and parallelism. Therefore, dedicated GNN accelerators have become a necessary and feasible path to solve these problems.

[0003] Simulation frameworks are important in processor design for verifying, analyzing, and optimizing hardware performance by providing a virtual environment, which helps reduce development costs and time, especially when designing dedicated GNN processors. Existing GNN accelerator simulation frameworks mainly rely on general-purpose CPU platforms and cannot efficiently simulate the actual performance of dedicated hardware, especially on large-scale graph datasets; due to the limited computing power of the CPU platform, it often causes significant performance bottlenecks during the simulation process, thus affecting the verification and optimization of accelerator design.

[0004] Related simulation schemes in the prior art include CPU-based frameworks and general-purpose GPU-based frameworks, such as network simulators like Booksim. However, the generality and performance of these frameworks often cannot fully meet the requirements of dedicated GNN processors, especially in aspects such as simulating the computation and resource management of GNN processors.

[0005] Existing graph neural network accelerator simulation schemes have the following problems and deficiencies: 1. Performance bottleneck: Traditional CPU simulations cannot provide sufficient computing power, especially when dealing with large-scale graph data, which is prone to becoming a performance bottleneck; 2. Insufficient simulation accuracy: Existing simulation frameworks cannot comprehensively simulate the actual behavior of dedicated GNN processors and cannot fully evaluate their computation and resource management performance; 3. Low efficiency: Existing GPU accelerator simulation frameworks cannot fully utilize the parallel computing advantages of GPUs when simulating dedicated hardware, resulting in low simulation efficiency. These problems mainly stem from the insufficient simulation ability of existing frameworks for dedicated GNN processor hardware and the failure to fully utilize the parallel computing power of modern GPUs to improve simulation efficiency.

[0006] In summary, it is obvious that the prior art has inconveniences and defects in actual use, so it is necessary to improve it. Summary of the Invention

[0007] In view of the above deficiencies, the object of the present invention is to provide an acceleration method, system, storage medium and electronic device for a graph neural network processor simulation framework, which is used to solve the performance bottleneck problem existing in the simulation of dedicated GNN processors in the prior art.

[0008] To solve the above technical problems, in a first aspect, the present invention provides an acceleration method for a graph neural network processor simulation framework, including the steps of:

[0009] Construct a GPU-based simulation framework, decompose the computing tasks of the dedicated GNN processor into multiple nodes, and each node corresponds to a computing unit of the GPU;

[0010] Rearrange and optimize the operators of the graph neural network to preferentially execute the vertex feature transformation operator and delay the execution of the edge data distribution operator in the simulation framework;

[0011] Adopt a kernel fusion strategy to merge multiple GPU kernels into a single kernel;

[0012] Based on the GNAS (Gather-Network-ApplyEdge-Scatter) programming model, divide the workload into thread blocks centered on edges, each thread block parallelly processes the computing tasks of multiple edges, and accumulates the results to the target vertex through reduction and atomicAdd function operations.

[0013] Further, the vertex feature transformation operator is a linear transformation, the edge data distribution operator is an arithmetic operation, and the two satisfy the distributive law and the commutative law.

[0014] Further, the kernel fusion strategy is to merge feature transformation, edge weight calculation and result accumulation into a single GPU kernel, and the intermediate results are stored in the GPU register.

[0015] Further, the step size of each thread block for parallelly processing edges simultaneously is max(T / m, 1); where T and m are the thread block size and the feature length respectively.

[0016] Further, the computing tasks of each thread block for parallelly processing multiple edges include:

[0017] In the thread block, consecutive threads are used to process consecutive entries of the feature vector of each edge, and computing tasks are dynamically allocated through thread indices.

[0018] In a second aspect, the present invention also provides an acceleration system for a graph neural network processor simulation framework, including:

[0019] A framework construction unit for constructing a GPU-based simulation framework, which decomposes the computing tasks of a dedicated GNN processor into multiple nodes, and each node corresponds to a computing unit of the GPU;

[0020] A rearrangement optimization unit for rearranging and optimizing the operators of the graph neural network to preferentially execute vertex feature transformation operators and delay the execution of edge data distribution operators in the simulation framework;

[0021] A kernel fusion unit for merging multiple GPU kernels into a single kernel by adopting a kernel fusion strategy;

[0022] A processing unit for dividing the workload into thread blocks centered on edges based on the GNAS programming model. Each thread block processes the computing tasks of multiple edges in parallel, and accumulates the results to the target vertex through reduction and atomicAdd function operations.

[0023] In a third aspect, an embodiment of the present invention provides a storage medium for storing a computer program for executing the acceleration method for the simulation framework for a graph neural network processor described in any one of the above.

[0024] In a fourth aspect, an embodiment of the present invention provides an electronic device, including a storage medium, a processor, and a computer program stored on the memory and executable on the processor. The processor implements the acceleration method for the simulation framework for a graph neural network processor described above when executing the computer program.

[0025] Through efficient GPU programming and optimization algorithms, the present invention can significantly improve the simulation efficiency of dedicated GNN processors, reduce redundant calculations and memory transfers during the simulation process, thereby greatly improving the overall efficiency of the design and evaluation of dedicated GNN processors. Moreover, the simulation method based on GPU acceleration in the present invention can adapt to large-scale graph data sets and provide more accurate performance predictions for the design of future GNN accelerators. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a schematic flowchart of the acceleration method for the simulation framework for a graph neural network processor provided in Embodiment 1 of the present invention;

[0027] Figure 2 It is a schematic diagram of redundant calculation using the existing method;

[0028] Figure 3 It is a schematic diagram of redundant calculation by the acceleration method for the simulation framework for a graph neural network processor provided in Embodiment 1 of the present invention;

[0029] Figure 4Schematic diagram of the optimization of the kernel fusion algorithm for the acceleration method of the graph neural network processor simulation framework provided in the first embodiment of the present invention;

[0030] Figure 5 Schematic diagram of the structure of the acceleration system for the graph neural network processor simulation framework provided in the second embodiment of the present invention;

[0031] Figure 6 Schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present invention. Detailed implementation manners

[0032] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0033] It should be noted that the references to "one embodiment", "embodiment", "example embodiment", etc. in this specification mean that the described embodiment may include specific features, structures or characteristics, but not every embodiment must include these specific features, structures or characteristics. In addition, such expressions do not refer to the same embodiment. Further, when combining an embodiment to describe specific features, structures or characteristics, it has been shown that it is within the knowledge of those skilled in the art to combine such features, structures or characteristics into other embodiments whether or not there is an explicit description.

[0034] In addition, in the specification and subsequent claims, certain terms are used to refer to specific components or parts. Those of ordinary skill in the art should understand that manufacturers may use different nouns or terms to refer to the same component or part. The specification and subsequent claims do not use the difference in name as a way to distinguish components or parts, but use the difference in function of components or parts as the criterion for distinction. The terms "comprising" and "including" mentioned throughout the specification and subsequent claims are open-ended terms, and should be interpreted as "including but not limited to". In addition, the term "connected" herein includes any direct and indirect means of electrical connection. Indirect means of electrical connection include connection through other devices.

[0035] When conducting the design and performance evaluation of a dedicated GNN accelerator, existing simulation techniques cannot fully meet the efficient simulation requirements of dedicated GNN processors. The main problem is that existing simulation frameworks based on the CPU platform always execute graph neural network operators in a serial and serialized manner, resulting in the computation and memory transfer during the simulation becoming bottlenecks. In addition, the redundant computations and memory accesses in existing frameworks have not been effectively optimized. The graphics processing unit (GPU) can utilize its numerous threads to optimize serial computation into parallel computation, which can greatly improve the execution speed and efficiency. Based on this discovery, this application proposes an acceleration method for the simulation framework of a graph neural network processor to solve the performance bottleneck problem existing in the prior art when simulating a dedicated GNN processor.

[0036] Before describing the embodiments of this application in detail, first briefly describe the technical concept of this application: Graph operator rearrangement: Prioritize the execution of vertex feature transformation (NN operator) and delay data distribution (Scatter operator) to avoid duplicate computation of vertex features and reduce the computational complexity; Kernel fusion: Merge multiple GPU kernels into a single kernel, reduce the memory transfer of intermediate data, and use registers to store intermediate results to reduce latency; GPU parallel mapping: Divide thread blocks centered on edges and dynamically allocate computational tasks to maximize the GPU throughput. Through the above optimizations, redundant computations and memory accesses are significantly reduced, the simulation efficiency of large-scale graph data is improved, and an efficient and accurate verification tool is provided for the design of dedicated GNN accelerators.

[0037] Next, combine specific embodiments to describe the specific principle of the acceleration method of the simulation framework for a graph neural network processor of this application.

[0038] Figure 1 The acceleration method for the simulation framework of a graph neural network processor provided in Embodiment 1 of the present invention is shown, including the following steps:

[0039] S101: Construct a GPU-based simulation framework, decompose the computational tasks of the dedicated GNN processor into multiple nodes, and each node corresponds to a computing unit of the GPU. That is, design a system containing multiple dedicated GNN processor nodes, decompose the computational tasks of the dedicated GNN processor into multiple nodes, and each node corresponds to a computing unit of the GPU (such as a streaming multiprocessor SM). Each node is responsible for computing a part of the data in the graph; data exchange between nodes is carried out through an efficient communication protocol (such as NVLink or PCIe). During simulation, utilize the GPU to parallelly process the computational tasks of these nodes and simultaneously optimize the memory bandwidth utilization of the nodes. At the same time, design a GPU kernel program to map the computational tasks of the dedicated GNN processor nodes to multiple computing units of the GPU.

[0040] The simulation framework constructed in step S101 can provide an infrastructure for parallel processing for subsequent optimization, ensuring that computing tasks can be distributed to multiple thread blocks of the GPU.

[0041] S102: Rearrange and optimize the operators of the graph neural network to preferentially execute the vertex feature transformation operator in the simulation framework and delay the execution of the edge data distribution operator. That is, adjust the operator execution order to preferentially execute the vertex feature transformation (NN operator) and delay the execution of the edge data distribution (Scatter operator) to reduce the repeated calculation of vertex features.

[0042] Specifically, the vertex feature transformation operator is a linear transformation, and the edge data distribution operator is an arithmetic operation, and both satisfy the distributive law and the commutative law.

[0043] In the optimization of reducing operator redundancy, the core idea is that redundant calculations mainly originate from sharing the same source vertex and performing repeated feature transformations on it. In this embodiment, by delaying the execution of the Scatter operator and preferentially performing the NN operator transformation, the repeated calculation can be significantly reduced, and only one transformation of the same feature is required. Figure 2 Shows the original calculation process with redundant calculations, Figure 3 is the optimized calculation process; in the optimization scheme, first calculate the feature transformation results φ(v1), φ(v2), and φ(v3) of the source vertices, then map these results to the corresponding target vertices through Scatter, and finally complete the calculations of g(φ(v1), φ(v2)) and g(φ(v1), φ(v3)). In this way, the NN operator is first executed for feature transformation, and then Scatter is used to distribute the output to the target vertices. Throughout the process, the number of calculations of the function g is still |e| times, but the function φ with a large calculation overhead only needs to be calculated |v| times. In most cases, g is an arithmetic operator and φ is a linear operator, and both usually satisfy the distributive law and the commutative law; therefore, this redundant calculation can always be effectively eliminated through reasonable operator rearrangement, thereby improving the overall calculation efficiency.

[0044] S103: Adopt a kernel fusion strategy to merge multiple GPU kernels into a single kernel. For kernel fusion acceleration, kernel fusion is one of the most popular methods to reduce kernel startup time and unnecessary data movement, where a series of GPU kernels are fused into a single kernel, and intermediate results can be stored in on-chip registers without being written back to memory.

[0045] Specifically, the kernel fusion strategy combines feature transformation, edge weight calculation, and result accumulation into a single GPU kernel, and the intermediate results are stored in GPU registers. In this embodiment, multiple independent GPU kernels (such as feature transformation, edge weight calculation, result accumulation) are combined into a single kernel, reducing the memory transfer of intermediate data and using GPU registers to store intermediate results; thus, the kernel startup latency can be reduced and the parallel efficiency of the GPU can be improved.

[0046] S104: Based on the GNAS programming model, the workload is divided into thread blocks centered on edges. Each thread block processes the calculation tasks of multiple edges in parallel, and the results are accumulated to the target vertex through reduction and atomicAdd function operations. Specifically, in this embodiment, based on the GNAS programming model, thread blocks are divided centered on edges, and the stride of each thread block for parallel processing of edges is max(T / m, 1); where T and m are the thread block size and feature length respectively, dynamically adapting to the computational load.

[0047] In this embodiment, the single kernel after kernel fusion directly processes data reduction, avoiding writing intermediate results back to memory.

[0048] Within the thread block, partial sums are accumulated through parallel reduction (such as tree summation), and scalar multiplication is performed in combination with edge weights.

[0049] The atomicAdd function (an atomic operation function) is used to accumulate the final result to the feature vector of the target vertex, ensuring the atomicity and data consistency of multi-threaded concurrent writing.

[0050] This programming model is based on the redundancy elimination algorithm described above. The processing of the NN layer of the neural network is inserted into the traditional GAS programming paradigm. The fused GNAS kernel divides the workload into thread blocks in an edge-centered manner. For the thread block size T and feature length m, each thread block processes max(T / m, 1) edges.

[0051] See Figure 4 , in a specific implementation, the kernel fusion algorithm uses an input feature matrix $H_{in}$ with dimensions N*m and a graph input in COO format, sets T = 256 to maintain high utilization of the GPU's SM; each thread block will simultaneously process edges with a stride of max(T / m, 1); within the thread block, consecutive threads are used to process consecutive entries of the feature vector of each edge, and computational tasks are dynamically allocated through thread indices. In this example, T = 4, m = 4, and the thread block works on 2 edges: Edge 1-3 and Edge 2-3 ;

[0052] In a specific implementation process, first, all 4 source vertex feature elements to be processed are selected through thread indexing and thread block indexing. At the same time, the corresponding threads will select the corresponding calculation vectors from the input fully connected weight matrix for scalar-vector multiplication. The 4 threads respectively correspond to different colors in Figure 2. The result calculated in this step is recorded as the partial sum.

[0053] Further, a reduction sum is performed on the partial sum, and then a vector-scalar multiplication is performed with the corresponding edge weight; finally, the multiplication result is accumulated into the feature vector storing the target vertex using atomicAdd.

[0054] In this embodiment, by simulating and transferring the dedicated GNN processor to the GPU and optimizing the kernel program, the simulation efficiency can be effectively improved, the transfer of intermediate data can be reduced, and the computational bottleneck in the simulation process can be reduced;

[0055] Based on the same inventive concept, Figure 5 An acceleration system 100 for a simulation framework of a graph neural network processor according to Embodiment 2 of the present invention is shown. The system 100 includes a framework construction unit 10, a rearrangement optimization unit 20, a kernel fusion unit 30, and a processing unit 40, where:

[0056] The framework construction unit 10 is used to construct a simulation framework based on the GPU, decompose the computational tasks of the dedicated GNN processor into multiple nodes, and each node corresponds to a computational unit of the GPU; the rearrangement optimization unit 20 is used to rearrange and optimize the operators of the graph neural network to preferentially execute the vertex feature transformation operator and delay the execution of the edge data distribution operator in the simulation framework; the kernel fusion unit 30 is used to adopt a kernel fusion strategy to merge multiple GPU kernels into a single kernel; the processing unit 40 is used to divide the workload into thread blocks centered on edges based on the GNAS programming model, each thread block parallelly processes the computational tasks of multiple edges, and accumulates the results to the target vertex through reduction and atomicAdd function operations.

[0057] The acceleration system 100 for the simulation framework of the graph neural network processor provided in this Embodiment 2 has the same or similar technical effects as those in the above Embodiment 1. Its specific technical principle is the same as that in the above Embodiment 1, and will not be elaborated here.

[0058] In summary, through efficient GPU programming and optimization algorithms, the present invention can significantly improve the simulation efficiency of the dedicated GNN processor, reduce redundant calculations and memory transfers in the simulation process, thereby greatly improving the overall efficiency of the design and evaluation of the dedicated GNN processor. By making full use of the parallel computing power and high-bandwidth memory interface of the GPU, the present invention can complete the simulation of large-scale graph datasets in a short time, significantly improving the simulation accuracy and efficiency.

[0059] The present invention also provides a storage medium for storing a computer program for any one of the acceleration methods for the graph neural network processor simulation framework as Figures 1 to 4 described above. For example, computer program instructions, when executed by a computer, can, through the operations of the computer, call or provide the methods and / or technical solutions according to the present invention and achieve the same technical effects. To avoid repetition, details are not described herein again. The program instructions for calling the methods of the present invention may be stored in a fixed or removable storage medium, and / or transmitted through a data stream in a broadcast or other signal-bearing medium and / or stored in the storage medium of a computer device running according to the program instructions.

[0060] According to an embodiment of the present invention, the present invention also provides an electronic device 400 as Figure 6 shown. The electronic device 400 may optionally include a storage medium 200 for storing a computer program and a processor 300 for executing the computer program. When the computer program is executed by the processor 300, any one of the above acceleration methods for the graph neural network processor simulation framework is implemented, triggering the electronic device 300 to execute the methods and / or technical solutions based on the foregoing multiple embodiments and achieving the same technical effects. To avoid repetition, details are not described herein again. It should be noted that the electronic devices in the embodiments of the present invention include mobile electronic devices and non-mobile electronic devices. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer, a netbook, or a personal digital assistant, etc., and the non-mobile electronic device may be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present invention do not make specific limitations.

[0061] It should be noted that the present invention can be implemented in software and / or a combination of software and hardware. For example, it can be implemented using an application specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of the present invention can be executed by a processor to implement the above steps or functions. Similarly, the software program (including related data structures) of the present invention can be stored in a computer-readable recording medium, for example, a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. In addition, some steps or functions of the present invention can be implemented using hardware, for example, a circuit that cooperates with a processor to execute each step or function.

[0062] The present invention can be implemented on a computer as a computer-implemented method, or in dedicated hardware, or in a combination of both. The executable code or portions thereof for the method according to the present invention can be stored on a computer program product. Examples of computer program products include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Optionally, the computer program product includes non-temporary program code components stored on a computer-readable medium for performing the method according to the present invention when the program product is executed on a computer.

[0063] In an alternative embodiment, the computer program includes computer program code components suitable for performing all steps of the method according to the present invention when the computer program is run on a computer. Optionally, the computer program is embodied on a computer-readable medium.

[0064] It should be noted that, in this document, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising that element. In addition, it should be pointed out that the scope of the methods and apparatuses in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described method may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0065] Of course, the present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications according to the present invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims of the present invention.

Claims

1. An acceleration method for a graph neural network processor simulation framework, characterized in that: Includes steps: Build a GPU-based simulation framework to decompose the computational tasks of the dedicated GNN processor into multiple nodes, each of which corresponds to a computational unit of the GPU; Rearrange and optimize the operators of the graph neural network so as to give priority to executing the vertex feature transformation operator and postpone the edge data distribution operator in the simulation framework; Adopting the kernel fusion strategy to merge multiple GPU kernels into a single kernel; Based on the GNAS programming model, the workload is divided into thread blocks with edges as the center. Each thread block processes the computation tasks of multiple edges in parallel and accumulates the results to the target vertex through reduction and atomicAdd function operations.

2. The acceleration method for the graph neural network processor simulation framework according to claim 1, characterized in that: The vertex feature transformation operator is a linear transformation, and the edge data distribution operator is an arithmetic operation, and both satisfy the distributive law and the commutative law.

3. The acceleration method for the graph neural network processor simulation framework according to claim 1, characterized in that: The kernel fusion strategy is to merge feature transformation, edge weight calculation and result accumulation into a single GPU kernel, and the intermediate results are stored in GPU registers.

4. The acceleration method for the graph neural network processor simulation framework according to claim 1, characterized in that: The step length of the edges processed in parallel by each thread block is max(T / m, 1); wherein T and m are the thread block size and characteristic length respectively.

5. The acceleration method for the graph neural network processor simulation framework according to claim 1, characterized in that: The computational tasks of each thread block processing multiple edges in parallel include: The thread block uses continuous threads to process continuous entries of the feature vector of each edge, and dynamically allocates computing tasks through thread indexes.

6. An acceleration system for a graph neural network processor simulation framework, characterized in that: Included are: The framework building unit is used to build a GPU-based simulation framework, decomposing the computing tasks of the dedicated GNN processor into multiple nodes, each of which corresponds to a computing unit of the GPU; A rearrangement optimization unit, used for rearrangement optimization of the operators of the graph neural network, so as to give priority to executing the vertex feature transformation operator and postpone the execution of the edge data distribution operator in the simulation framework; A kernel fusion unit, used to adopt a kernel fusion strategy to merge multiple GPU kernels into a single kernel; The processing unit is used to divide the workload into thread blocks based on the GNAS programming model and with edges as the center. Each thread block processes the computation tasks of multiple edges in parallel and accumulates the results to the target vertex through reduction and atomicAdd function operations.

7. A storage medium, characterized in that: A computer program for storing a method for accelerating a graph neural network processor simulation framework according to any one of claims 1 to 6.

8. An electronic device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the acceleration method for the graph neural network processor simulation framework described in any one of claims 1 to 6 is implemented.