End-to-end compiling method of coarse-grained reconfigurable array based on hierarchical data stream
Through the end-to-end compilation method based on hierarchical data flow, the kernel code of CGRA is automatically processed, solving the problem of time-consuming and data transmission neglected under manual annotation method, and achieving efficient acceleration performance of CGRA.
Patent Information
- Application Number
- CN202510482796.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
In the end-to-end compilation of CGRA, the manual progma labeling method of labeling kernel code that needs to be accelerated, resulting in the time-consuming and labor-intensive code of the code editing process, and ignore the data transmission and link of the host-side code and the CGRA kernel-side code, which is unable to effectively exert the CGRA acceleration performance.
The end-to-end compilation method of a coarse-grained reconstructable array based on hierarchical data streams is adopted. By converting the code to be compiled into an optimized intermediate representation of MLIR, and converting it into a front-end data stream through the CGRA data stream operator, the front-end data stream is converted into a back-end data stream of CGRA based on the preset unloading strategy, and automated Kernel kernel code unloading is performed, kernel code data flow diagrams and CGRA mapping configuration information are generated, and executable files are finally generated.
There is no need for programmers to edit codes on the Kernel kernel code that needs to be accelerated. The automated compilation process greatly reduces manual intervention, realizes joint optimization between the host side and the CGRA side, and effectively exerts the optimization performance of CGRA.
Smart Images

Figure CN120010860A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic compilation, and in particular to an end-to-end compilation method of a coarse-grained reconfigurable array based on hierarchical data flow. Background Art
[0002] Coarse-Grained Reconfigurable Array (CGRA) is a computing architecture with the flexibility of a general-purpose processor and the efficiency of a dedicated integrated circuit. The general process of CGRA mainly consists of two major steps: one is the extraction of data flow graphs, and the other is the mapping and configuration generation of data flow graphs. The input source code at the beginning of CGRA compilation can be a general high-level language (such as C++) or a domain-specific language (DSL) to describe task information, and then the code is separated according to software and hardware. Generally speaking, CGRA is suitable for executing cyclic code, so the goal of division is to use the code of cyclic execution tasks as hardware code and the rest of the code as software code. After compilation, a general processor executable file is obtained. The hardware code requires front-end processing, code transformation and optimization, and then the optimized code is used to perform storage space allocation and time / space domain division. After the above series of transformations, the optimized intermediate expression form needs to be generated for task scheduling and control modules, and then the data flow graph needs to be generated through the final intermediate representation (IR). The obtained data flow graph needs to be used together with the hardware description information file as the input of the mapping module.
[0003] The end-to-end compilation of CGRA is mainly divided into two parts, namely the division, change, optimization, scheduling and mapping of the front-end code and the scheduling and mapping of the back-end. For the front-end part, the existing technical solutions in the field of coarse-grained reconfiguration only use manual progma annotation to annotate the corresponding kernel code that needs to be accelerated for specific applications. The kernel code annotated by the front-end is processed by the framework system (Low Level Virtual Machine, LLVM) of the architecture compiler to obtain the corresponding intermediate code, and then the intermediate code will be converted into a data flow graph and input into the corresponding CGRA mapper to map the corresponding configuration information. However, manual progma annotation requires programmers to spend more energy on code editing, and only generates data flow graphs for the annotated kernel code, while CGRA scheduling is mostly based on loops as the basic unit. Since the manual annotation method unloads each annotated code as an independent kernel code, the data transmission and link between the host side code and the CGRA kernel side code are ignored. Summary of the invention
[0004] The present invention provides an end-to-end compilation method for a coarse-grained reconfigurable array based on hierarchical data flow, so as to solve the problem in the prior art that, during the end-to-end compilation of CGRA, the kernel code that needs to be accelerated is marked by a manual progma marking method, which not only consumes time and effort in the code editing process, but also ignores the data transmission and linking between the host side code and the CGRA kernel side code, and cannot effectively exert the acceleration performance of CGRA.
[0005] The present invention provides an end-to-end compilation method for a coarse-grained reconfigurable array based on hierarchical data flow, the method comprising the following steps: Converting the programming code to be compiled into an optimized intermediate representation of MLIR, and converting the optimized intermediate representation into a front-end data stream through a CGRA data stream operator, wherein the data structure of the front-end data stream is a tree structure of multiple hierarchical data streams, and the hierarchical data streams are generated according to the nested loop structure in the programming code to be compiled; Convert the front-end data stream into a back-end data stream of CGRA according to a preset offloading strategy; Automated kernel code unloading is performed on the backend data flow to obtain a target kernel code, and a corresponding kernel code data flow graph is generated based on the target kernel code; Generate a host object file based on the host side code without the Kernel kernel code, and generate corresponding CGRA mapping configuration information according to the kernel code data flow graph, The host object file and the CGRA mapping configuration information are compiled to obtain an executable file of the programming code to be compiled.
[0006] In some embodiments, converting the optimized intermediate representation into a front-end data stream by using a CGRA data stream operator includes: determining a nested loop structure in the optimized intermediate representation; According to the task operator in the CGRA data flow operator, the nested loop structure is reused as a task node, wherein each same-level loop structure in the nested loop structure is reused as a task node; The optimized intermediate representation is layered according to the task nodes to obtain a corresponding hierarchical data flow.
[0007] In some embodiments, converting the front-end data stream into a back-end data stream of CGRA according to a preset offloading strategy includes: Performing a legality check on the front-end data stream, and when the front-end data stream meets the legal conditions included in the preset offloading policy, converting the front-end data stream into a back-end data stream of CGRA; The legal conditions include: The hardware parameter information of the Kernel kernel code in the front-end data stream meets the PE computing resource restrictions and I / O port restrictions of CGRA; The operators included in the Kernel kernel code in the front-end data stream meet the back-end hardware support of CGRA; In the Kernel kernel code in the front-end data stream, the ratio of the number of memory read and write operations to the total number of computing operations is not greater than a preset ratio threshold.
[0008] In some embodiments, the memory access mode of the backend data stream of the CGRA in the CGRA is a dual-layer access mode, and the dual-layer access mode includes: software dual-layer access and hardware dual-layer access; The software dual-layer access includes: Extract the memory subview corresponding to the backend data stream of the CGRA, and configure the L / S controller of the CGRA according to the cache access mode of the memory subview to transfer the block memory corresponding to the memory subview to the cache area of the CGRA; According to the memory access mode information of the L / S instruction in the node area, the corresponding data stream is read from the block memory of the cache area of the CGRA to the I / O port of the CGRA to perform PE operation, wherein the node area is the area storing the Kernel kernel code in the back-end data stream; The hardware dual-layer access includes: Read the target memory configuration information of the backend data stream to the memory configuration management module of the L / S controller and the I / O control of the CGRA respectively; The target memory configuration information is read from the memory configuration management module through CGRA, and a controller unit corresponding to CGRA is configured according to the target memory configuration information to perform memory access to the backend data stream.
[0009] In some embodiments, the performing of automated Kernel kernel code unloading on the backend data stream to obtain the target kernel code includes: In the task node of the hierarchical data flow in the backend data flow, determining a target node representing the innermost loop body code in the nested loop structure; The Kernel kernel code corresponding to the target node is uninstalled to obtain the target kernel code.
[0010] In some embodiments, generating a host object file based on the host side code without the Kernel kernel code includes: The host-side code without the Kernel code is reduced to the MLIR intermediate representation through the MLIR descent channel; Lowering the MLIR intermediate representation to an LLVM compiler dialect, and converting the LLVM compiler dialect to an LLVM compiler intermediate representation; Build host object files based on the LLVM compiler intermediate representation.
[0011] The present invention also provides an end-to-end compilation device for a coarse-grained reconfigurable array based on hierarchical data flow, the device comprising the following modules: A conversion module, used to convert the programming code to be compiled into an optimized intermediate representation of MLIR, and convert the optimized intermediate representation into a front-end data stream through a CGRA data stream operator, wherein the data structure of the front-end data stream is a tree structure of multiple hierarchical data streams, and the hierarchical data streams are generated according to the nested loop structure in the programming code to be compiled; The conversion module is further used to convert the front-end data stream into a back-end data stream of CGRA according to a preset offloading strategy; An unloading module, used to perform automatic Kernel kernel code unloading on the backend data flow, obtain a target kernel code, and generate a corresponding kernel code data flow graph based on the target kernel code; A generation module, used to generate a host object file based on the host side code without the Kernel kernel code, and generate corresponding CGRA mapping configuration information according to the kernel code data flow graph; A compiling module is used to compile the host object file and the CGRA mapping configuration information to obtain an executable file of the programming code to be compiled.
[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, an end-to-end compilation method for a coarse-grained reconfigurable array based on hierarchical data flow as described above is implemented.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described end-to-end compilation methods for a coarse-grained reconfigurable array based on hierarchical data flow.
[0014] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned end-to-end compilation methods for the coarse-grained reconfigurable array based on hierarchical data flow.
[0015] The end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow provided by the present invention realizes automatic Kernel code unloading through CGRA after the front-end data flow is converted into the back-end data flow, and optimizes the code through the hierarchical data flow based on the nested loop structure, which can handle various complex large-scale application codes, and does not require programmers to perform code editing work on the Kernel code that needs to be accelerated. Programmers do not need to be aware of the existence of CGRA, which saves time and effort. In addition, the host side code and the Kernel code are first separated and then linked and compiled, and finally an executable file is generated, which realizes the joint optimization of the host side and the CGRA side, and effectively exerts the optimization performance of CGRA. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced one by one below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0017] Figure 1 It is a flow chart of an end-to-end compilation method of a coarse-grained reconfigurable array based on hierarchical data flow provided by the present invention.
[0018] Figure 2 It is a schematic diagram of the framework of the end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow provided by the present invention.
[0019] Figure 3 It is an example diagram of the CGRA data flow operator provided by the present invention.
[0020] Figure 4 This is an example diagram of the visualization of the Kernel kernel code data flow graph provided by the present invention.
[0021] Figure 5 It is a schematic diagram of generating the hierarchical data stream provided by the present invention.
[0022] Figure 6 It is a structural schematic diagram of an end-to-end compilation device of a coarse-grained reconfigurable array based on hierarchical data flow provided by the present invention.
[0023] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0025] The end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow provided by the present invention can be deployed in compilers of various programming languages for implementation. In the compiler, the programming code to be compiled is optimized and a corresponding executable file is generated for execution by the computer. By implementing the end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow on the compiler, the present invention also realizes a hierarchical dataflow multi-level optimization compiler (Hierarchical Dataflow-Oriented CGRA Compiler, HDCC) for reconfigurable computing arrays.
[0026] Taking the compiler HDCC as an example, the end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow implemented on the compiler HDCC is described below in conjunction with the accompanying drawings. Figure 1 is a flow chart of an end-to-end compilation method of a coarse-grained reconfigurable array based on hierarchical data flow provided by the present invention, such as Figure 1 As shown, the method includes the following steps 101 to 105.
[0027] Step 101: Convert the programming code to be compiled into an optimized intermediate representation of MLIR, and convert the optimized intermediate representation into a front-end data stream through a CGRA data stream operator.
[0028] First, when the programmer edits the programming code in the compiler (HDCC), the programming code edited in the compiler is obtained as the programming code to be compiled. In an embodiment of the present invention, the compiler is implemented based on the multi-level intermediate representation (MLIR) compilation infrastructure, so the compiler here is called a multi-level intermediate representation compilation system. Of course, there are many programming languages and encoding codes, such as C / C++, Pytorch, etc.
[0029] After obtaining the programming code to be compiled, the programming code to be compiled is converted into an optimized intermediate representation of MLIR. Depending on the programming language, it will also be converted into different optimized intermediate representations of MLIR. Figure 2 As shown, Figure 2It is a schematic diagram of the framework of the end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow provided by the present invention. The programming code to be compiled can be the programming language C / C++ and the programming language Pytorch. For the programming language C / C++, the converted multi-level intermediate representation is Polygeist, and for the programming language Pytorch, the converted multi-level intermediate representation is Torch-MLIR.
[0030] Next, the optimized intermediate representation is converted to the front-end data stream through the CGRA data stream operator. However, before the conversion, the multi-level intermediate representation of some programming languages is converted to a dialect. Here, the Linalg dialect of the multi-level intermediate representation is further reduced to the Affine dialect through the MLIR built-in conversion. The specific reduction process can be achieved through the MLIR built-in optimization technology (MLIR Builtin Optimization).
[0031] Continue to see Figure 2 After the Linalg dialect is uniformly converted to the optimized intermediate representation of the Affine dialect, the optimized intermediate representation is converted to the front-end data stream through the CGRA data stream operator at the front end of the data stream. The CGRA data stream operator is specifically some processing functions for processing the Affine dialect. The processing function example can be as follows Figure 3 As shown, these CGRA data flow operators are used to perform some specific operations of the front-end data flow.
[0032] The process of converting to the front-end data flow is the process of lowering the optimized intermediate representation of the Affine dialect to the front-end data flow (Lower to Frontend Dataflow). The data structure of the front-end data flow is a tree structure of multiple hierarchical data flows, and the hierarchical data flow is generated according to the nested loop structure in the programming code to be compiled.
[0033] Specifically, at the data flow front end, cgra.task and cgra.dispatch in the CGRA data flow operator are first used to hierarchize the nested loop structure in the optimized intermediate representation of the Affine dialect to obtain multiple hierarchical data flows, and loop optimization and data flow optimization are performed on the hierarchical data flows to build the corresponding hierarchical data flow tree structure. Among them, loop optimization can be performed based on the number of processing elements (PE), the number of input / output (I / O) ports, and the cache size of CGRA in the CGRA hardware parameter information. Block optimization and loop unrolling can be implemented through MLIR. In the final tool command line option of MLIR, a description json file of the CGRA hardware parameter information can be specified, and relevant information can be read to perform block optimization and loop unrolling, so as to ensure that the converted front-end data flow is legal when it is subsequently converted to the back end.
[0034] Data flow optimization is to determine the load balance of different kernel codes in the optimized intermediate representation of the Affine dialect under the condition of meeting the CGRA hardware constraints, which can ensure that CGRA can execute different kernel codes simultaneously to the maximum extent and make full use of all computing resources. In other words, it ensures that different load kernels can obtain different numbers of PE resources according to their own kernel size. The number of computing operations of each kernel is counted to evaluate the number of PEs and time required for the corresponding kernel, and the resources of each kernel code are scheduled based on these data.
[0035] Step 102: Convert the front-end data stream into a back-end data stream of CGRA according to a preset offloading strategy.
[0036] like Figure 2 As shown, after the optimized intermediate representation is converted into the front-end data flow at the front end of the data flow, the front-end data flow is lowered to the back-end data flow (Lower to Backend Dataflow), that is, the front-end data flow is lowered to the back-end data flow, specifically, the intermediate representation of the front-end data flow is converted to the intermediate representation of the back-end of the data flow.
[0037] At the back end of the data stream, the front-end data stream needs to be converted into the back-end data stream of CGRA according to the preset offloading strategy. However, before the conversion, the intermediate representation of the front-end data stream should be checked for legitimacy. This process will read the description json file of the CGRA hardware parameter information, and judge whether the corresponding Kernel code can be successfully offloaded to CGRA based on the PE data, I / O port and other information in the file. At the same time, it is also necessary to check whether the operators contained in the Kernel code can be supported by the back-end hardware of CGRA. Even if the current Kernel code can meet all hardware restrictions, this only ensures that CGRA can correctly execute the current Kernel code. However, since the execution efficiency of CGRA is greatly affected by the mapping effect of the CGRA mapper, not all Kernel codes can obtain better execution performance than the CPU. Therefore, it is allowed to use heuristic methods for heuristic offloading according to the characteristics of different Kernel codes. Therefore, corresponding offloading strategies can be formulated according to the characteristics of Kernel codes, such as the proportion of memory read and write operations in the number of computing operations, the proportion of computing resources required for memory read and write operations in the total allocated computing resources, etc.
[0038] After the intermediate representation of the front-end data stream has been checked for legality and is determined to meet the corresponding offloading policy, the intermediate representation of the front-end data stream can be converted into the intermediate representation of the back-end data stream of CGRA at the back-end of the data stream.
[0039] Step 103: Automated kernel code unloading is performed on the backend data flow to obtain the target kernel code, and a corresponding kernel code data flow graph is generated based on the target kernel code.
[0040] After the intermediate representation conversion of the backend data stream of CGRA is completed, the Kernel code is automatically unloaded on the backend data stream to obtain the target kernel code. Figure 2 As shown, first the intermediate representation of the backend data stream is lowered to CGRA (Lower to CGRA), with the purpose of converting the intermediate representation of the backend data stream belonging to the Affine dialect into the corresponding coarse-grained reconfigurable array dialect (CGRA Dialect). Then the automatic kernel code unloading is performed, that is, the automatic kernel code unloading (Auto Kernel Offloading) is performed on the backend data stream. During the unloading process, the hardware and software separation is performed on the current compiler HDCC to separate the kernel code from the host side (CPU side) code. As shown Figure 2As shown, by unloading the kernel code, the host code (MLIR Host Code) without the kernel code belonging to the software aspect and the target kernel code (MLIR Kernels) belonging to the hardware aspect are separated.
[0041] Specifically, the intermediate representation of the backend data flow is still multiple hierarchical data streams, which are constructed based on the nested loop structure. Therefore, during the unloading process, the separation of the MLIR host-side code and the Kernel kernel code obtained by the compiler HDCC unloading can be performed according to each hierarchical data stream.
[0042] In terms of hardware, the corresponding kernel code data flow graph can be generated based on the target kernel code (MLIR Kernels), and the target kernel code obtained by unloading is input into the data flow graph generation module of the compiler HDCC to generate the corresponding kernel code data flow graph. For example, the data flow graph of the kernel kernel code can be seen in Figure 4 .
[0043] The data flow graph generation module analyzes the intermediate representation of the MLIR of the target kernel code in the cgra.node / launch area, and hierarchically analyzes the loop structure information and memory access mode of the nested loops in the target kernel code and the related operators to be executed through the MLIR built-in interface. The corresponding MLIR IR is mapped with the target architecture instruction parameters to obtain the data flow graph information that can replace the original MLIR IR nested loops.
[0044] The embodiment of the present invention provides an implementation scheme for directly generating a data flow graph containing the information required by the CGRA mapping module using the MLIR intermediate representation, rather than directly dropping the MLIR intermediate representation to the LLVM compiler and then generating the data flow graph. Since the LLVM intermediate representation is more low-level, the representation of the relevant information has been converted into the most basic instruction representation in the LLVM intermediate representation. Therefore, when obtaining information such as memory addresses and memory offsets, the LLVM intermediate representation is more difficult to analyze and obtain. Compared with the scheme of generating a data flow graph based on the LLVM intermediate representation, the scheme of directly generating a data flow graph using the MLIR intermediate representation in the embodiment of the present invention is simpler and easier to expand.
[0045] Step 104: Generate a host object file based on the host side code without the Kernel kernel code, and generate corresponding CGRA mapping configuration information according to the kernel code data flow graph.
[0046] In step 103, when the kernel code is uninstalled, the host side code without the kernel code is separated. Here, on the host side (CPU side) of the software, a host object file can be generated based on the host side code without the kernel code. Figure 2 As shown in the figure, the host side code (MLIR Host Code) without the Kernel kernel code can be descended to the LLVM side (Lower to LLVM) through the descent path, converted into LLVM intermediate representation (LLVM IR), and then through the LLVM backend generation tool, the LLVM compiler intermediate representation corresponding to the host side code is generated into the corresponding host object file, that is, the host module Object file.
[0047] On the hardware side, the CGRA generates the kernel code data flow graph for the target kernel code (MLIR Kernels), and then builds the corresponding CGRA mapping configuration information. The target kernel code data flow graph is saved in DOT format after generation, and converted into JSON format through the dot-Tdot_json command of graphviz and used together with the CGRA architecture information file as the input of the mapping unit. The data flow graph can also be visualized (such as Figure 4 ). The kernel code data flow graph is mapped through the coarse-grained reconfigurable array mapping module (CGRA Mapper) to obtain the coarse-grained reconfigurable array mapping configuration information (Configuration), and the corresponding CGRA interface is encapsulated into a C language call function according to the coarse-grained reconfigurable array mapping configuration information. After the host side code descent is completed, the corresponding CGRA call interface can be generated through the cgra.node interface. This CGRA call interface is used to call this C language call function, so as to obtain the CGRA mapping configuration information to configure CGRA, including PE configuration and memory / IO port configuration, for subsequent compilation execution.
[0048] Step 105: compile the host object file and the CGRA mapping configuration information to obtain an executable file of the programming code to be compiled.
[0049] like Figure 2As shown, after the host object file and the coarse-grained reconfigurable array mapping configuration information are generated through step 104, the last step is the compilation and linking process. According to the CGRA mapping module (CGRA Mapper), the CGRA mapping configuration information (Configuration) can be used to generate a corresponding configuration template (Config Template), and then the host object file and the configuration template are input into the host compilation tool chain (for example, it can be RISCV-GNU-Toolchain), and the two are first linked, and then the linked file is compiled to obtain an executable file of the programming code to be compiled for execution by the computer host. The executable file under the coarse-grained reconfigurable array is as follows Figure 2 As shown, in the coarse-grained reconfigurable array, the PE processing unit is controlled simultaneously by the memory controller and the configuration controller, so that when the executable file is executed on the computer host, the PE processing unit is controlled to perform related calculations.
[0050] In the embodiment of the present invention, after the front-end data stream is converted to the back-end data stream, the automatic Kernel code unloading is realized through CGRA, and the code is optimized through the hierarchical data stream based on the nested loop structure, which can handle various complex large-scale application codes, and programmers do not need to edit the code execution of the Kernel code that needs to be accelerated. Programmers do not need to be aware of the existence of CGRA, which saves time and effort. In addition, the host side code and the Kernel code are separated and then linked and compiled, and finally an executable file is generated, which realizes the joint optimization of the host side and the CGRA side from the hardware and software aspects, and effectively exerts the optimization performance of CGRA.
[0051] In some embodiments, at the data flow front end, the optimized intermediate representation is converted into a front-end data flow through the CGRA data flow operator, and the main process is to determine the corresponding hierarchical data flow. Specifically, the nested loop structure in the optimized intermediate representation is first determined, and the compiler HDCC can be used to query and determine the nested loop structure in the optimized intermediate representation. The nested loop structure is a multi-layer nested loop structure code of the programming code to be compiled. After being converted into the optimized intermediate representation of MLIR, the compiler HDCC can determine the nested loop structure in the optimized intermediate representation one by one.
[0052] Next, according to the task operator in the CGRA data flow operator, the nested loop structure is reused as a task node, and each same-level loop structure in the nested loop structure is reused as a task node.
[0053] Here, the CGRA data flow operators can be the two task operators cgra.task and cgra.dispatch. Through cgra.task, the nested loop structure can be reused as a task node Node, while cgra.dispatch is responsible for counting cgra.task and distributing it.
[0054] like Figure 5 As shown, since there may be multiple nested loop structures in the programming code to be compiled, a complete nested loop structure in the optimized intermediate representation of MLIR can be represented as a cgra.schedule function through cgra.dispatch. There are two nested loop structures cgra.node in this function, which means that two cgra.tasks need to be reused. Each cgra.task corresponds to a nested loop structure cgra.node. The two nested loop structures are at the same level and there is no nested relationship. Therefore, the two cgra.tasks are also at the same level. Because each loop structure at the same level is reused as a task node Node, two task nodes Node are expanded under the task list (Schedule) corresponding to the cgra.schedule function, and the two task nodes Node are at the same level, thus forming a tree structure.
[0055] like Figure 5 As shown in the figure, each cgra.task corresponds to a nested loop structure cgra.node. After entering the next layer of the nested loop, the task node Node corresponding to this cgra.node will be used as the task list (Schedule) corresponding to the cgra.schedule function in the next layer. Figure 5 As shown, the two task nodes Node are used as two task lists Schedule respectively.
[0056] After constructing the cgra.schedule function through cgra.dispatch, the corresponding cgra.task is distributed to the next layer of loop cgra.node. At the same time, the corresponding multiple task nodes Node are expanded under each task list (Schedule), so that the tree structure continues to extend downward until the innermost code of the nested loop structure is reached, that is, the bottom layer of the tree structure is reached.
[0057] Through the above steps, the optimized intermediate representation is layered according to the task nodes to obtain the corresponding hierarchical data flow. Using the task nodes, the optimized intermediate representation of the entire MLIR is converted into a multi-layer tree structure according to the hierarchical data flow, realizing the conversion of the nested loop structure in the optimized intermediate representation into multiple hierarchical data flows as the front-end data flow.
[0058] In an embodiment of the present invention, at the front end of the data flow, the optimized intermediate representation of MLIR of the programming code to be compiled is hierarchically processed, and the optimized intermediate representation of MLIR is converted into a hierarchical data structure according to a nested loop structure, thereby providing a larger optimization space for the programming code.
[0059] In some embodiments, the front-end data stream is converted into the back-end data stream of CGRA according to a preset unloading policy, including: performing a legality check on the front-end data stream, and when the front-end data stream meets the legal conditions included in the preset unloading policy, converting the front-end data stream into the back-end data stream of CGRA.
[0060] The purpose of the legality check is to ensure that the subsequent kernel code can be successfully unloaded to CGRA. The preset unloading strategy includes the legal conditions that need to be met during the legality check. The legal conditions specifically include the following three points, which are explained one by one below.
[0061] First, the hardware parameter information of the Kernel code in the front-end data stream meets the PE computing resource restrictions and I / O port restrictions of CGRA. That is, the PE computing resources and I / O ports that execute the Kernel code cannot exceed the restrictions. The second point is to ensure that the operators contained in the Kernel code in the front-end data stream meet the back-end hardware support of CGRA, that is, the operators contained in the Kernel code can be supported by the back-end hardware of CGRA. If the hardware conditions do not support the operators, unloading cannot be achieved.
[0062] In general, if the above two points are met, it can be determined that the front-end data stream is legal, but judging it to be legal only ensures that CGRA can correctly execute the current Kernel kernel code. Since the execution efficiency of CGRA is greatly affected by the mapping effect of the CGRA mapper, not all Kernels can obtain better execution performance than the CPU. The compiler HDCC allows the use of heuristic unloading according to different code characteristics. For example, some programming codes have many calculation operations, and some programming codes occupy a large amount of memory. Therefore, in some embodiments, the unloading of the Kernel kernel code must also meet the third point, that is, in the Kernel kernel code in the front-end data stream, the ratio of the number of memory read and write operations to the total number of calculation operations is not greater than the preset ratio threshold. In this way, the number of memory read and write operations can be limited, and more PE calculation operations of the I / O port can be executed, so that CGRA can obtain better execution performance than the CPU.
[0063] This unloading strategy is to perform heuristic unloading of Kernel kernel code according to different characteristics of Kernel kernel code, which can be expressed as:
[0064] In the above formula, Indicates whether the front-end data stream can be offloaded to CGRA after being converted to the back-end data stream of CGRA. Indicates the number of memory read and write operations. Represents the total number of computation operations, Indicates the preset ratio threshold, It means that the first and second points of the above legal conditions are met.
[0065] In the embodiment of the present invention, when converting the front-end data stream into the back-end data stream of CGRA, an offload strategy is formulated according to whether the data stream can be offloaded to CGRA, and the legitimacy of the front-end data stream is checked, which can not only ensure that the data stream can be successfully offloaded to CGRA, but also further screen the data stream when comparing the performance of CGRA and CPU, so that the execution performance of the data stream that meets the legal conditions on CGRA is better than that on a traditional CPU.
[0066] In some embodiments, when converting the front-end data stream to the back-end data stream of CGRA, considering that the intermediate representation of the front-end data stream is non-regionally isolated, the Node area of the Kernel kernel code needs to be isolated from the outside during the conversion process. Therefore, for external variables outside the Node area, they need to be used as parameters of the Node node of the back-end data stream. In addition, regarding the memory access of external variables, since the CGRA side needs to determine the corresponding access mode during compilation, such external variables can obtain the corresponding sub-memory area by extracting the memory sub-view, which can avoid accessing these external variables. However, this also causes corresponding external additional storage during the conversion of the front-end data stream, and correspondingly causes huge memory access overhead.
[0067] Based on the above scenario, when converting the front-end data stream into the back-end data stream of CGRA, the embodiment of the present invention sets the memory access mode of the back-end data stream of CGRA in CGRA to a double-layer access mode, and the double-layer access mode includes two aspects: software double-layer access and hardware double-layer access.
[0068] For software two-layer access, at the first layer, the memory subview corresponding to the backend data stream of CGRA is first extracted, and the read / write (Load / Store, L / S) controller of CGRA is configured according to the cache access mode of the memory subview to transfer the block memory corresponding to the memory subview to the cache area of CGRA. The cache area can be the scraptch banks module defined in CGRA. At the second layer, according to the memory access mode information of the L / S instruction in the node area, the corresponding data stream is read from the block memory of the cache area of CGRA to the I / O port of CGRA for PE operation. The node area here is the Node area that stores the Kernel kernel code in the backend data stream.
[0069] Specifically, the first layer first extracts the memory subview corresponding to the backend data stream of CGRA, and determines the parameters such as unloading (offset), size (size), and stride (stride) of the cache access mode. Specifically, it can be to extract the parameters in the affine.store function in the backend data stream. These parameter information are used to configure the L / S controller of CGRA, and the block memory corresponding to the memory subview is transferred to the cache area of CGRA, such as the scraptch banks module defined in CGRA. Then the second layer further configures the I / O controller of CGRA according to the memory access mode information of the L / S instruction in the Node area, and reads the corresponding data stream from the block memory of the scraptch banks module in CGRA to the corresponding I / O port of CGRA for calculation by the PE processing unit.
[0070] For hardware double-layer access, the target memory configuration information of the back-end data stream is first read to the memory configuration management module of the L / S controller and the I / O control of CGRA respectively, and then the target memory configuration information is read from the memory configuration management module through CGRA, and the controller unit corresponding to CGRA is configured according to the target memory configuration information to perform memory access to the back-end data stream.
[0071] Specifically, hardware dual-layer access is mainly for the implementation of the configuration access mode of Load / Store Controller and CGRA I / OControl. One is used for read and write control of data flow, and the other is used for calculation and processing control of data flow. Both parts have a ConfigMem module as a memory configuration management module. Through the corresponding call interface on the host side, the target memory configuration information of the back-end data flow is first read into the ConfigMem modules of these two parts. Then, when the CGRA execution call interface is started, the corresponding target memory configuration information can be read from the ConfigMem module to configure the corresponding controller unit in CGRA for memory access.
[0072] In the embodiment of the present invention, a two-layer access mode is designed from the software and hardware aspects respectively for the memory access method of the back-end data flow of CGRA in CGRA. Specifically, by configuring the L / S controller and I / O controller of CGRA, two-layer storage access is realized, which effectively reduces the storage access overhead and the additional storage usage overhead.
[0073] In some embodiments, when performing automated Kernel code unloading on the backend data stream, in order to achieve separation between the host side code and the Kernel code, the method used is to perform automated Kernel code unloading on the backend data stream to obtain the target kernel code. Specifically: first, in the task node of the hierarchical data stream in the backend data stream, determine the target node representing the innermost loop body code in the nested loop structure, and then unload the Kernel code corresponding to the target node in the backend data stream to obtain the target kernel code.
[0074] The backend data stream is composed of multiple hierarchical data streams, which are layered according to the nested loop structure to form a tree structure. Each nested loop structure is reused as a task node Node. Therefore, when unloading, first determine the target node representing the innermost loop body code in the nested loop structure in the task node of the hierarchical data stream in the backend data stream. The innermost loop body code has no dependencies between loops, and there is no nested loop structure inside. Therefore, these codes are the lowest task nodes in the task nodes of the hierarchical data stream, which can also be said to be the leaf nodes of the tree structure. These task nodes belonging to the leaf nodes of the tree structure are the target nodes of the innermost loop body code. The other middle-layer task nodes in the tree structure are loop statements and loop condition codes, and do not belong to the kernel kernel code.
[0075] Of course, the bottom-level task nodes may also exist at the same level, so there may be multiple leaf nodes in the tree structure. These bottom-level and same-level leaf nodes all belong to the target nodes of the innermost loop body code.
[0076] Therefore, at the back end of the data flow, the kernel code corresponding to the target node is directly determined as the final unload code. Therefore, when the back end data flow is unloaded to CGRA, the kernel code corresponding to the target node is directly unloaded to obtain the target kernel code. And the entire unloading process is automatically performed by the compiler HDCC, without the need for manual execution of other operations.
[0077] In the embodiment of the present invention, when the backend data stream is unloaded to CGRA, the host side code and the Kernel code are effectively separated for the backend data stream to determine the software / hardware division scheme. This realizes the automatic extraction of the Kernel code, and the programmer does not need to manually edit and annotate the Kernel code, which saves time and effort. In addition, the automatic separation of the host side code and the Kernel code also lays a foundation for the subsequent joint optimization of the host side and the CGRA side.
[0078] In some embodiments, to generate an executable file of the code to be compiled, after the host side code and the kernel code are separated, a host object file is generated based on the host side code without the kernel code, and the host side code is generally a CPU side code related to the software. Therefore, generating a host object file also realizes code optimization on the software side.
[0079] First, the host-side code without the kernel code needs to be dropped to the LLVM (Low Level Virtual Machine) compiler side. Since the back-end data stream has been converted to the CGRA dialect when unloaded to CGRA, when it is dropped to the LLVM compiler, the host-side code without the kernel code is first dropped to the MLIR intermediate representation through the MLIR drop channel, and then the MLIR intermediate representation is dropped to the LLVM compiler dialect, and finally the LLVM compiler dialect is converted to the LLVM compiler intermediate representation. Of course, these drop operations and conversion operations are completed on the back-end generation tool of the LLVM compiler.
[0080] Finally, the host object file is built based on the LLVM compiler intermediate representation. The host module Object file is generated based on the LLVM compiler intermediate representation using the LLVM compiler backend generation tool, which is used to link with the CGRA mapping configuration information and compile the executable file of the code to be compiled.
[0081] In the embodiment of the present invention, for the host side code without the Kernel kernel code, the MLIR intermediate representation is first converted into the LLVM compiler intermediate representation, and then the host object file is implemented on the LLVM side, thereby realizing the independent optimization process of the host side code and the Kernel kernel code, facilitating the subsequent linking with the CGRA mapping configuration information constructed by the Kernel kernel code, and realizing the final software / hardware partitioning solution.
[0082] The end-to-end compilation device of the coarse-grained reconfigurable array based on the hierarchical data stream provided by the present invention is described below. The end-to-end compilation device of the coarse-grained reconfigurable array based on the hierarchical data stream described below and the end-to-end compilation method of the coarse-grained reconfigurable array based on the hierarchical data stream described above can correspond to each other.
[0083] like Figure 6 As shown, the end-to-end compilation device of the coarse-grained reconfigurable array based on hierarchical data flow includes: a conversion module 601, an unloading module 602, a generation module 603, and a compilation module 604. Specifically, the conversion module 601 is used to convert the programming code to be compiled into an optimized intermediate representation of MLIR, and convert the optimized intermediate representation into a front-end data flow through a CGRA data flow operator, wherein the data structure of the front-end data flow is a tree structure of multiple hierarchical data flows, and the hierarchical data flow is generated according to the nested loop structure in the programming code to be compiled; the conversion module 601 is also used to convert the front-end data flow into a back-end data flow of CGRA according to a preset unloading strategy; The unloading module 602 is used to perform automatic Kernel kernel code unloading on the backend data flow to obtain the target kernel code, and generate a corresponding kernel code data flow graph based on the target kernel code; the generating module 603 is used to generate a host object file based on the host side code without the Kernel kernel code, and generate corresponding CGRA mapping configuration information according to the kernel code data flow graph; The compiling module 604 is used to compile the host object file and the CGRA mapping configuration information to obtain an executable file of the programming code to be compiled.
[0084] It should be noted that the beneficial effects of the end-to-end compilation device of the coarse-grained reconfigurable array based on the hierarchical data flow here and the end-to-end compilation method of the coarse-grained reconfigurable array based on the hierarchical data flow mentioned above can correspond to each other, so the beneficial effects of the end-to-end compilation device of the coarse-grained reconfigurable array based on the hierarchical data flow are not repeated here.
[0085] Figure 7 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 7As shown, the electronic device may include: a processor (processor) 710 , a communication interface (Communications Interface) 720 , a memory (memory) 730 and a communication bus 740 , wherein the processor 710 , the communication interface 720 , and the memory 730 communicate with each other through the communication bus 740 . The processor 710 can call the logic instructions in the memory 730 to execute an end-to-end compilation method for a coarse-grained reconfigurable array based on a hierarchical data flow, the method comprising: converting the programming code to be compiled into an optimized intermediate representation of MLIR, and converting the optimized intermediate representation into a front-end data flow through a CGRA data flow operator, wherein the data structure of the front-end data flow is a tree structure of multiple hierarchical data flows, and the hierarchical data flow is generated according to the nested loop structure in the programming code to be compiled; converting the front-end data flow into a back-end data flow of CGRA according to a preset unloading strategy; performing automatic kernel code unloading on the back-end data flow to obtain a target kernel code, and generating a corresponding kernel code data flow graph based on the target kernel code; generating a host object file based on the host-side code without the kernel code, and generating corresponding CGRA mapping configuration information according to the kernel code data flow graph, and compiling the host object file and the CGRA mapping configuration information to obtain an executable file of the programming code to be compiled.
[0086] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0087] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow provided by the above methods, the method comprising: converting the programming code to be compiled into an optimized intermediate representation of MLIR, and converting the optimized intermediate representation into a front-end data stream through a CGRA data flow operator, wherein the data structure of the front-end data stream is a tree structure of multiple hierarchical data streams, and the hierarchical data stream is generated according to the nested loop structure in the programming code to be compiled; converting the front-end data stream into a back-end data stream of CGRA according to a preset unloading strategy; performing automatic kernel kernel code unloading on the back-end data stream to obtain a target kernel code, and generating a corresponding kernel code data flow graph based on the target kernel code; generating a host object file based on the host side code without the kernel kernel code, and generating corresponding CGRA mapping configuration information according to the kernel code data flow graph, and compiling the host object file and the CGRA mapping configuration information to obtain an executable file of the programming code to be compiled.
[0088] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when executed by a processor to execute the end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow provided by the above methods, the method comprising: converting the programming code to be compiled into an optimized intermediate representation of MLIR, and converting the optimized intermediate representation into a front-end data flow through a CGRA data flow operator, wherein the data structure of the front-end data flow is a tree structure of multiple hierarchical data flows, and the hierarchical data flow is generated according to the nested loop structure in the programming code to be compiled; converting the front-end data flow into a back-end data flow of CGRA according to a preset unloading strategy; performing automatic kernel code unloading on the back-end data flow to obtain a target kernel code, and generating a corresponding kernel code data flow graph based on the target kernel code; generating a host object file based on the host-side code without the kernel code, and generating corresponding CGRA mapping configuration information according to the kernel code data flow graph, and compiling the host object file and the CGRA mapping configuration information to obtain an executable file of the programming code to be compiled.
[0089] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art may understand and implement it without creative effort.
[0090] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An end-to-end compilation method for a coarse-grained reconfigurable array based on hierarchical data flow, characterized in that: The method comprises: Converting the programming code to be compiled into an optimized intermediate representation of MLIR, and converting the optimized intermediate representation into a front-end data stream through a CGRA data stream operator, wherein the data structure of the front-end data stream is a tree structure of multiple hierarchical data streams, and the hierarchical data streams are generated according to the nested loop structure in the programming code to be compiled; Convert the front-end data stream into a back-end data stream of CGRA according to a preset offloading strategy; Automated kernel code unloading is performed on the backend data flow to obtain a target kernel code, and a corresponding kernel code data flow graph is generated based on the target kernel code; Generate a host object file based on the host side code without the Kernel kernel code, and generate corresponding CGRA mapping configuration information according to the kernel code data flow graph, The host object file and the CGRA mapping configuration information are compiled to obtain an executable file of the programming code to be compiled.
2. The end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow according to claim 1 is characterized in that: The step of converting the optimized intermediate representation into a front-end data stream by using a CGRA data stream operator includes: determining a nested loop structure in the optimized intermediate representation; According to the task operator in the CGRA data flow operator, the nested loop structure is reused as a task node, wherein each loop structure at the same level in the nested loop structure is reused as a task node; The optimized intermediate representation is layered according to the task nodes to obtain a corresponding hierarchical data flow.
3. The end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow according to claim 1 is characterized in that: Converting the front-end data stream into a back-end data stream of CGRA according to a preset offloading strategy includes: Performing a legality check on the front-end data stream, and when the front-end data stream meets the legal conditions included in the preset offloading policy, converting the front-end data stream into a back-end data stream of CGRA; The legal conditions include: The hardware parameter information of the Kernel kernel code in the front-end data stream meets the PE computing resource restrictions and I / O port restrictions of CGRA; The operators included in the Kernel kernel code in the front-end data stream meet the back-end hardware support of CGRA; In the Kernel kernel code in the front-end data stream, the ratio of the number of memory read and write operations to the total number of computing operations is not greater than a preset ratio threshold.
4. The end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow according to claim 3 is characterized in that: The memory access mode of the backend data stream of the CGRA in the CGRA is a double-layer access mode, and the double-layer access mode includes: software double-layer access and hardware double-layer access; The software dual-layer access includes: Extract the memory subview corresponding to the backend data stream of the CGRA, and configure the L / S controller of the CGRA according to the cache access mode of the memory subview to transfer the block memory corresponding to the memory subview to the cache area of the CGRA; According to the memory access mode information of the L / S instruction in the node area, the corresponding data stream is read from the block memory of the cache area of the CGRA to the I / O port of the CGRA to perform PE operation, wherein the node area is the area storing the Kernel kernel code in the back-end data stream; The hardware dual-layer access includes: Read the target memory configuration information of the backend data stream to the memory configuration management module of the L / S controller and the I / O control of the CGRA respectively; The target memory configuration information is read from the memory configuration management module through CGRA, and a controller unit corresponding to CGRA is configured according to the target memory configuration information to perform memory access to the backend data stream.
5. The end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow according to claim 1, characterized in that: The step of performing automated Kernel kernel code unloading on the backend data stream to obtain a target kernel code includes: In the task node of the hierarchical data flow in the backend data flow, determining a target node representing the innermost loop body code in the nested loop structure; The Kernel kernel code corresponding to the target node is uninstalled to obtain the target kernel code.
6. The end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow according to claim 1, characterized in that: The generating of the host object file based on the host side code without the Kernel kernel code comprises: The host-side code without the Kernel code is reduced to the MLIR intermediate representation through the MLIR descent channel; Lowering the MLIR intermediate representation to an LLVM compiler dialect, and converting the LLVM compiler dialect to an LLVM compiler intermediate representation; Build host object files based on the LLVM compiler intermediate representation.
7. An end-to-end compilation device for a coarse-grained reconfigurable array based on hierarchical data flow, characterized in that: The device comprises: A conversion module, used to convert the programming code to be compiled into an optimized intermediate representation of MLIR, and convert the optimized intermediate representation into a front-end data stream through a CGRA data stream operator, wherein the data structure of the front-end data stream is a tree structure of multiple hierarchical data streams, and the hierarchical data streams are generated according to the nested loop structure in the programming code to be compiled; The conversion module is further used to convert the front-end data stream into a back-end data stream of CGRA according to a preset offloading strategy; An unloading module, used to perform automatic Kernel kernel code unloading on the backend data flow, obtain a target kernel code, and generate a corresponding kernel code data flow graph based on the target kernel code; A generation module, used to generate a host object file based on the host side code without the Kernel kernel code, and generate corresponding CGRA mapping configuration information according to the kernel code data flow graph; A compiling module is used to compile the host object file and the CGRA mapping configuration information to obtain an executable file of the programming code to be compiled.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow is implemented as claimed in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow as claimed in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the end-to-end compilation method of the coarse-grained reconfigurable array based on hierarchical data flow as claimed in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Data reuse memory access conflict elimination method for coarse-grained reconfigurable structure
CN112631610A
Compiler flow logic for reconfigurable architecture
CN114450661A
Multi-level parallelism development method for multi-array coarse-grained reconfigurable architecture
CN116048521A
Method and device for realizing calculation acceleration based on CGRA and storage medium
CN119225976A
Task execution method based on domain-specific language and software development tool chain
CN119311253A
Cited By
Simulation program generation method, simulation test method and related equipment
CN120315689A