Efficient Execution Method and System for Coarse-Grained Reconfigurable Array Data Flow Processors

By adopting the decoupled PE architecture design in the CGRA architecture, the execution of Codelet nodes is divided into multiple stages, which solves the problem of inefficiency in the execution of Codelet models by the existing CGRA architecture, and achieves higher parallelism and component utilization.

CN116303226BActive Publication Date: 2025-07-01INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310159302.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-14
Publication Date
2025-07-01
Estimated Expiration
2043-02-14

AI Technical Summary

Technical Problem

When executing Codelet models, the existing CGRA architecture cannot fully utilize the advantages of the Codelet model, resulting in low execution efficiency and low component utilization.

Method used

An efficient execution method of coarse-grained reconfigurable array data stream processor is proposed. By decoupling the PE architecture design, the execution of the Codelet node is divided into multiple stages, independently scheduling the components of each stage, and improving the parallelism of nodes and component utilization.

Benefits of technology

It improves the parallelism of nodes and instructions of the program on CGRA, improves program execution efficiency, shortens execution time, and fully improves the utilization rate of CGRA components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116303226B_ABST
    Figure CN116303226B_ABST
Patent Text Reader

Abstract

The present invention provides an efficient execution method and system for a coarse-grained reconfigurable array data flow processor, including: nodes in the directed data flow graph of the program to be executed are code segments, and the connections are dependencies between nodes; PEs of the coarse-grained reconfigurable array data flow processor load the configuration information, operation instructions, and operands of each node from the global cache; schedule the nodes whose predecessor dependencies have been satisfied as the current node to start execution, and divide the code segment of the current node into multiple execution stages; schedule the next loop of the current node to start execution, and during execution, monitor that the components of the coarse-grained reconfigurable array data flow processor corresponding to the next stage of the current node are idle, then the current node enters the next execution stage, and use the components of the coarse-grained reconfigurable array data flow processor to execute its next execution stage; after running through the loops of all nodes in the directed data flow graph, output the current running result from the global cache of the coarse-grained reconfigurable array data flow processor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer architecture, and particularly relates to an efficient design method and system for a data flow processor. Background Art

[0002] The von Neumann architecture is the architecture used by most computer chips today. The characteristic of this architecture is that both instructions and data are stored in memory, read sequentially, and executed sequentially. In the von Neumann architecture, a program is represented as a sequence of instructions, and the way a von Neumann machine runs is to sequentially execute this string of instructions. This execution model is called the control flow model.

[0003] In the 1970s, a new data flow model was proposed. Its core idea is that a program can be represented as a directed data flow graph (DFG). In this graph, nodes represent instructions, the connections represent data, and the direction of the connections represents the dependency relationship of data between instructions. For any node (i.e., instruction), as long as the operands it depends on are ready, it can start execution. Compared with the traditional control flow model, this data flow model has the following advantages:

[0004] 1. The data flow model allows data to flow between instruction nodes, avoiding frequent storage and access, and can greatly reduce the memory access time in computationally intensive application scenarios, improving the efficiency of program execution;

[0005] 2. The data flow model distributes instructions among multiple processing units, and each instruction can be issued and executed as soon as its operands are ready. Compared with the von Neumann architecture, it greatly increases the instruction-level parallelism;

[0006] 3. The data flow model models a program as a directed graph data flow graph, which means that compared with the execution model of the GPU, the data flow model has the advantage of efficiently processing programs with complex dependencies.

[0007] The Coarse-Grained Reconfigurable Array (CGRA) is a spatial computing architecture that can execute programs in a data flow model. A CGRA generally consists of a Network-on-Chip connection, a Processing Element (PE) array, a Host, buffers, etc. The CGRA has better programmability than the FPGA (Field Programmable Gate Array) and better power consumption performance and more general parallel capabilities than the GPU. Before executing a program, a CGRA generally has a process of instruction mapping first. That is, mapping the DFG representing the entire program onto the PE array - this determines at which time point and on which PE each instruction will be executed.

[0008] The traditional way of executing data flow programs using a CGRA is strictly based on fine-grained instructions: that is, during instruction mapping, mapping is performed on an instruction-by-instruction basis; during the execution process, every time an instruction is executed, the generated data needs to be transmitted through the Network-on-Chip to the PE that requires this data. The advantages of this fine-grained execution mode are: on the one hand, the modeling of instruction mapping is simpler and more direct, and the optimization granularity is finer; on the other hand, the control logic of the PE is relatively simple. Its disadvantage is that there is a transmission overhead every time an instruction is executed, and the overhead is relatively large.

[0009] The Codelet model is a coarse-grained partitioning of a program in the field of high-performance computing. Compared with the traditional fine-grained data flow model based on instructions, it is a coarse-grained data flow model with code segments as units - each node is a code segment rather than an instruction. The benefit brought by this model is that the number of nodes can be greatly reduced, and the overhead of instruction mapping is also greatly reduced.

[0010] However, if the Codelet model is directly applied to the existing CGRA architecture, only one instruction is still executed at a time. This execution efficiency is very low: during the execution of an instruction, all components in the PE are alternately in an idle state, and other instructions must also wait. Summary of the Invention

[0011] The purpose of the present invention is to solve the problem that the existing operation mode of the CGRA architecture cannot fully utilize the advantages brought by the Codelet model, and proposes a design of PE decoupling and a specific CGRA architecture that can efficiently execute the Codelet model.

[0012] In view of the deficiencies of the prior art, the present invention proposes an efficient execution method for a coarse-grained reconfigurable array data flow processor, which includes:

[0013] Step 1: Obtain the program to be executed represented by a directed data flow graph, where each node in the directed data flow graph is a code segment, and the connection direction between nodes represents the dependency relationship of data between code segments; the PEs of the coarse-grained reconfigurable array data flow processor load the configuration information, operation instructions, and operands of each node from its global cache.

[0014] Step 2: Schedule the nodes whose predecessor dependencies are satisfied as the current nodes to start execution, and divide the code segments of the current nodes into multiple execution stages.

[0015] Step 3: Schedule the next loop of the current node to start execution. During execution, monitor that the components of the coarse-grained reconfigurable array data flow processor corresponding to the next stage of the current node are idle. Then the current node enters the next execution stage, and use the components of the coarse-grained reconfigurable array data flow processor to execute its next execution stage.

[0016] Step 4: After executing the current loop, transmit the execution result of the current loop to the PEs in the coarse-grained reconfigurable array data flow processor that depend on the current node.

[0017] Step 5: Determine whether all the loops of the nodes in the directed data flow graph have been run. If so, end the operation and output the current operation result from the global cache of the coarse-grained reconfigurable array data flow processor; otherwise, execute Step 2 again.

[0018] In the efficient execution method of the coarse-grained reconfigurable array data flow processor, the PEs of the coarse-grained reconfigurable array data flow processor have a reading component, a computing component, and a storage component corresponding to the reading stage, the computing stage, and the storage stage respectively.

[0019] In the efficient execution method of the coarse-grained reconfigurable array data flow processor, the PEs of the coarse-grained reconfigurable array data flow processor include:

[0020] An instruction cache for storing instructions to be executed;

[0021] An operand register for storing operands;

[0022] A router for data exchange within the PE;

[0023] A computing component including arithmetic and logical operation components;

[0024] A reading component that fetches data from the global cache of the coarse-grained reconfigurable array data flow processor and stores it in the operand register inside the PE;

[0025] A storing component for storing the data in the operand register back to the global cache;

[0026] A data transfer component for transferring data to the required PEs.

[0027] A controller for controlling the operation of the entire PE.

[0028] The efficient execution method of the coarse-grained reconfigurable array data flow processor, wherein the controller includes:

[0029] A kernel table for recording the configuration information of each node, generated by the compiler, including the base address information and the number of loop iterations of the node in each execution stage;

[0030] A status table for recording the status information of the node, including whether all dependent nodes of the node have been executed and the current execution stage of the node;

[0031] A scheduling component for scheduling the execution of nodes;

[0032] A message processing component for receiving messages sent by other PEs and performing corresponding processing.

[0033] The present invention also proposes an efficient execution system of a coarse-grained reconfigurable array data flow processor, which includes:

[0034] An initial module for obtaining a program to be executed represented by a directed data flow graph, and each node in the directed data flow graph is a code segment, and the connection direction between nodes represents the dependency relationship of data between code segments; the PEs of the coarse-grained reconfigurable array data flow processor load the configuration information, operation instructions and operands of each node from its global cache;

[0035] An execution module for scheduling nodes whose predecessor dependencies are satisfied as the current node to start execution, and dividing the code segment of the current node into multiple execution stages;

[0036] A monitoring module for scheduling the start of the next loop of the current node, and when monitoring that the component of the coarse-grained reconfigurable array data flow processor corresponding to the next stage of the current node is idle during execution, the current node enters the next execution stage, and uses the component of the coarse-grained reconfigurable array data flow processor to execute its next execution stage;

[0037] A transmission module for, after executing the current loop, transmitting the execution result of the current loop to the PEs in the coarse-grained reconfigurable array data flow processor that depend on the current node;

[0038] A judgment module for judging whether all nodes' loops in the directed data flow graph have been run. If so, end the operation and output the current operation result from the global cache of the coarse-grained reconfigurable array data flow processor. Otherwise, schedule the execution module again.

[0039] The efficient execution system of the coarse-grained reconfigurable array data flow processor, wherein the PE of the coarse-grained reconfigurable array data flow processor has a reading component, a computing component, and a storage component corresponding to the reading stage, the computing stage, and the storage stage respectively.

[0040] The efficient execution system of the coarse-grained reconfigurable array data flow processor, wherein the PE of the coarse-grained reconfigurable array data flow processor includes:

[0041] An instruction cache for storing instructions to be executed;

[0042] An operand register for storing operands;

[0043] A router for data exchange within the PE;

[0044] A computing component including arithmetic and logical operation components;

[0045] A reading component that fetches data from the global cache of the coarse-grained reconfigurable array data flow processor and stores it in the operand register inside the PE;

[0046] A storing component for storing the data in the operand register back into the global cache;

[0047] A data transfer component for transferring data to the required PE;

[0048] A controller for controlling the operation of the entire PE.

[0049] The efficient execution system of the coarse-grained reconfigurable array data flow processor, wherein the controller includes:

[0050] A kernel table for recording the configuration information of each node, generated by a compiler, including the base address information and the number of loop iterations of the node in each execution stage;

[0051] A status table for recording the status information of the node, including whether all the dependent nodes of the node have been executed and the execution stage where the node is currently located;

[0052] A scheduling component for scheduling the execution of the node;

[0053] A message processing component for receiving messages sent by other PEs and performing corresponding processing.

[0054] The present invention also provides a storage medium for storing a program for executing the efficient execution method of any one of the coarse-grained reconfigurable array data flow processors.

[0055] The present invention also provides a client for the efficient execution system of any one of the coarse-grained reconfigurable array data flow processors.

[0056] As can be seen from the above solution, the advantages of the present invention are as follows: Compared with the prior art, the present invention can improve the node and instruction parallelism of program execution on the CGRA, thereby improving the efficiency of program instructions and shortening the program execution time. On the other hand, the present invention can also fully improve the component utilization rate of the CGRA. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a comparison diagram between a fine-grained data flow model with traditional instructions as units and a Codelet coarse-grained data flow model;

[0058] Figure 2 It is a schematic diagram of the PE of the CGRA architecture proposed by the present invention;

[0059] Figure 3 It is a schematic diagram of the process of the CGRA executing the Codelet model proposed by the present invention;

[0060] Figure 4 It is a schematic diagram of the decoupled PE design. DETAILED DESCRIPTION OF THE INVENTION

[0061] In order to improve the component utilization rate of the CGRA running in the Codelet model and give full play to the advantages of the Codelet model, the present invention proposes an optimized data flow execution model (Dataflow Execution Model), and on this basis, proposes a corresponding design of decoupled PE (Decoupled PE) and a specific CGRA architecture implementation. This design can fully schedule multiple nodes mapped to one PE and give full play to the advantages of the Codelet model. Specifically, the present invention includes the following key technical points:

[0062] Key point 1, the design of the decoupled PE architecture; dividing the Codelet nodes and the components of the PE into multiple stages, and the multiple stages can be executed without coupling. The components of each stage can be independently scheduled, which improves the freedom and parallelism of multi-node scheduling on one PE and also improves the component utilization rate of the PE;

[0063] Key point 2, a specific implementation of applying the decoupled PE architecture design; dividing the execution and components of the nodes into multiple stages. By maintaining the status tables of each node, the components of each stage only need to schedule the nodes in this stage. The stages of the nodes are decoupled, and each PE can run multiple Codelet nodes in parallel, improving the node parallelism and component utilization rate.

[0064] In order to make the above features and effects of the present invention more clearly and understandably described, specific embodiments are given below and detailed descriptions are made in conjunction with the accompanying drawings of the specification as follows.

[0065] The decoupled design of the PE proposed by the present invention is as shown in the appendix Figure 4 : The entire execution process of a node is divided into N execution stages. The corresponding executions are also divided into N types of execution components according to these stages. Among them, the execution components are the components inside the PE. For example, the components in the PE can be divided into Load components, computing components, and Store components. Each node is divided according to the stages and enters each stage in a pipeline-like manner in turn, and also executes on various types of components in turn. However, there are essential differences between its design and the idea of the pipeline:

[0066] ● The pipeline is for the pipelining of instructions, while the design of the present invention is for the pipelining of nodes;

[0067] ● The pipeline requires that the time of each stage is equal, while in the design of the present invention, the execution time of each node in each stage is arbitrary;

[0068] ● The pipeline is tightly coupled, and an instruction cannot advance to the next pipeline stage before its predecessor instruction advances to the next pipeline stage. However, the design of the present invention is decoupled, which will be specifically described below;

[0069] ● The pipeline requires that each instruction goes through all stages, but the implementation of the present invention determines whether a stage can be skipped by judging whether the number of instructions in this stage is greater than 0.

[0070] Here, the appendix Figure 4 is used to specifically explain the decoupled design of the PE: If designed according to the pipeline, after node 2 stage 1 finishes execution at t0, since the component 2 in its next stage is still occupied by node 1, it has to insert a wait and cannot enter stage 2 until t1, and only then can node 3 be transferred to component 1 to execute stage 1. However, due to the decoupled design of the components in each stage of the PE in the design of the present invention, after node 2 finishes executing stage 1 at t0, it can release the execution component and let node 3 start executing. This design reduces useless waiting, improves the utilization rate of components, also improves the parallelism of nodes and shortens the program execution time.

[0071] It is worth mentioning that the decoupled design of each component of the PE is not limited to a certain specific design - that is, the decoupled design proposed by the present invention can be achieved through multiple implementations. For example, it can be achieved by maintaining a status table or by maintaining a queue of multiple states.

[0072] The specific implementation of the application of the decoupled PE design proposed by the present invention is as shown in the appendix Figure 2 and has the following components / recorded information:

[0073] ● Inst Buffer: Instruction buffer, which stores instructions to be executed;

[0074] ●Operand RegFile: The operand register file stores operands;

[0075] ●Router: Routing is used to transfer data and messages of this PE, that is, for the transfer and exchange of data within the PE;

[0076] ●CAL Pipline: The computing component adopts a superscalar design and contains some arithmetic and logical operation components. It corresponds to the CAL Stage in the four execution stages (LD, CAL, FLOW, ST);

[0077] ●LOAD Stream: The reading component fetches data from the global buffer of the CGRA and stores it in the OperandRegFile inside the PE. It corresponds to the LD Stage in the four stages;

[0078] ●Store Stream: The data storage component is used to store the data in the Operand RegFile back to the global buffer of the CGRA. It corresponds to the ST Stage in the four stages;

[0079] ●Dataflow Unit: The data transfer component is used to transfer data to the required PE. It corresponds to the Flow Stage in the four stages;

[0080] ●Controller: The controller is used to control the operation of the entire PE. The entire Controller also includes:

[0081] ■Kernel Table: Records the configuration information of each node, generated by the compiler, including the base address information and loop count of the node in the four execution stages, etc.;

[0082] ■Status Table: Records the status information of the node, including whether all dependent nodes of this node have been executed and the execution stage that this node is in;

[0083] ■Scheduler: The scheduling component is used to schedule the execution of nodes;

[0084] ■ACKPort: The message processing component is used to receive messages sent by other PEs and perform corresponding processing. For example, when a certain node A of another PE finishes execution, after this PE receives this message, it can start the execution of the subsequent node B and return an ack message to the PE that sent the message.

[0085] The specific working process of the present invention:

[0086] In step S101, the PE loads the configuration information, operation instructions, and operands of each Node from the global buffer of the CGRA.

[0087] In step S102, the scheduler determines which nodes to schedule for execution by judging through the status table which nodes' predecessor dependencies are all satisfied. Whether the predecessor dependencies of a node are all satisfied can be known from the control information sent by other PEs in step S107.

[0088] In step S103, the scheduler starts to schedule the next cycle of each node for execution. When scheduling for the first time, it is to execute the first cycle. The cycle corresponds to the program requirements. If there are loop operations in the program, the corresponding nodes may have loops. The number of loops is determined according to the corresponding process after being compiled by the compiler.

[0089] In step S104, if the component corresponding to the next stage of the node (such as the load component for fetching data and the store component for storing data within the PE) is idle, then the scheduler starts to schedule the node to enter the next execution stage and executes it using the corresponding execution component. When scheduling for the first time, it is the Load stage, and it is scheduled to start execution on the Load Stream component.

[0090] A node sequentially enters each stage and finally completes the execution of this node (one cycle). The stages are divided into LD, CAL, FLOW, and ST. So the execution process is to enter the LD, CAL, FLOW, and ST stages in sequence, and each stage needs to be executed in the LD or CAL or FLOW or ST component corresponding to this stage. When scheduling this node for the first time, this node enters the LD stage, so it is scheduled to be executed in the LD component.

[0091] In step S105, it is judged whether there are remaining stages for the node. If so, step S104 is executed again; otherwise, step S106 is executed.

[0092] In step S106, after a node finishes executing one cycle, it needs to send information to the PEs that depend on this node to notify these PEs that this node has finished execution.

[0093] In step S107, it is judged whether all the loops of all the nodes in the data flow graph have been run. If so, the operation ends and the current operation result is output; otherwise, step S103 is executed again. The output here is the end of the calculation of the entire CGRA, but the calculation result is stored in the internal storage of the CGRA. It needs to be transferred back to the memory that can be read by the computer.

[0094] The following is a system embodiment corresponding to the method embodiment above. This embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0095] The present invention also proposes an efficient execution system for a coarse-grained reconfigurable array data flow processor, which includes:

[0096] An initial module for obtaining a program to be executed represented by a directed data flow graph, where each node in the directed data flow graph is a code segment, and the connection direction between nodes represents the dependency relationship of data between code segments; the PEs of the coarse-grained reconfigurable array data flow processor load the configuration information, operation instructions, and operands of each node from its global cache.

[0097] An execution module for scheduling nodes whose predecessor dependencies are satisfied as the current node to start execution, and dividing the code segment of the current node into multiple execution stages.

[0098] A monitoring module for scheduling the start of the next loop of the current node. When executing, if it monitors that the components of the coarse-grained reconfigurable array data flow processor corresponding to the next stage of the current node are idle, the current node enters the next execution stage, and uses the components of the coarse-grained reconfigurable array data flow processor to execute its next execution stage.

[0099] A transmission module for transmitting the execution result of the current loop to the PEs in the coarse-grained reconfigurable array data flow processor that depend on the current node after the current loop is executed.

[0100] A judgment module for judging whether all the loops of the nodes in the directed data flow graph have been run. If so, it ends the operation and outputs the current operation result from the global cache of the coarse-grained reconfigurable array data flow processor; otherwise, it schedules the execution module again.

[0101] In the efficient execution system of the coarse-grained reconfigurable array data flow processor, the PEs of the coarse-grained reconfigurable array data flow processor have a reading component, a computing component, and a storage component corresponding to the reading stage, the computing stage, and the storage stage respectively.

[0102] In the efficient execution system of the coarse-grained reconfigurable array data flow processor, the PEs of the coarse-grained reconfigurable array data flow processor include:

[0103] An instruction cache for storing instructions to be executed.

[0104] An operand register for storing operands.

[0105] A router for data exchange within the PE;

[0106] A computing component that includes arithmetic and logic operation components;

[0107] A reading component that fetches data from the global cache of the coarse-grained reconfigurable array data flow processor and stores it in the operand register inside the PE;

[0108] A storing component for storing the data in the operand register back into the global cache;

[0109] A data transferring component for transferring data to the required PE;

[0110] A controller for controlling the operation of the entire PE.

[0111] The efficient execution system of the coarse-grained reconfigurable array data flow processor, wherein the controller includes:

[0112] A kernel table for recording the configuration information of each node, generated by the compiler, including the base address information and the number of loop iterations of the node in each execution stage;

[0113] A status table for recording the status information of the node, including whether all the dependent nodes of the node have been executed and the execution stage where the node is currently located;

[0114] A scheduling component for scheduling the execution of the node;

[0115] A message processing component for receiving messages sent by other PEs and performing corresponding processing.

[0116] The present invention also proposes a storage medium for storing a program for executing the efficient execution method of any one of the coarse-grained reconfigurable array data flow processors.

[0117] The present invention also proposes a client for the efficient execution system of any one of the coarse-grained reconfigurable array data flow processors.

Claims

1. An efficient execution method for a coarse-grained reconfigurable array data flow processor, characterized in that, including: Step 1: Obtain the program to be executed represented by a directed data flow graph, where each node in the directed data flow graph is a code segment, and the connection direction between nodes represents the dependency relationship of data between code segments; the PEs of the coarse-grained reconfigurable array data flow processor load the configuration information, operation instructions, and operands of each node from its global cache. Step 2: Schedule the nodes whose predecessor dependencies are satisfied as the current nodes to start execution, and divide the code segments of the current nodes into multiple execution stages. Step 3: Schedule the next loop of the current node to start execution. During execution, monitor whether the components of the coarse-grained reconfigurable array data flow processor corresponding to the next stage of the current node are idle. If so, the current node enters the next execution stage, and use the components of the coarse-grained reconfigurable array data flow processor to execute its next execution stage. Step 4: After executing the current loop, transfer the execution result of the current loop to the PEs in the coarse-grained reconfigurable array data flow processor that depend on the current node. Step 5: Determine whether all the loops of all nodes in the directed data flow graph have been run. If so, end the run and output the current running result from the global cache of the coarse-grained reconfigurable array data flow processor. Otherwise, execute Step 2 again.

2. The efficient execution method of the coarse-grained reconfigurable array data flow processor according to claim 1, wherein In the PEs of the coarse-grained reconfigurable array data flow processor, there are reading components, computing components, and storing components corresponding to the reading stage, computing stage, and storing stage respectively.

3. The efficient execution method of the coarse-grained reconfigurable array data flow processor according to claim 1, characterized in that, The PE of the coarse-grained reconfigurable array data flow processor includes: An instruction cache for storing instructions to be executed. An operand register for storing operands. A router for data exchange within the PE. A computing component containing arithmetic and logic operation components. A reading component that fetches data from the global cache of the coarse-grained reconfigurable array data flow processor and stores it in the operand register inside the PE. A storing component for storing the data in the operand register back to the global cache. A data transfer component for transferring data to the required PEs. A controller for controlling the operation of the entire PE.

4. The efficient execution method of the coarse-grained reconfigurable array data flow processor according to claim 3, characterized in that, The controller includes: A kernel table for recording the configuration information of each node, generated by the compiler, including the base address information and the number of loop iterations of the node in each execution stage. A status table for recording the status information of the node, including whether all the dependent nodes of the node have been executed and the execution stage where the node is currently located. A scheduling component for scheduling the execution of the node. A message processing component for receiving messages sent by other PEs and performing corresponding processing.

5. An efficient execution system for a coarse-grained reconfigurable array data flow processor, characterized in that, including: An initial module for obtaining the program to be executed represented by a directed data flow graph, where each node in the directed data flow graph is a code segment, and the connection direction between nodes represents the dependency relationship of data between code segments; the PEs of the coarse-grained reconfigurable array data flow processor load the configuration information, operation instructions, and operands of each node from its global cache. An execution module for scheduling the nodes whose predecessor dependencies are satisfied as the current nodes to start execution, and dividing the code segments of the current nodes into multiple execution stages. The monitoring module is used to schedule the start of the next cycle of the current node. When executing, if it monitors that the coarse-grained reconfigurable array data flow processor component corresponding to the next stage of the current node is idle, the current node enters the next execution stage, and the coarse-grained reconfigurable array data flow processor component is used to execute its next execution stage; The transmission module is used to, after completing the current cycle, transmit the execution result of the current cycle to the PE in the coarse-grained reconfigurable array data flow processor that depends on the current node; The judgment module is used to judge whether all nodes in the directed data flow graph have completed their cycles. If so, the operation ends, and the current operation result is output from the global cache of the coarse-grained reconfigurable array data flow processor. Otherwise, the execution module is scheduled again.

6. The efficient execution system of the coarse-grained reconfigurable array data flow processor according to claim 5, characterized in that, In the PE of the coarse-grained reconfigurable array data flow processor, there are a reading component, a computing component, and a storage component corresponding to the reading stage, computing stage, and storage stage respectively.

7. The efficient execution system of the coarse-grained reconfigurable array data flow processor according to claim 5, characterized in that, The PE of the coarse-grained reconfigurable array data flow processor includes: An instruction cache for storing instructions to be executed; An operand register for storing operands; A router for data exchange within the PE; A computing component containing arithmetic and logical operation components; A reading component that fetches data from the global cache of the coarse-grained reconfigurable array data flow processor and stores it in the operand register inside the PE; A data storage component for storing the data in the operand register back into the global cache; A data transfer component for transferring data to the required PE; A controller for controlling the operation of the entire PE.

8. The efficient execution system of the coarse-grained reconfigurable array data flow processor according to claim 7, characterized in that The controller includes: A kernel table for recording the configuration information of each node, generated by the compiler, including the base address information and the number of loops of the node in each execution stage; A status table for recording the status information of the node, including whether all dependent nodes of the node have been executed and the execution stage where the node is currently located; A scheduling component for scheduling the execution of the node; A message processing component for receiving messages sent by other PEs and performing corresponding processing.

9. A storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the efficient execution method of the coarse-grained reconfigurable array data flow processor described in any one of claims 1-4 are implemented.

10. A client for the efficient execution system of the coarse-grained reconfigurable array data flow processor described in any one of claims 5 to 8.

Citation Information

Patent Citations

  • Fine-grained code automatic generation method and system based on multi-view code features

    CN113342318A

  • Near-memory computing system based on data-driven coarse-grained reconfigurable array

    CN114398308A