Data stream array accelerator based on forward network and hybrid routing

By introducing forward-pass networks and hybrid routing in the data flow accelerator, the problems of communication delay and uneven bandwidth utilization are solved, and more efficient data flow graph execution is achieved.

CN120729786APending Publication Date: 2025-09-30INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510887813.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing data flow accelerators suffer from nonlinear growth of communication delay with network scale and uneven bandwidth utilization, which leads to poor data flow flexibility and affects overall execution efficiency.

Method used

A data flow array accelerator design based on forward delivery network and hybrid routing is adopted, combining fixed-direction and variable-direction routing. The forward delivery Mesh network is used to reduce the number of communication hops and delay, and the data packet priority sorting routing algorithm is used to optimize the transmission order.

Benefits of technology

It reduces communication delay, improves routing flexibility and overall network efficiency, reduces waiting delay of downstream nodes, and improves the execution efficiency of data flow graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120729786A_ABST
    Figure CN120729786A_ABST
Patent Text Reader

Abstract

The invention provides a data flow array accelerator based on a forward network and hybrid routing, which is used for executing a data flow diagram, and comprises a storage unit, a plurality of computing units, an on-chip network and a plurality of network nodes, the physical links are established by adopting a Mesh network structure and are used for transmitting data packets between the computing units and between the computing units and the storage units; each computing unit is configured with a routing module, each routing module is configured with at most four routing ports, each routing port is configured to be a fixed routing port or a variable routing port, a data packet acquired by the fixed routing port is processed by the computing unit and then is forwarded to the on-chip network through the specified routing port, and the data packet acquired by the variable routing port is forwarded to the on-chip network through the variable routing port after being processed by the computing unit. The data packet acquired by the variable routing port is processed by the computing unit and then is forwarded to the network-on-chip by dynamically selecting the routing port according to a preset mode, and experiments show that the average routing conflict rate of the data flow array can be reduced, and the overall performance of the computing array is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer science and technology, in particular to the field of data flow accelerator design, and more particularly to a data flow array accelerator based on a forwarding network and hybrid routing. Background Art

[0002] In existing dataflow accelerators, arrays are composed of many processing elements (PEs), often containing dozens or even hundreds of them. The design of the network-on-chip (NoC) directly impacts the communication efficiency between PEs and the system's energy efficiency. With the continuous expansion of accelerator scale and the increasing demand for throughput and computing power, traditional bus-based data exchange methods have resulted in low data transmission rates, becoming a bottleneck limiting accelerator performance improvements. Data communication between PEs via an interconnected NoC has become a widely adopted technical solution. Currently, mesh NoC topologies and their variants are widely used in dataflow acceleration arrays.

[0003] However, existing accelerators still have some shortcomings. First, communication latency increases nonlinearly. For example, in a 128×128 mesh network, data transmission between edge PEs requires 254 hops ((128−1)+(128−1)=254), which will cause significant data request latency during program execution. Second, bandwidth utilization is uneven. When the execution load is unevenly distributed across the PE array, the data transmission bandwidth and the number of network transmission requests required by each PE are different, resulting in extremely high link utilization in some areas of the mesh network and extremely low link utilization in other areas. Third, data flow is inflexible. When data flow graphs are distributed across a mesh network, data flows across multiple PEs often lead to poor execution efficiency of the entire data flow graph.

[0004] In summary, the communication delay of existing array accelerators increases nonlinearly with the network scale, resulting in high latency in data transmission between execution computing units. In addition, the bandwidth utilization of existing array accelerators is uneven. When the load is unevenly distributed, some links are congested while others are idle, resulting in poor data flow flexibility, which in turn affects the overall execution efficiency.

[0005] It should be noted that this background information is provided solely to introduce relevant information of the present invention to facilitate understanding of the technical solution of the present invention. It does not necessarily constitute prior art. In the absence of evidence demonstrating that the relevant information was disclosed prior to the filing date of the present invention, the relevant information should not be considered prior art. Summary of the Invention

[0006] Therefore, the object of the present invention is to overcome the above-mentioned defects of the prior art and provide a data flow array accelerator based on forward delivery network and hybrid routing.

[0007] The purpose of the present invention is achieved through the following technical solutions:

[0008] According to a first aspect of the present invention, a data flow array accelerator based on a forward network and hybrid routing is proposed, which is used to execute a data flow graph. The accelerator includes a storage unit and multiple computing units. The accelerator also includes: an on-chip network, which is a physical link for transmitting data packets between computing units and between computing units and storage units established using a mesh network structure; each computing unit is configured with a routing module, each routing module is configured with at most four routing ports, and each routing port is configured as a fixed routing port or a variable routing port, wherein data packets obtained by the fixed routing port are forwarded to the on-chip network through a designated routing port after being processed by the computing unit, and data packets obtained by the variable routing port are forwarded to the on-chip network through a dynamically selected routing port according to a preset method after being processed by the computing unit.

[0009] Preferably, the Mesh network structure includes an ordinary Mesh network and a forward Mesh network, wherein the ordinary Mesh network is a physical link connecting two adjacent computing units or adjacent storage units and computing units, and the forward Mesh network is a physical link directly connecting two computing units across multiple hops or a computing unit and a storage unit across multiple hops.

[0010] Preferably, the number of hops spanned by each physical link in the forward Mesh network is greater than or equal to 2 hops and less than a preset hop threshold, wherein the preset hop threshold is determined by the main frequency of the computing unit and the scale of the array accelerator.

[0011] Preferably, the routing module is configured as follows: for a routing module with 4 routing ports, 2 fixed-direction routing ports and 2 variable-direction routing ports are configured; for a routing module with 3 routing ports, 2 fixed-direction routing ports and 1 variable-direction routing port are configured; for a routing module with 2 routing ports, 2 fixed-direction routing ports are configured; wherein, each routing port of each routing module is connected to the on-chip network.

[0012] Preferably, the computing unit is configured to: periodically perform the following steps: each routing port of the routing module in each computing unit obtains the data packets required by the computing unit from the on-chip network respectively; the routing module in each computing unit sorts all the data packets obtained by each routing port of its own to obtain a data packet sequence for each routing port; each computing unit processes the data packets in the data packet sequence of each routing port of its own routing module in turn, and the routing module of the computing unit forwards the processed data packets to the on-chip network.

[0013] Preferably, the routing module is configured to sort all data packets obtained by each routing port in the following manner: for all data packets obtained by each routing port, sort them in descending order based on the loop level at which the data carried by each data packet is located to obtain the initial data packet sequence of each routing port; for the initial data packet sequence of each routing port, arrange the data packets at the same loop level in descending order according to depth to obtain the final data packet sequence of each routing port.

[0014] Preferably, in the computing unit, the preset method includes: obtaining the target computing unit to which the data packet processed by the computing unit needs to be transmitted; determining the transmission path of the data packet in the on-chip network based on the obtained target computing unit; and using the routing port in the routing module connected to the determined transmission path as the routing port for forwarding the data packet.

[0015] According to a second aspect of the present invention, a program execution method is provided for executing a specified program. The method comprises: step S1, obtaining a data flow graph compiler, and compiling the specified program using the data flow graph compiler to obtain a data flow graph of the specified program; step S2, obtaining an array accelerator according to any one of claims 1 to 7, and configuring the routing ports of each routing module of the array accelerator as fixed-direction routing ports or variable-direction routing ports based on the data flow graph of the specified program; and step S3, mapping the data flow graph of the specified program to the configured array accelerator, so that the array accelerator executes the specified program.

[0016] According to a third aspect of the present invention, a computer-readable storage medium stores a computer program thereon, wherein the computer program can be executed by a processor to implement the steps of the method according to the second aspect of the present invention.

[0017] According to the fourth aspect of the present invention, an electronic device is proposed, comprising: one or more processors; and a memory, wherein the memory is used to store executable instructions; the one or more processors are configured to implement the steps of the method described in the second aspect of the present invention by executing the executable instructions.

[0018] Compared with the prior art, the advantages of the present invention are:

[0019] This paper proposes an array accelerator design scheme that reduces communication hops and latency through a Mesh array with a forwarding network, using the Mesh to provide physical transmission links. This scheme combines hybrid routing with fixed and variable directions to reduce on-chip resource consumption while improving routing flexibility (fixed directions are simple and stable, while variable directions are dynamically adjustable). Furthermore, a packet priority-based sorting routing algorithm is used to optimize the transmission order in concurrent multi-packet scenarios, reducing waiting delays at downstream nodes caused by unreachable data, thereby improving overall network efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The embodiments of the present invention are further described below with reference to the accompanying drawings, in which:

[0021] Figure 1 Schematic diagram of a traditional Mesh network and a Mesh network with forward delivery according to an embodiment of the present invention;

[0022] Figure 2 Schematic diagram of a combined forward delivery network according to an embodiment of the present invention;

[0023] Figure 3 A schematic diagram of a combined hybrid routing structure according to an embodiment of the present invention;

[0024] Figure 4 An example diagram of data packet sorting according to an embodiment of the present invention;

[0025] Figure 5 Schematic diagram of program execution steps of a combined data stream acceleration array according to an embodiment of the present invention. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below through specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0027] As mentioned in the background technology section, the communication delay of existing array accelerators increases nonlinearly with the network scale, resulting in high latency in data transmission between execution computing units. In addition, the bandwidth utilization of existing array accelerators is uneven. When the load is unevenly distributed, some links are congested while other links are idle, resulting in poor data flow flexibility, which in turn affects overall execution efficiency.

[0028] To address the above issues, the present invention proposes an array accelerator design scheme. This scheme reduces the number of communication hops and delays through a Mesh array with a forward delivery network, and uses the Mesh to provide a physical transmission link. It combines hybrid routing with fixed and variable directions to reduce on-chip resource consumption while improving routing flexibility (fixed directions are simple and stable, and variable directions are dynamically adjustable). It also adopts a packet priority-based sorting routing algorithm to optimize the transmission order in multi-packet concurrent scenarios, reduce the waiting delay caused by data failure at downstream nodes, and thus improve overall network efficiency.

[0029] According to one embodiment of the present invention, the present invention provides an array accelerator, which includes a storage unit and multiple computing units. The accelerator also includes: an on-chip network, which is a physical link for transmitting data packets between computing units and between computing units and storage units established using a mesh network structure; each computing unit is configured with a routing module, each routing module is configured with at most four routing ports, and each routing port is configured as a fixed routing port or a variable routing port, wherein data packets obtained by the fixed routing port are forwarded to the on-chip network through a designated routing port after being processed by the computing unit, and data packets obtained by the variable routing port are forwarded to the on-chip network through a dynamically selected routing port according to a preset method after being processed by the computing unit.

[0030] In order to better understand the present invention, the present invention is described in detail below with reference to specific embodiments.

[0031] According to one embodiment of the present invention, in the present invention, the Mesh network structure includes an ordinary Mesh network and a forward Mesh network, wherein the ordinary Mesh network is a physical link connecting two adjacent computing units or adjacent storage units and computing units, and the forward Mesh network is a physical link directly connecting two computing units across multiple hops or a computing unit and a storage unit across multiple hops.

[0032] In order to facilitate the understanding of the Mesh network structure used in the on-chip network of the present invention, the following Figure 1 Detailed description.

[0033] Ordinary Mesh structure network such as Figure 1As shown in (a), in a typical Mesh structure, each PE is connected to the surrounding PEs, and only the PEs in the first column and first row can directly access storage. If PE11 needs to access data from storage A, it can only do so through PE01. Similarly, traditional Mesh networks are single-hop, meaning that if PE00 needs to transmit data to PE21, it must pass through PE10 and PE20 before reaching PE21. This results in significant communication and transmission delays, and data packets must be temporarily stored within PEs while passing through PE10 and PE20, placing some cache pressure on the routing module. This situation is even more pronounced in complex routing scenarios where a single PE sends and receives multiple data packets simultaneously, leading to greater latency and performance degradation in the array's execution.

[0034] According to one embodiment of the present invention, a forward mesh network is added to the array accelerator of the present invention, see Figure 1 (b) This figure is an example of a 2-hop forward delivery network. Since the array accelerator has added a forward delivery Mesh network, in addition to the PEs in the first column and the first row, all PEs in the second row and the second column can also directly access storage A and storage B. This can effectively reduce the latency of memory access in applications with a lot of memory access. During data communication transmission, if PE00 needs to send a data packet to PE02, it does not need to be transferred through PE01, but can be sent directly to PE02. This reduces the number of hops required for data communication and can effectively reduce the latency of data communication. After adding the forward delivery network, it does not affect the function of the traditional Mesh network. Figure 1 As can be seen in (b), the present invention still retains the common 1-hop Mesh network as an important part of the network function. It should be noted that the forward Mesh network can be 2-hop, 3-hop, or even more. When the array is large, a forward Mesh network with more hops is often selected. Moreover, the forward Mesh network is combinable. In a large-scale array, the PE array can be divided into different blocks, each block can use a forward network with different hops, and forward networks can also be used between blocks. See Figure 2 Blocks composed of fewer PEs can use a 2-hop or 3-hop forward-passing mesh network. Blocks composed of more PEs use a 4-hop forward-passing mesh network design to reduce internal data communication latency. A forward-passing mesh network design is also introduced between different blocks to accelerate data communication between blocks. The invention incorporates a forward-passing mesh network into the array accelerator to reduce the number of hops during data transmission, thereby improving data transmission efficiency between PEs (computing units).

[0035] According to one embodiment of the present invention, in the present invention, the number of hops spanned by each physical link in the forward delivery Mesh network is greater than or equal to 2 hops and less than a preset hop threshold, wherein the preset hop threshold is determined by the main frequency of the computing unit and the scale of the array accelerator, and the preset hop threshold is the maximum distance (maximum hop number) that a data packet can be transmitted in one computing cycle of an array accelerator.

[0036] It should be noted that in an on-chip network without the addition of a forward-delivery Mesh network, each PE can only receive data from the four PEs closest to it (upper and lower functions). After adding a forward-delivery Mesh network, a PE can receive data from multiple PEs in addition to the data from the four PEs closest to it. The number of data packets obtained by each PE will increase, which undoubtedly poses a higher-level challenge to the processing efficiency of the routing module.

[0037] In order to improve the processing efficiency of the routing module, according to one embodiment of the present invention, the present invention adopts a hybrid routing in which fixed direction routing and variable direction routing coexist. The structure is as follows: Figure 3As shown in the figure, each hybrid routing module consists of a data packet unpacking module, which is responsible for extracting information and then sending it to a fixed direction routing or a variable direction routing for subsequent processing. The data packet is then sent from the data packet buffer module to the destination PE. (1) Fixed direction routing uses a preset path, such as downward propagation or rightward propagation. The advantages are simple structure, low area and power consumption; the disadvantages are that the routing path is fixed and can only be configured before the program is executed. The direction cannot be changed during execution. If congestion occurs, it can only wait for the previous data packet to be sent. Similarly, if a failure occurs during execution, the data packet cannot be sent and the upper layer protocol needs to change the routing path for retransmission. (2) The advantage of variable direction routing is that it can dynamically adjust the transmission direction during program execution, which increases the flexibility of routing; the disadvantages are that the structure is relatively complex, and the area and power consumption are high. For PEs in the array, the mesh structure has four directions of routing (east, south, west, and north). PEs located on the four edges of the mesh may only have two or three directions of routing. Therefore, the present invention employs the following hybrid routing configuration: For a routing module with four routing ports, two fixed-direction routing ports and two variable-direction routing ports are configured; for a routing module with three routing ports, two fixed-direction routing ports and one variable-direction routing port are configured; and for a routing module with two routing ports, two fixed-direction routing ports are configured. Furthermore, each routing port of each routing module is connected to the on-chip network. Fixed-direction routing is simple, consumes fewer on-chip resources, and the data transmission direction cannot be changed after configuration. However, variable-direction routing is complex, consumes more on-chip resources, and allows for dynamic adjustment of data transmission direction. Compared to exclusively fixed-direction routing, variable-direction routing allows for dynamic adjustment of routing direction based on actual application execution. For example, if a routing module is focused almost entirely on transmitting data to east-bound PEs, only one of the four fixed-direction routes can transmit data to the east-bound PE, resulting in data packet congestion. In the above scenario, two fixed-direction routes and two variable-direction routes can flexibly change the direction of data packets sent in the variable-direction route to the east. This can reduce data packet congestion and lower data communication latency. Combining the two types of routing can reduce on-chip resource consumption while increasing routing flexibility. Routing determines the direction of transmission. It should be noted that the routing mentioned above does not include the data path between PEs and storage; it only refers to data communication between PEs.

[0038] According to one embodiment of the present invention, in the present invention, each computing unit periodically performs the following steps to perform the calculations in the data flow graph: each routing port of the routing module in each computing unit obtains the data packets required by the computing unit from the on-chip network respectively; the routing module in each computing unit sorts all the data packets obtained by each routing port of its own to obtain a data packet sequence for each routing port; each computing unit processes the data packets in the data packet sequence of each routing port of its own routing module in turn, and the routing module of the computing unit forwards the processed data packets to the on-chip network.

[0039] According to one embodiment of the present invention, all packets received by each routing port are sorted separately in the following manner: For each routing port, all packets are sorted based on the loop level at which the data carried by each packet is located, from highest to lowest, to obtain the initial packet sequence for each routing port; for each routing port's initial packet sequence, packets at the same loop level are sorted from highest to lowest depth to obtain the final packet sequence for each routing port. In a data flow graph with multiple loops, nodes with lower loop levels have fewer iterations and are located deeper in the data flow graph. Ensuring this portion has a high priority prevents premature entry into continuous iterations and reduces packet injection in the interconnect. In complex routing scenarios with multiple packets arriving simultaneously, packets can be forwarded in an appropriate order. The computing unit processes the packets based on the acquired packet sequence, reducing delays at downstream nodes caused by the non-arrival of required data, making the data flow graph execute more smoothly and alleviating data congestion.

[0040] For example, see Figure 4 Suppose that in one clock cycle, PE23 receives two packets, packet A and packet B, simultaneously in the northbound direction. Packet A is a load packet for memory access, sent from PE13 to PE23. The data it carries is the data currently being stored, and is transmitted to PE23 via PE03 and PE13. Packet B is a flow packet, sent by PE03 to PE23. These two packets need to be sorted according to the priority algorithm. Packet A is in loop 5, and packet B is in loop 3. Packet B's loop level is smaller than that of packet A, so packet B has a higher priority than packet A. PE23 processes packet B first, then packet A. In this example, there are no multiple packets in the same loop level, so there is no need to further sort them according to the depth of the data flow graph.

[0041] According to one embodiment of the present invention, in the computing unit of the present invention, the preset method includes: obtaining the target computing unit to which the data packet processed by the computing unit needs to be transmitted; determining the transmission path of the data packet in the on-chip network based on the obtained target computing unit; and using the routing port in the routing module connected to the determined transmission path as the routing port for forwarding the data packet.

[0042] According to one embodiment of the present invention, a program execution method implemented based on the data flow array accelerator is proposed. The method includes: step S1, obtaining a data flow graph compiler, and using the data flow graph compiler to compile a specified program to obtain a data flow graph of the specified program; step S2, obtaining the data flow array accelerator proposed by the present invention, and configuring the routing port of each routing module of the array accelerator as a fixed-direction routing port or a variable-direction routing port based on the data flow graph of the specified program; step S3, mapping the data flow graph of the specified program to the configured array accelerator, so that the array accelerator executes the specified program.

[0043] According to one embodiment of the present invention, the present invention proposes an array accelerator design scheme that reduces communication hops and latency by introducing a mesh array with a forward delivery network. This scheme combines fixed-direction and variable-direction hybrid routing to reduce on-chip resource consumption while improving routing flexibility (fixed-direction is simple and stable, while variable-direction is dynamically adjustable). Furthermore, a packet priority-based sorting routing algorithm is employed to optimize the transmission order in concurrent multi-packet scenarios, reducing waiting delays at downstream nodes due to unreached data, thereby improving overall network efficiency.

[0044] To verify the beneficial effects of the present invention, the inventors conducted comparative experiments comparing the performance of the array accelerator proposed in this invention with that of the NVIDIA GPU Jetson NX. The experimental results show that introducing a forward-passing Mesh network into the accelerator can reduce the average routing conflict rate of the data flow array. The conflict rate reduction ratio for complex data-dependent flow applications is 54%, and the conflict rate reduction ratio for typical DSP applications is 41%. This result is attributed to the reduction in routing hops and the improvement in efficiency based on priority forwarding. After introducing the forward-passing Mesh network into the accelerator, the results on a typical DSP algorithm test set show that the overall performance of the computing array is improved by 1.11 to 2.19 times. Compared with the NVIDIA GPU Jetson NX, the DFUL using the forward-passing Mesh network design described in this patent achieved a 3.14-fold efficiency improvement. At the same time, after combining the hybrid routing optimization solution proposed in this invention, the forward-passing network only introduced 2.58% of on-chip resource overhead compared to the traditional Mesh network.

[0045] It should be noted that although the above describes the various steps in a specific order, it does not mean that the steps must be performed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order as long as the required functions can be achieved.

[0046] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0047] A computer-readable storage medium may be a tangible device that holds and stores instructions used by an instruction execution device. Computer-readable storage media may include, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove having instructions stored thereon, and any suitable combination thereof.

[0048] While various embodiments of the present invention have been described above, the above descriptions are intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A data flow array accelerator based on a forward pass network and hybrid routing, for executing a data flow graph, the accelerator comprising a storage unit and a plurality of computing units, characterized in that: The accelerator further comprises: On-chip network, which is a physical link for transmitting data packets between computing units and between computing units and storage units established using a mesh network structure; Each computing unit is configured with a routing module, each routing module is configured with at most four routing ports, and each routing port is configured as a fixed routing port or a variable routing port. The data packets obtained by the fixed routing port are processed by the computing unit and then forwarded to the on-chip network through the designated routing port. The data packets obtained by the variable routing port are processed by the computing unit and then forwarded to the on-chip network through the dynamically selected routing port in a preset manner.

2. The data stream array accelerator according to claim 1, characterized in that: The Mesh network structure includes a common Mesh network and a forward Mesh network, wherein the common Mesh network is a physical link connecting two adjacent computing units or adjacent storage units and computing units, and the forward Mesh network is a physical link directly connecting two computing units across multiple hops or a computing unit and a storage unit across multiple hops.

3. The data stream array accelerator according to claim 1, characterized in that: The number of hops spanned by each physical link in the forward mesh network is greater than or equal to 2 hops and less than a preset hop threshold, wherein the preset hop threshold is determined by the main frequency of the computing unit and the scale of the array accelerator.

4. The data stream array accelerator according to claim 1, characterized in that: The routing module is configured to: For a routing module with 4 routing ports, configure 2 fixed-direction routing ports and 2 variable-direction routing ports; For a routing module with three routing ports, configure two fixed-direction routing ports and one variable-direction routing port. For a routing module with two routing ports, configure two fixed-direction routing ports; Each routing port of each routing module is connected to the on-chip network.

5. The data stream array accelerator according to claim 1, characterized in that: The computing unit is configured to periodically perform the following steps: Each routing port of the routing module in each computing unit obtains the data packet required by the computing unit from the on-chip network; The routing module in each computing unit sorts all the data packets obtained by each routing port of the computing unit to obtain the data packet sequence of each routing port; Each computing unit processes the data packets in the data packet sequence of each routing port of its own routing module in turn, and the routing module of the computing unit forwards the processed data packets to the on-chip network.

6. The data stream array accelerator according to claim 5, characterized in that: The routing module is configured to sort all data packets obtained by each routing port in the following manner: For all the data packets obtained by each routing port, sort them in descending order based on the loop level of the data carried by each data packet to obtain the initial data packet sequence of each routing port; For the initial data packet sequence of each routing port, the data packets at the same cycle layer are arranged in descending order of depth to obtain the final data packet sequence of each routing port.

7. The data stream array accelerator according to claim 5, characterized in that: In the calculation unit, the preset method includes: Obtaining the target computing unit to which the data packet processed by the computing unit needs to be transmitted; Determine a transmission path of the data packet in the on-chip network based on the acquisition target computing unit; The routing port in the routing module that is connected to the determined transmission path is used as the routing port for forwarding data packets.

8. A program execution method for executing a specified program, characterized in that: The method comprises: Step S1, obtaining a data flow graph compiler, and using the data flow graph compiler to compile a specified program to obtain a data flow graph of the specified program; Step S2: obtaining the array accelerator according to any one of claims 1 to 7, and configuring the routing port of each routing module of the array accelerator as a fixed-direction routing port or a variable-direction routing port based on a data flow graph of a specified program; Step S3: Map the data flow graph of the specified program to the configured array accelerator, so that the array accelerator executes the specified program.

9. A computer-readable storage medium, characterized in that A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of the method according to claim 8.

10. An electronic device, characterized in that: include: one or more processors; as well as a memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method of claim 8 by executing the executable instructions.