Data processing array interface having an interface style with a plurality of direct memory access circuits

The integration of multiple DMA circuits and interface styles within the array interface of an IC data processing array addresses the challenges of low data throughput and high latency, achieving improved performance through enhanced data transfer capabilities.

JP2025516711APending Publication Date: 2025-05-30XILINX INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024567543
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-17
Filing Date
2023-03-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing integrated circuits (ICs) with data processing arrays face challenges in achieving high data throughput and reducing latency in data transfers due to limitations in their array interfaces.

Method used

The proposed solution involves an integrated circuit (IC) with a data processing array that includes a plurality of computing tiles arranged in a grid, coupled with an array interface that features multiple interface styles. Each interface style includes a stream switch connected to adjacent compute or memory tiles, and multiple direct memory access (DMA) circuits connected to a network-on-chip (NoC) via independent communication channels.

Benefits of technology

This configuration enhances data throughput and reduces latency by allowing simultaneous data transfers and independent communication channels for each DMA circuit, thereby improving the overall performance of applications executed on the data processing array.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025516711000001_ABST
    Figure 2025516711000001_ABST
Patent Text Reader

Abstract

An integrated circuit (IC) can include a data processing array that includes a plurality of computing tiles arranged in a grid. The IC can include an array interface coupled to the data processing array. The array interface includes a plurality of interface tiles. Each interface tile includes a plurality of direct memory access circuits. The IC can include a network-on-chip (NoC) coupled to the array interface. Each direct memory access circuit is communicatively linked to the NoC via an independent communication channel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to integrated circuits, and more particularly to an array interface for a data processing array, the array interface including an interface style having a plurality of direct memory access circuits.

Background Art

[0002] Integrated circuits (ICs) have evolved over time to provide increasingly sophisticated computing architectures. Some ICs utilize a computing architecture that includes a single processor, while other ICs include multiple processors. Further, other ICs include multiple processors arranged in an array. Such ICs can provide significant computing power and a high degree of parallelism that far exceeds the capabilities of single-processor architectures and even multi-core processor architectures.

Summary of the Invention

[0003] In one or more exemplary implementations, an integrated circuit (IC) can include a data processing array that includes a plurality of computing tiles arranged in a grid. The IC can include an array interface coupled to the data processing array. The array interface includes a plurality of interface styles. Each interface style includes a plurality of direct memory access circuits. The IC can include a network-on-chip (NoC) coupled to the array interface. Each direct memory access circuit is communicatively linked to the NoC via an independent communication channel.

[0004] In one or more exemplary implementations, an array interface for an IC can include multiple interface styles. Each interface style can include a stream switch connected to an adjacent compute tile of the data processing array or another stream switch of an adjacent memory tile of the data processing array. The stream switch is also connected to the stream switch of at least one adjacent interface style of the array interface. Each interface style can include a stream multiplexer-demultiplexer connected to the stream switch of the same interface style. Each interface style can include multiple direct memory access circuits. Each direct memory access circuit is connected to the stream multiplexer-demultiplexer of the same interface style and the NoC. The direct memory access circuit is configured to convert a data stream received from the stream multiplexer-demultiplexer into a memory-mapped transaction transmitted to the NoC. The direct memory access circuit is also configured to convert a memory-mapped transaction received from the NoC into a data stream transmitted to the stream multiplexer-demultiplexer.

[0005] This summary section is provided merely to introduce certain concepts and is not provided to identify any key or essential features of the claimed subject matter. Other features of the configurations of the present invention will become apparent from the accompanying drawings and the following detailed description.

[0006] The configurations of the present invention are illustrated by way of example in the accompanying drawings. However, the drawings should not be construed as limiting the configurations of the present invention to the specific implementations shown. Various aspects and advantages will become apparent upon review of the following detailed description and with reference to the drawings.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0008] The present disclosure relates to integrated circuits (ICs), and more particularly to an array interface for a data processing (DP) array, including an interface style having a plurality of direct memory access (DMA) circuits. The DP array may include a plurality of tiles. The tiles may include compute tiles or a mix of compute tiles and memory tiles. The DP array is configurable to perform desired compute activity by loading configuration data into the DP array. Once configured, the DP array can perform compute activity. The configuration data loaded into the DP array, sometimes referred to as an “application,” may specify various operating parameters of the DP array, including, but not limited to, specific kernels to be executed by the compute tiles, connectivity between various tiles of the DP array, and the like.

[0009] When performing compute activity, the DP array depends on the array interface to obtain input data from other circuits and / or systems and provide the generated output data to such other circuits and / or systems. The other circuits and / or systems may be located within the same IC as the DP array or within a system external to the IC in which the DP array is located. According to the configuration of the present invention described within the present disclosure, the array interface is configured to include additional DMA circuits that increase the data throughput both inside and outside of the DP array. The additional DMA circuits facilitate transferring a greater amount of data per unit time both inside and outside of the DP array. This enhanced ability of the array interface leads to a reduction in latency in data transfers to and from the DP array, thereby increasing the overall throughput of the applications executed on the DP array.

[0010] Further aspects of the configuration of the present invention are described below with reference to the figures. FIG. 1 illustrates an exemplary system 100. In one or more exemplary implementations, system 100 is an electronic system that is fully implemented within a single IC. System 100 may be implemented within a single IC package. In one aspect, system 100 is implemented using a single die disposed within a single IC package. In another aspect, system 100 is implemented using two or more interconnected dies disposed within a single IC package.

[0011] System 100 includes a DP array 102, an array interface 104, a network-on-chip (NoC) 106, and a plurality of subsystems 108, 110, 112, 114, and 116. In this example, the DPE array 102 is coupled to the NoC 106 through the array interface 104. The subsystems 108 - 116 are coupled to the NoC 106. The DP array 102 may communicate with one or more or all of the subsystems 108 - 116 via the array interface 104 and the NoC 106.

[0012] It should be understood that the number of subsystems illustrated in the example of FIG. 1 is for illustrative purposes. System 100 may include more or fewer subsystems than those shown. The different subsystems may represent any combination of various different types of electronic subsystems and / or circuits. For illustrative purposes, the examples of subsystems 108 - 116 may include, but are not limited to, any combination such as a processor system, programmable logic, hardwired circuit blocks (e.g., application-specific circuit blocks), etc.

[0013] Figure 2 illustrates another exemplary implementation of system 100. In the example of Figure 2, the DP array 102, the array interface 104, and the NoC 106 are implemented within the same IC. In this example, an array controller 202 is illustrated that can also be located within the same IC as the DP array 102. The memory 204 and the processor 206 may or may not be disposed within the same IC as the DP array 102. In some examples, the IC implementation of system 100 includes the memory 204 and / or the processor 206. In other examples, the IC implementation of system 100 does not include the memory 204 and / or the processor 206, and the memory 204 and / or the processor 206 are implemented within a device that is external to the IC containing the DP array 102 and coupled to the IC.

[0014] The DP array 102 is formed from a plurality of different types of circuit blocks referred to as tiles. As defined within this disclosure, the term "array tile" means a circuit block included in the DP array. The array tiles of the DP array 102 may include only compute tiles, or a mix of compute tiles and memory tiles. The compute tiles and the memory tiles are hardwired and programmable. The array interface 104 includes a plurality of interface tiles. As defined within this disclosure, the term "interface tile" means a circuit block included in the array interface. The interface tiles of the array interface 104 communicatively link the array tiles of the DP array 102 to communicate with circuits external to the DP array 102, regardless of whether such circuits are located on the same die, a different die within the same IC package, or external to the IC package. The interface tiles are also hardwired and programmable. As defined within this disclosure, the term "tile" means an array tile and / or an interface tile. Thus, a tile may refer to a compute tile, a memory tile, or an interface tile.

[0015] The array controller 202 is optionally included in the system 100. The array controller 202 is communicatively linked to the DP array 102 and also to the array interface 104. In one aspect, the array controller 202 is dedicated to controlling the operation of the DP array 102 and the array interface 104. The array controller 202 may be implemented as a state machine (e.g., a hardened controller) or as a processor. Regardless of whether it is implemented as a state machine or a processor, the array controller 202 may be implemented as a hardwired circuit block or using programmable logic.

[0016] The NoC 106 is a programmable interconnect network for sharing data between endpoint circuits within the IC. The endpoint circuits can be located within the DP array 102 or within other subsystems 108 - 116. In the example of FIG. 2, one or both of the memory 204 and / or the processor 206 are an example of a subsystem such as the subsystems 108 - 116 of FIG. 1. The NoC 106 can include a high-speed data path with dedicated switching. The NoC 106 also supports network connections. In one example, the NoC 106 includes one or more horizontal paths, one or more vertical paths, or both horizontal and vertical paths. The NoC 106 is an example of a common infrastructure available within the IC for connecting selected components and / or subsystems.

[0017] Generally, the programmability of NoC 106 means that the nets to be routed through NoC 106 are unknown until the design (e.g., application) is created for implementation within system 100. NoC 106 may be programmed by loading configuration data into internal configuration registers that define how the elements within NoC 106, such as switches and interfaces, are configured and operate to pass data between switches and between the ingress / egress (e.g., interface) circuits of NoC 106 that connect to endpoint circuits. NoC 106 is fabricated as part of system 100 (e.g., hardwired) and is programmed to establish network connectivity between different master circuits and different slave circuits of the user circuit design.

[0018] In one or more examples, NoC 106 may implement a selected data path used when configuring system 100 upon power-up. Upon power-up, NoC 106 cannot implement a data path or route for implementing the user application therein until NoC 106 is configured using configuration data that at least specifies the path of the nets for connecting the endpoint circuits of the user application or design.

[0019] Memory 204 may be implemented as a random-access memory (RAM). In one or more exemplary implementations, memory 204 may be implemented within the same integrated circuit (IC) that includes DP array 102, and may be embedded, for example. For example, memory 204 may be a RAM circuit implemented on the same die as DP array 102, or on a different die within the same IC package. For example, memory 204 may be implemented as a High Bandwidth Memory (HBM). In another aspect, memory 204 is external to the IC that includes DP array 102. For example, memory 204 may be one or more RAM modules communicatively linked to the IC that includes DP array 102 (e.g., located on the same circuit board as the IC).

[0020] In one aspect, processor 206 is implemented within the same IC that includes DP array 102 and is embedded, for example. Processor 206 may be implemented as a hardwired processor within the IC, or may be implemented using programmable logic. In another aspect, processor 206 is external to the IC that includes DP array 102. In that case, processor 206 may be part of another data processing system, such as a host computer communicatively linked to the IC that includes DP array 102.

[0021] In the example of FIG. 2, the DP array 102 and the array interface 104 may operate under the control of another circuit. That is, another circuit such as the processor 206 and / or the array controller 202 may control the configuration and operation of the DP array 102 and / or the array interface 104 over time. If the system 100 includes both the processor 206 and the array controller 202, the processor 206 may execute an application and provide instructions, such as tasks or jobs, to the array controller 202. The array controller 202 may execute the instructions to control the configuration and / or operation of the DP array 102. In other configurations, the array controller 202 may be omitted so that the processor 206 controls the configuration and / or operation of the DP array 102. In that case, when the processor 206 is implemented within the same IC as the DP array 102 and the array interface 104, it may include one or more direct connections to the DP array 102 and / or the array interface 104.

[0022] Figure 3 illustrates an exemplary implementation of the DP array 102 and the array interface 104. In this example, the DP array 102 includes compute tiles 302 and memory tiles 306. In the example of Figure 3, the compute tiles 302 and the memory tiles 306 are arranged in a grid having a plurality of rows and columns. The interface tiles 304 are arranged in rows, and individual interface tiles 304 are aligned with the columns of the grid configuration of the DP array 102. The compute tiles 302 include compute tiles 302-1, 302-2, 302-3, 302-4, 302-5, 302-6, 302-7, 302-8, 302-9, 302-10, 302-11, 302-12, 302-13, 302-14, 302-15, 302-16, 302-17, and 302-18. The interface tiles 304 include interface tiles 304-1, 304-2, 304-3, 304-4, 304-5, and 304-6. The memory tiles 306 include memory tiles 306-1, 306-2, 306-3, 306-4, 306-5, and 306-6. In this example, each tile is coupled to adjacent tiles to the left (west), right (east), above (north), and below (south) when the tile is so positioned.

[0023] The example of Figure 3 is provided for illustrative purposes only. The number of tiles in a given column and / or row, the number of tiles included in the DP array 102 and / or the array interface 104, the sequence or order of tile types (e.g., memory and compute tiles) in a column and / or row are for illustrative purposes and not limiting. Other configurations having various numbers of tiles, rows, columns, mixtures of tile types, etc. may be included. For example, the rows of Figure 3 are homogeneous with respect to tile type, but the columns are not. In other configurations, the rows may be heterogeneous with respect to tile type, but the columns may be homogeneous. Further, additional rows of memory tiles 306 may be included in the DP array 102. Such rows of memory tiles 306 may be grouped together without interposing rows of compute tiles 302, or the rows of compute tiles 302 may be dispersed throughout the DP array 102 such that they interpose between rows of memory tiles 306 or groups of rows.

[0024] In another exemplary implementation of the DP array 102, the memory tile 306 may be omitted such that the bottom row of the compute tile 302 is directly coupled to the interface tile 304. For example, when the memory tile 306 is omitted, the interface tile 304-1 will be directly connected to the compute tile 302-3, etc. In such a case, in various exemplary implementations described herein, instead of the memory tile 306, data may be read from the memory 204 and written to the memory 204. However, including the memory tile 306 can increase the data throughput of the DP array 102 in that data can be stored closer to the compute tile 302 without the need to continuously read data from and / or write data to the external RAM of the DP array 102.

[0025] In the example of FIG. 3, each of the interface tiles 304 includes a plurality of DMA circuits 308. For illustrative purposes, each interface tile 304 includes two DMA circuits 308. The interface tile 304-1 includes the DMA circuits 308-1, 308-2, the interface tile 304-2 includes the DMA circuits 308-3, 308-4, the interface tile 304-3 includes the DMA circuits 308-5, 308-6, the interface tile 304-4 includes the DMA circuits 308-7, 308-8, the interface tile 304-6 includes the DMA circuits 308-9, 308-10, and the interface tile 304-6 includes the DMA circuits 308-11, 308-12.

[0026] Each of the DMA circuits 308 can communicate with the NoC 106 via an independent communication channel as shown. In one aspect, each DMA circuit 308 is connected to the interface circuit of the NoC 106. Each DMA circuit 308 may be connected to a different independent interface circuit of the NoC 106. The DMA circuits 308 within the same interface style 304 can operate simultaneously such that each can transmit one or more data streams to the other simultaneously. This enables each interface style 304 to transmit data to the NoC 106 simultaneously using both DMA circuits 308 therein and / or to transmit data streams to the DP array 102 simultaneously using both DMA circuits 308 therein.

[0027] FIG. 4 illustrates an exemplary implementation of the compute tile 302. The example of FIG. 4 is provided to illustrate features on a particular architecture of the compute tile 302 and is not generally intended to limit the form of the DP array 102 or the architecture of the compute tile 302. Some connections between components and / or tiles are omitted to facilitate illustration.

[0028] In this example, each compute tile 302 includes a core 402, a RAM 404, a stream switch 406, a memory-mapped switch 408 (e.g., abbreviated as the "MM" switch in the figure), an event broadcast switch 410, event logic 412, control registers 414, and a direct memory access (DMA) circuit 434. The core 402 includes a processor 420 and a program memory 422. The control registers 414 can be written by the memory-mapped switch 408 to control the operation of various components included in the compute tile 302. Although not shown, each memory component of the compute tile 302 (e.g., program memory 422, control registers 414, and RAM 404) can be read and / or written via the memory-mapped switch 408 for configuration and / or initialization purposes.

[0029] Processor 420 can be any of a variety of different processor types. In one aspect, processor 420 is implemented as a vector processor. Program memory 422 may be loaded with one or more sets of executable instructions, generally referred to as "kernels", for example, by loading configuration data. Compute tile 302 is capable of performing data processing operations and manipulating large amounts of data by executing kernels.

[0030] Each core 402, such as processor 420, is directly connected to RAM 404 located within the same compute tile 302 through memory interface 432. Within the present disclosure, when a memory interface is used by circuitry within the same tile to access RAM, the memory interface is referred to as a "local memory interface". Memory interface 432-1 is an example of a local memory interface because the processor 420 within the same tile uses the memory interface to access RAM 404. In comparison, a memory interface used by circuitry external to the tile to access RAM 404 is referred to as an adjacent memory interface. Memory interfaces 432-2, 432-3, and / or 432-4 are examples of adjacent memory interfaces because such memory interfaces are used by circuitry within other adjacent tiles to access RAM 404.

[0031] Thus, each processor 420 can access (e.g., read and / or write) the RAM 404 within the same compute tile 302 and one or more other RAMs 404 within adjacent tiles via standard read and write operations directed to the memory interface. The processor 420 can execute program code stored in the program memory 422. The RAM 404 is configured to store application data. The RAM 404 can be read and / or written via the memory mapped switch 408 for setup and / or initialization purposes. The RAM 404 can be read and / or written by the processor 420 and / or by the DMA circuit 434 during runtime.

[0032] The DMA circuit 434 can read and write data to the RAM 404 located within the same compute tile 302. The DMA circuit 434 can receive data from a source external to the compute tile 302 via the stream switch 406 and store such data in the RAM 404. The DMA 434 can read data from the RAM 404 and output the data to the stream switch 406 to transfer it to one or more other destinations external to the compute tile 302.

[0033] Each core 402, e.g., processor 420, can be directly connected via a memory interface to a RAM 404 located within an adjacent compute tile 302 (e.g., in a north, south, east, and / or west direction). Thus, processor 420 can directly access such other adjacent RAM 404 in the same manner that processor 420 can access a RAM 404 located within the same compute tile 302 without initiating a read or write transaction via a stream switch 406 and / or without using a DMA circuit 434. As an illustrative example, the processor 420 of compute tile 302-5 can perform reads and / or writes to the RAM 404 located within compute tiles 302-5, 302-2, 302-4, and 302-6 without submitting a read or write transaction via a stream switch 406 and / or without using a DMA circuit 434. However, it should be understood that processor 420 can initiate read and write transactions to the RAM 404 of any other compute tile 302 and / or memory tile 306 via a stream switch 406 and a DMA circuit 434.

[0034] Processor 420 may also include a direct connection, referred to as a cascade connection (not shown), to a processor 420 of an adjacent core (e.g., in a north, south, east, and / or west direction) that enables direct sharing of data stored in an internal register (e.g., an accumulator register) of processor 420 with other processors 420. This means that data stored in one or more internal registers of one processor 420 may be directly transferred to one or more internal registers of a different processor 420 without first writing such data to a RAM 404 and / or without transmitting such data through a stream switch 406.

[0035] The event logic 412 may also be configured by the control register 604. In the example of FIG. 4, the event logic 412 is coupled to the event broadcast switch 410. The configuration data loaded into the control register 414 defines specific events that may be locally detected within the compute tile 302.

[0036] For each control register 414, the event logic 412 is capable of detecting various different events that originate from and / or pertain to the DMA circuit 434, the memory-mapped switch 408, and / or the stream switch 406. Examples of events may include, but are not limited to, DMA completion transfers, unlocks (e.g., of the RAM 404 not shown), locks (e.g., of the RAM 404), or other events related to the start or end of data flow through the compute tile 302. The event logic 412 may be configured by the control register 414 within the compute tile 302. The event logic 412 may provide such events to the event broadcast switch 410. As shown, the event broadcast switch 410 is coupled to the event broadcast switch (if available) of each of the tiles adjacent to the west, east, north, and south. Thus, the event broadcast switch 410 is capable of propagating events internally generated by the event logic 412 and / or events received from outside the compute tile 302 to other tiles within the DP array 102 and / or the array circuit 104.

[0037] In one or more exemplary implementations, the event broadcast switch 410 can collect broadcast events from one or more or all directions. In some cases, the event broadcast switch 410 can perform a logical “OR” of signals and transfer the result in one or more or all directions. Each output from the event broadcast switch 410 may include a bitmask configurable by configuration data loaded into the control register 414. The bitmask determines which events are broadcast individually in each direction. Such a bitmask can, for example, eliminate unwanted or duplicate propagation of events.

[0038] FIG. 5 illustrates an exemplary implementation of the memory tile 306. The example of FIG. 5 is provided to illustrate features on a particular architecture of the memory tile 306 and is not generally intended to limit the form of the DP array 102 or the architecture of the memory tile 306. Some connections between components and / or tiles are omitted to facilitate illustration. Generally, the memory tile 306 is characterized by the absence of a processor therein.

[0039] Each memory tile 306 includes a stream switch 406, a memory-mapped switch 408, a DMA circuit 502, a RAM 504, an event broadcast switch 410, event logic 412, and / or a control register 414. The control register 414 may be written by the memory-mapped switch 408 to control the operation of the various components illustrated in the memory tile 306. Although not shown, each memory component (e.g., RAM 504 and control register 414) of the memory tile 306 may be read and / or written via the memory-mapped switch 408 for configuration and / or initialization purposes.

[0040] Each DMA circuit 502 of the memory tile 306 is coupled to the RAM 504 within the same memory tile 306 via the local memory interface 532-1 and can be coupled to one or more RAMs 504 of other adjacent memory tiles 306. In the example of FIG. 5, each DMA circuit 502 can access (e.g., read and / or write to) the RAM 504 included within the same memory tile 306 via the local memory interface 532-1. The RAM 504 includes the adjacent memory interfaces 532-2 and 532-3 through which the DMA circuits of the east and west memory tiles 306 can access the RAM 504. For example, the DMA circuit 502 of the memory tile 306-2 can access the RAM 504 of the memory tile 306-1 and / or the RAM 504 of the memory tile 306-3. The DMA circuit 502 can place the data read from the RAM 504 onto the stream switch 406 and can write the data received via the stream switch to the RAM 504.

[0041] The event broadcast switch 410 and the event logic 412 within the memory tile 306 may operate substantially the same as the event broadcast switch 410 and the event logic 412 described in connection with FIG. 4. Similarly, if the memory mapped switch 408 is used for the purpose of configuring and initializing the memory tile 306, the stream switch 406 is used to transfer data during runtime.

[0042] FIG. 6 illustrates an exemplary implementation of interface style 304 of array interface 104. In the example of FIG. 6, interface style 304 includes a stream switch 406, a memory-mapped switch 602, a control register 604 (e.g., abbreviated as "CRS" 604 in FIG. 6), event logic 606, an event broadcast switch 608, a control, debug, and trace (CDT) circuit 610, an interrupt handler 612, a bridge 614, a stream multiplexer-demultiplexer 616, a plurality of DMA circuits 618, a plurality of NoC stream interfaces 620, and a plurality of switches 622.

[0043] Memory-mapped switch 602 may be connected to memory-mapped switches in tiles adjacent to the west, east, and north. For example, memory-mapped switch 602 may include one or more memory-mapped interfaces, and the memory-mapped interfaces have masters and slaves that connect to memory-mapped switches in tiles adjacent to the west, east, and north. Different from the memory-mapped switch 408 of the array tile, the memory-mapped switch 602 of interface style 304 is capable of communicating with memory-mapped switches within adjacent west and east interface styles 304.

[0044] Thus, memory-mapped switch 602 can move data (e.g., configuration, control, and / or debug data) from one interface style 304 to another interface style 304 to reach a particular or correct interface style 304 and direct that data northward to a particular target array tile or multiple array tiles. For example, if a memory-mapped transaction is received from another circuit within a particular interface style 304, memory-mapped switch 602 can disperse the transaction horizontally, e.g., to other interface styles 304 within array interface 104.

[0045] In this example, the control register 604 is connected to the memory-mapped switch 602 and may be read and / or written through the memory-mapped switch 602. Through the memory-mapped switch 602, configuration data may be loaded into the control register 604 to control various functions and operations executed by various components within the interface style 304. For illustrative purposes, the control register 604 is not shown as being coupled to various circuit blocks of the interface style 304. However, it should be understood that the control register 604 may be connected to such circuit blocks via control signals in order to control the operating characteristics of various circuit blocks of the interface style 304 (e.g., event logic 606, event broadcast switch 608, CDT 610, interrupt handler 612, stream switch 406, bridge 614, stream multiplexer-demultiplexer 616, DMA circuit 618, NoC stream interface 620, and / or switch 622).

[0046] The memory-mapped switch 602 is coupled to the NoC 106 via the bridge 614. The bridge 614 is capable of converting a memory-mapped data transfer from the NoC 106 corresponding to configuration data, control data, and / or debug data into memory-mapped data that may be received by the memory-mapped switch 602.

[0047] The event logic 606 may be constituted by the control register 604. In the example of FIG. 6, the event logic 606 is coupled to an event broadcast switch 608 and a CDT circuit 610. The configuration data loaded into the control register 604 defines specific events that may be locally detected within the interface style 304. The event logic 606 is capable of detecting various different events that occur from and / or relate to the DMA circuit 618, the memory mapped switch 602, the stream switch 406, and / or the NoC stream interface 620 for each control register 604. Examples of events may include, but are not limited to, DMA completion transfers, unlocking, locking, or other events related to the start or end of data flow through the interface style 304. The event logic 606 may provide such events to the event broadcast switch 608 and / or the CDT circuit 610. In another exemplary implementation, the event logic 606 may not have a direct connection to the CDT circuit 610, but rather is connected to the CDT circuit 610 via the event broadcast switch 608.

[0048] The event broadcast switch 608 and the event logic 606 may operate in the same manner as the event broadcast switch and logic described in relation to the compute tile 302 and the memory tile 306. In the example of FIG. 6, the event broadcast switch 608 provides an interface between an event broadcast network formed from the event broadcast switches within the DP array 102 and the interface style 304. The event broadcast switch 608 is coupled to the event broadcast switches in the tiles adjacent to the west, east, and north of the interface style 304.

[0049] The event broadcast switch 608 can send events generated internally by the event logic 606, events received from other interface styles 304, and / or events received from array tiles (e.g., compute tile 302 and / or memory tile 306) to other tiles. Since events may be broadcast between interface styles 304, an event can traverse through the event broadcast switch 608 in an interface style 304 to a specific interface style 304 and then traverse to one or more targets (e.g., intended) array tiles within the DP array 102, thereby reaching any array tile within the DP array 102.

[0050] In another example, an event may also be sent from the event broadcast switch 608 to other circuit blocks and / or subsystems within the same IC. The event data may be transferred to the CDT 610 and transferred onto the stream switch 406 for transmission through the NoC 106. For example, the CDT circuit 610 can packetize the received event and send the packetized event to the stream switch 406.

[0051] In one or more exemplary implementations, the event broadcast switch 608 can collect broadcast events from one or more or all directions. In some cases, the event broadcast switch 608 can perform a logical "OR" of the signals and transfer the result to one or more or all directions (e.g., including the CDT circuit 610). Each output from the event broadcast switch 608 may include a bitmask configurable by the configuration data loaded into the control register 604. The bitmask determines which events are broadcast individually to each direction. Such a bitmask can, for example, eliminate unwanted propagation or duplicate propagation of events.

[0052] The interrupt handler 612 is coupled to the event broadcast circuitry 608 and is capable of receiving events broadcast from the event broadcast circuitry 608. In one or more exemplary implementations, the interrupt handler 612 may be configured by configuration data loaded into the control register 604 to generate an interrupt to the NoC 106 or another interconnect on the IC in response to selected events and / or combinations of events from the event broadcast switch 608 (e.g., events generated by the DP array 102 and / or events generated within the interface style 304). The interrupt handler 612 may be capable of generating an interrupt to circuitry and / or subsystems of the same IC that are external to the DP array 102 and / or the array interface 104 based on the configuration data. In another aspect, the interrupt may be transmitted to a system external to the IC in which the DP array 102 and / or the array interface 104 are implemented. In one aspect, the interrupt may be transmitted directly to such other circuitry and / or subsystems via an interrupt circuit (not shown). In another aspect, the interrupt handler 612 may be connected to the NoC 106 and transmit the interrupt via the NoC 106.

[0053] The stream switch 406 is connected to the stream switches in the tiles adjacent to the west, east, and north. The stream switch 406 is also coupled to the DMA circuit 618 and the NoC stream interface 620 through a stream multiplexer-demultiplexer (M-D) 616. The stream switch 406 may be configurable by configuration data loaded into the control register 604. The stream switch 406 may be configured, for example, to support packet switching operations and / or circuit switching operations based on the configuration data. Further, the configuration data defines the specific tile with which the stream switch 406 communicates, regardless of whether it is an array tile or another interface tile. Generally, referring to the DP array 102 and the array interface 104, the configuration data loaded into the various tiles defines the logical connections established between the stream switches 406 during operation (e.g., runtime operation). For example, referring to the interface tile 304, the configuration data defines a specific array tile within the column of array tiles immediately above the interface tile 304 with which the stream switch 406 communicates.

[0054] Stream M-D616 may include a multiplexer that receives data streams from DMA circuit 618-1, DMA circuit 618-2, NoC stream interface 620-1, and / or NoC stream interface 620-2. Stream M-D616 can multiplex the received data streams and transfer the multiplexed data streams to stream switch 406. Stream switch 406 selectively transmits different ones of the received data streams to adjacent tiles to the west, east, or north. Stream M-D616 may include a demultiplexer that receives data streams from stream switch 406 and selectively demultiplexes the data streams to transmit different ones of the received data streams to DMA circuit 618-1, DMA circuit 618-2, NoC stream interface 620-1, and / or NoC stream interface 620-2. For example, stream M-D616 may be programmed by configuration data stored in control register 604 that indicates which data streams to route to DMA circuit 618-1, DMA circuit 618-2, NoC stream interface 620-1, and / or NoC stream interface 620-1.

[0055] The DMA circuit 618 converts the data stream from the stream M-D616 into a memory-mapped transaction transmitted onto the NoC 106. Similarly, the DMA circuit 618 converts the memory-mapped transaction received from the NoC 106 into a data stream transferred to the stream M-D616. The NoC stream interface 620 transfers the data stream received from the NoC 106 to the stream M-D616. Similarly, the NoC stream interface 620 transfers the data stream received from the stream M-D616 to the NoC 106. For example, the NoC stream interface 620 transfers the data stream received from the stream M-D through the data portion of the communication bus (e.g., without using the address bus portion of the bus) to the NoC 106 via the switch 622. The address information may be embedded, for example, in the data stream itself.

[0056] Each of the switches 622 receives data from the attached DMA circuit 618 or the attached NoC stream interface 620, arbitrates between the data, and transfers it to the NoC 106. In this example, the switch 622 is coupled to the NoC master circuit (NMC) 650. The NoC slave circuit (NSC) 652 is connected to the NoC stream interface 620 and the bridge 614. In one or more exemplary implementations, the DMA circuit 618 may include a hardware synchronization circuit that is used to synchronize one or more channels included in each of the DMA circuits 618 with a master that polls and drives lock requests. For example, the master may be another circuit or subsystem within an IC in which the interface style 304 is disposed. The master may also receive an interrupt generated by the hardware synchronization circuit within the DMA circuit 618.

[0057] In one or more exemplary implementations, the DMA circuit 618 can access the memory. As discussed, one or more of the subsystems 108-116 may implement such a memory. Alternatively, the memory may be located external to the IC. For example, the DMA circuit 618 and / or the NoC stream interface 620 can receive a data stream from an array tile of the DP array 102 and transmit the data stream through the NoC 106 to the memory. Similarly, the DMA circuit 618 can receive memory-mapped data originating from the memory and / or other subsystems and provide the data in the form of a data stream to the stream M-D616 for transmission to other tiles. The NoC stream interface 620 can receive a data stream originating from the memory and / or other subsystems and provide the data stream to the stream M-D616 for transmission to other tiles.

[0058] In one or more exemplary implementations, the DMA circuit 618 may include security bits that may be set using some of the security registers included in the interface style 304. The security registers may be included in the control register 604 or may be implemented separately and independently from the control register 604. In one aspect, the memory and / or subsystem communicatively linked to the interface style 304 via the NoC 106 may be divided into different regions or partitions, called DP arrays 102 or some of its array tiles called partitions, where only certain regions of the memory and / or subsystem are permitted to be accessed. The security bits in the DMA circuit 618 may be set such that the array tiles of the DP array 102 can access only the specific regions of the memory and / or subsystem permitted for each security bit via the DMA circuit 618. For example, an application implemented in the DP array 102 or its partition (e.g., a subset of array tiles) may be limited to using this mechanism to access only specific regions of the memory and / or subsystem, may be limited to reading only from specific regions of the memory, and / or may be limited from writing to the memory.

[0059] The security bits in the DMA circuit 618 that control access to the memory and / or other subsystems may be implemented to control the entire DP array 102 or may be implemented in a more granular manner where access to the memory and / or other subsystems is specified and / or controlled on a per-array tile basis.

[0060] The CDT circuit 610 is capable of performing control operations, debug operations, and trace operations within the interface style 304. Regarding debugging, each of the registers located within the interface style 304 is mapped onto a memory map accessible via the memory-mapped switch 602. The CDT circuit 610 may include circuits such as, for example, trace hardware, a trace buffer, performance counters, and / or stall logic. The trace hardware of the CDT circuit 610 is capable of collecting trace data. The trace buffer of the CDT circuit 610 is capable of buffering trace data. The CDT circuit 610 is further capable of outputting the trace data to the stream switch 406.

[0061] In one or more exemplary implementations, the CDT circuit 610 is capable of collecting data, such as trace data and / or debug data, packetizing such data, and then outputting the packetized data through the stream switch 406. For example, the CDT circuit 610 is capable of outputting the packetized data and providing such data to the stream switch 406. The control register 604 can be read or written during debugging via a memory-mapped transaction through the memory-mapped switch 602 of each interface style. Similarly, the performance counters within the CDT circuit 610 can be read or written during profiling via a memory-mapped transaction through the memory-mapped switch 602 of each interface style 304.

[0062] In one or more exemplary implementations, the CDT circuit 610 can receive any event propagated by the event broadcast switch 608, or an event selected for each bitmask utilized by the event broadcast switch 608 coupled to the CDT circuit 610. The CDT circuit 610 can further receive events generated by the event logic 606. For example, the CDT circuit 610 can receive broadcast events from the array tile, interface style 304. The CDT circuit 610 can pack a plurality of such events into a packet, for example packetize them, and associate the packetized events with timestamps. The CDT circuit 610 can further transmit the packetized events to a destination external to the interface style 304 through the stream switch 406. The events can be transmitted to the NoC 106 via the stream switch 406 and the stream M-D 616, and / or via the DMA circuit 618 and / or the NoC stream interface 620.

[0063] In the examples of FIGS. 4, 5, and 6, the stream switch 406 may be implemented using a greater number of ports and channels to accommodate the increased connectivity of the interface style 304 to the NoC 106. The increased number of ports and channels enables the stream switch 406 to ensure that a higher level of data throughput provided by the interface style 304 continues into the DP array 102.

[0064] Figure 7 illustrates another exemplary implementation of interface style 304 of the array interface 104. In the example of Figure 7, switches 702 are added. As shown, switches 702-1 and 702-2 are added. Switch 702-1 is connected to switches 622-1 and NMC650-2. Switch 702-1 also connects to switch 702-2 in the interface style adjacent to the east. Switch 702-2 is connected to switches 622-2 and NMC650-2. Switch 702-2 also connects to switch 702-1 in the interface style adjacent to the west.

[0065] The switches 702 provide additional flexibility for routing the data flow entering and leaving the DP array 102 with respect to the NoC 106. That is, since the nets for the user application are routed through the NoC 106 and use the corresponding entry / exit points for the NMC650 and the NSC652 respectively, including the switches 702 provides greater flexibility when routing the nets through the NoC 106. The additional switching capabilities provided within the interface style 304 by including the switches 702 can reduce routing congestion that may occur by enabling data from a particular interface style 304 to enter the NoC 106 via a different interface circuit than would be the case if the switches 702 were omitted. The NMC650 is an example of an interface circuit of the NoC 106.

[0066] For example, data from the DMA circuit 618 or the NoC stream interface 620 may enter the NoC 106 based on the programming of the switch 702-1, via the NMC 650-1, or via the NMC 650 within the interface style adjacent to the east. Thus, programming the switches 702-1 and 702-2 within the tile adjacent to the east provides a selection mechanism for determining where the DMA circuit 618-1 and / or the NoC stream interface 620-1 couple to the NoC 106. In this regard, the switch 702 may be referred to as a "selector switch" in that programming such a switch selects the connection points of several circuits (e.g., the DMA circuit 618 and / or the NoC stream interface 620) to the NoC 106. The DMA circuit 618-1 and / or the NoC stream interface 620-1 may enter the NoC 106 via the NMC 650-1 or via another NMC 650 (not shown) located under the interface style 304 adjacent to the east. Thus, the switch 702 is programmable to connect different circuits among the interface circuits of the NoC 106 to different tiles among the interface tiles 304. This flexibility allows implementation tools to avoid overly congested interface circuits when routing nets into, out of, and / or through the NoC 106.

[0067] Referring to FIGS. 6 and 7, both exemplary interface styles 304 provide increased connectivity from the DP array 102 to the NoC 106. Since both exemplary interface styles 304 can be connected to the NoC 106 via different NMCs 650 and NSCs 652, routing congestion can be reduced. By including multiple DMA circuits 618 and multiple NoC stream interfaces 620, a larger amount of data can be transferred between the DP array 102 and the NoC 106 for each interface style 304. This means that the amount of data transferred overall between the NoC 106 and the DP array 102 (e.g., for a given one or more applications implemented on the DP array 102) becomes significantly larger. As discussed, for example, the ability to transfer a larger amount of data between any of the subsystems 108-116 and the DP array 102 increases data throughput and reduces latency. The data may be moved from memory to the DP array 102, for example, using a bandwidth that is twice the bandwidth if a single DMA circuit 618 and / or a single NoC stream interface 620 were included in the interface style 304.

[0068] Referring to both FIGS. 6 and 7, routing congestion is reduced by enabling data directed to a particular interface style 304 to exit the NoC 106 via the NSC652-1 or NSC652-2.

[0069] The flexibility added by including the switch 702 of FIG. 7 also helps to reduce additional congestion that may occur in the NoC 106 and access contention to the NMC 650 and / or NSC 652 by including multiple DMA circuits 618 and multiple NoC stream interfaces 620 that support a larger number of connections to the NoC 106 for each interface style. For example, if the bandwidth of a particular DMA circuit 618 and / or NoC stream interface 620 of a selected interface style 304 is not fully utilized, the switch 702 connected to that DMA circuit 618 and NoC stream interface 620 may be programmed to receive data from adjacent DMA circuits and / or stream switches.

[0070] FIG. 8 illustrates an exemplary implementation of the connectivity between interface styles 304 of the array interface 104 achieved by the switch 702. The example of FIG. 7 illustrates how each switch 702-1 (except for the switch 702-1 of the rightmost interface style 304-6 in FIG. 3) can be communicatively linked to the switch 702-2 of the interface style 304 adjacent to the east. Each switch 702-2 (except for the first switch 702-2 of the interface style 304-1) is connected to the switch 702-1 of the interface style 304 adjacent to the west. In this regard, the switch 702-2 of the interface style 304-1 may be omitted. Similarly, the switch 702-1 of the rightmost interface style 304 (e.g., the interface style 304-6) may be omitted.

[0071] In the example of FIG. 8, each interface style 304 is aligned with a column of array tiles 802 (e.g., compute tiles 302, or a combination of compute tiles 302 and memory tiles 306). That is, one interface style 304 is aligned with one column of array tiles 802. Each interface style 304 includes two DMA circuits (not shown).

[0072] FIG. 9 illustrates an exemplary implementation of the NoC 106. The NoC 106 includes an NMC 650, an NSC 652, a network 914, a NoC peripheral interconnect (NPI) 910, and a register 912. Each NMC 650 is an ingress circuit that connects an endpoint circuit to the NoC 106. Each NSC 652 is an egress circuit that connects the NoC 106 to an endpoint circuit. The NMC 650 is connected to the NSC 652 through the network 914. In one example, the network 914 includes a NoC packet switch 906 (NPS) and a routing 908 between the NoC packet switches 906. Each NoC packet switch 906 performs switching of NoC packets. The NoC packet switches 906 are connected to each other through the routing 908 and are also connected to the NMC 650 and the NSC 652 to implement a plurality of physical channels. The NoC packet switches 906 also support a plurality of virtual channels for each physical channel.

[0073] NPI910 includes circuitry for programming NMC650, NSC652, and NoC packet switch 906. For example, NMC650, NSC652, and NoC packet switch 906 can include registers 912 that determine their functionality. NPI910 includes peripheral interconnects coupled to registers 912 to program the registers 912 to set functionality. Registers 912 within NoC106 support interrupts, Quality of Service (QoS), error handling and reporting, transaction control, power management, and address mapping control. Registers 912 can be initialized to a usable state before being reprogrammed, such as by using write requests to write to registers 912. Configuration data for NoC106, DP array 102, and / or array interface 104 can be stored in non-volatile memory (NVM) as part of a programming device image (PDI) that specifies configuration data for various components of the IC and / or subsystems (such as those including DP array 102 and / or array interface 104), and provided to NPI910 to program NoC106 and / or other endpoint circuitry.

[0074] NMC650 is a traffic entry point. NSC652 is a traffic exit point. NMC650 and NSC652 are examples of interface circuitry of NoC106. Endpoint circuitry coupled to NMC650 and NSC652 can be enhancement circuitry. In another aspect, one or more of NMC650 and / or NSC652 can be circuitry implemented in programmable logic (such as when the programmable logic is included in the IC as one or more of subsystems 108 - 116). A given endpoint circuitry can be coupled to two or more NMC650s or two or more NMC650s.

[0075] FIG. 10 is a block diagram showing the connection between endpoint circuits within an IC through NoC 106 according to an example. In this example, endpoint circuit 1002 is connected to endpoint circuit 1004 through NoC 106. Endpoint circuit 1002 is a master circuit coupled to NMC 650 of NoC 106. Endpoint circuit 1004 is a slave circuit coupled to NSC 652 of NoC 106. Each of endpoint circuits 1002 and 1004 can be any of various circuits, such as, for example, interface style 304 and / or any of various subsystems 108 - 116. Examples of different subsystems may include, but are not limited to, a processor system, programmable logic, hardwired circuit blocks (such as digital - to - analog converters, analog - to - digital converters, transceivers, etc.). For example, one or more of endpoint circuits 1002 and / or 1004 may be processor 206 and / or memory 204.

[0076] Network 914 includes a plurality of physical channels 1006. Physical channels 1006 are implemented by programming NoC 106. Each physical channel 1006 includes one or more NoC packet switches 906 and associated routing 908. NMC 650 is connected to NMC 650 through at least one physical channel 1006. Physical channels 1006 can also have one or more virtual channels 1008.

[0077] The connection through network 914 uses a master - slave configuration. In one example, the most basic connection on network 914 includes a single master connected to a single slave. However, in other examples, more complex structures can be implemented.

[0078] The present disclosure concludes with claims that define novel features, but the various features described within the present disclosure are believed to be better understood when considered in conjunction with the drawings and the description. The processes, machines, manufactures, and any variations thereof described herein are provided for illustrative purposes. The specific structural and functional details described within the present disclosure are not to be construed as limiting, but rather as a representative basis for the claims and for teaching one skilled in the art to variously employ the features described in virtually any appropriately detailed structure. Further, the terms and phrases used within the present disclosure are not intended to be limiting, but rather are intended to provide an understandable description of the features being described.

[0079] For simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some of the elements may be exaggerated relative to other elements for clarity. Further, where appropriate, reference numerals may be repeated among the figures to indicate corresponding, similar, or like features.

[0080] As defined herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0081] As defined herein, the term “about” means nearly correct or close to an exact value or quantity but not exact. For example, the term “about” may mean that the recited characteristic, parameter, or value is within a predetermined amount of the exact characteristic, parameter, or value.

[0082] As defined herein, the terms "at least one", "one or more", and "and / or" are open-ended expressions that are both conjunctive and disjunctive in operation, unless explicitly stated otherwise. For example, each of the expressions "at least one of A, B, and C", "at least one of A, B, or C", "one or more of A, B, and C", "one or more of A, B, or C", and "A, B, and / or C" means A alone, B alone, C alone, the combination of A and B, the combination of A and C, the combination of B and C, or the combination of A, B, and C.

[0083] As defined herein, the term "automatically" means without human intervention. As defined herein, the term "user" means a human being.

[0084] As defined herein, the term "when" means "sometimes" or "when" or "in response to" or "in response to" depending on the context. Thus, the phrases "when it is determined that" or "when [the described condition or event] is detected" can be interpreted, depending on the context, as "when it is determined that" or "in response to determining that", or "when [the described condition or event] is detected" or "in response to detecting [the described condition or event]" or "in response to detecting [the described condition or event]".

[0085] As defined herein, the term "in response to" and similar phrases as described above, e.g., "when", "sometimes", or "when", mean to readily respond or react to an action or event. The response or reaction is performed automatically. Thus, when a second action is performed "in response to" a first action, there is a causal relationship between the occurrence of the first action and the occurrence of the second action. The term "in response to" indicates a causal relationship.

[0086] As defined herein, the term "output" means storing in a physical memory element, e.g., in a device, writing to a display or other peripheral output device, transmitting or sending to another system, exporting, etc.

[0087] As defined herein, the term "real time" means a level of processing responsiveness such that a user or system perceives that a particular process or decision is made sufficiently instantaneously, or enables a processor to keep up with some external process.

[0088] As defined herein, the term "substantially" means that the recited characteristics, parameters, or values need not be achieved exactly, but that deviations or variations, including for example tolerances, measurement errors, measurement precision limitations, and other factors known to those of skill in the art, may occur in amounts that do not preclude the effect that the characteristic is intended to provide.

[0089] The terms first, second, etc. may be used herein to describe various elements. These elements should not be limited by these terms. Because these terms are only used to distinguish one element from another, unless stated otherwise or the context clearly indicates otherwise.

[0090] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems and methods according to various aspects of the present disclosure. In some alternative implementations, the operations recited in the blocks may occur out of the order recited in the figures. For example, two blocks shown in succession may be executed substantially simultaneously or the blocks may sometimes be executed in the reverse order depending on the functionality involved. In other instances, the blocks may generally be executed in numerical ascending order, and in still other instances, one or more of the blocks may be executed in various orders, the results being stored and used in subsequent blocks or other blocks that do not immediately follow. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a dedicated hardware-based system that performs the specified functions or acts or combinations of dedicated hardware and computer instructions.

[0091] One or more exemplary implementations include an IC. The IC can include a DP array including a plurality of computing tiles arranged in a grid. The IC can include an array interface coupled to the DP array. The array interface includes a plurality of interface tiles. Each interface tile includes a plurality of DMA circuits. The IC can include a NoC coupled to the array interface. Each DMA circuit is communicatively linked to the NoC via an independent communication channel.

[0092] The foregoing and other implementations can each optionally include, alone or in combination, one or more of the following features. Some exemplary implementations include all of the following features in combination.

[0093] In one aspect, the plurality of DMA circuits of each of the plurality of interface tiles are configured to simultaneously transfer a plurality of data streams to the data processing array.

[0094] In another aspect, each interface style includes a stream switch coupled to a plurality of DMA circuits.

[0095] In another aspect, each stream switch is connected to at least one other stream switch disposed within an interface style adjacent to the west or an interface style adjacent to the east. Each stream switch is connected to another stream switch within an array tile adjacent to the north.

[0096] In another aspect, the DMA circuit is configured to convert a data stream received from the stream switch into a memory-mapped transaction to be transmitted to the NoC. The DMA circuit is configured to convert a memory-mapped transaction received from the NoC into a data stream to be transmitted to the stream switch.

[0097] In another aspect, each interface style includes a stream multiplexer-demultiplexer coupled to each DMA circuit of the interface style. Each stream multiplexer-demultiplexer is configured to selectively pass a data stream received from a plurality of direct memory access circuits to the stream switch and direct a stream received from the stream switch to a selected one of the plurality of direct memory access circuits.

[0098] In another aspect, each interface style includes a plurality of NoC stream interfaces corresponding to the plurality of DMA circuits. Each NoC stream interface is configured to receive a data stream from the NoC, transmit the data stream to the stream multiplexer-demultiplexer, and receive a data stream from the stream switch and transmit the data stream to the NoC.

[0099] In another aspect, a grid of a plurality of compute tiles includes a plurality of columns, each of the plurality of columns including one or more of the plurality of compute tiles. Each column of the grid includes a selected compute tile of one or more of the plurality of compute tiles having a stream switch connected to a stream switch of an interface tile aligned with the column.

[0100] In another aspect, a grid of a plurality of compute tiles includes a plurality of memory tiles. The grid includes a plurality of columns, each of the plurality of columns including one or more of the plurality of compute tiles and one or more of the plurality of memory tiles. Each column of the grid includes a selected memory tile of one or more of the plurality of memory tiles having a stream switch connected to a stream switch of an interface tile aligned with the column.

[0101] In another aspect, each direct memory access circuit is connected to a NoC master circuit of the NoC and configured to transmit transactions through the NoC via the NoC master circuit.

[0102] In another aspect, each interface tile includes a selector switch programmable to connect a selected DMA circuit within the same interface tile or a selected DMA circuit within an adjacent interface tile to a selected interface circuit of the plurality of interface circuits of the NoC.

[0103] In another aspect, the selector switch is programmable to connect different interface circuits of the plurality of interface circuits of the NoC to different interface tiles of the interface tiles.

[0104] In one or more exemplary implementations, an array interface for an IC can include multiple interface styles. Each interface style can include a stream switch connected to an adjacent compute tile of the DP array or another stream switch of an adjacent memory tile of the DP array. The stream switch is also connected to the stream switch of at least one adjacent interface style of the array interface. Each interface style can include a stream multiplexer-demultiplexer connected to the stream switch. Each interface style can include multiple DMA circuits. Each DMA circuit is connected to the stream multiplexer-demultiplexer and the NoC. The DMA circuit is configured to convert a data stream received from the stream multiplexer-demultiplexer into a memory-mapped transaction transmitted to the NoC, and convert a memory-mapped transaction received from the NoC into a data stream transmitted to the stream multiplexer-demultiplexer.

[0105] Each of the foregoing and other implementations can optionally include one or more of the following features, alone or in combination. Some exemplary implementations include all of the following features in combination.

[0106] In one aspect, each DMA circuit is communicatively linked to the NoC via an independent communication channel. Each interface style of the multiple interface styles provides multiple independent connections to the NoC.

[0107] In another aspect, each stream multiplexer-demultiplexer is configured to selectively pass a data stream received from multiple DMA circuits to the stream switch, and selectively pass a data direct stream received from the stream switch to a selected one of the multiple DMA circuits.

[0108] In another aspect, each interface style can include a plurality of NoC stream interfaces corresponding to a plurality of DMA circuits. Each NoC stream interface is configured to receive a data stream from the NoC and transmit the data stream to a stream multiplexer-demultiplexer. Each NoC stream interface is configured to receive a data stream from the stream multiplexer-demultiplexer and transmit the data stream to the NoC.

[0109] In another aspect, each DMA circuit is connected to a NoC master circuit of the NoC and is configured to transmit transactions through the NoC via the NoC master circuit.

[0110] In another aspect, each interface style includes a selector switch that is programmable to connect a selected DMA circuit within the same interface style or a selected DMA circuit within an adjacent interface style to a selected interface circuit among a plurality of interface circuits of the NoC.

[0111] In another aspect, the selector switch is programmable to connect different interface circuits of the NoC to different interface styles among the interface styles.

[0112] In another aspect, the plurality of DMA circuits are configured to simultaneously transmit a plurality of data streams to a data processing array via the NoC stream interfaces.

Claims

**Claim 1** An integrated circuit comprising: A data processing array including a plurality of computing tiles arranged in a grid; An array interface coupled to the data processing array, the array interface including a plurality of interface styles, each interface style including a plurality of direct memory access circuits; A network-on-chip (NoC) coupled to the array interface, wherein each direct memory access circuit is communicatively linked to the NoC via an independent communication channel. **Claim 2** The integrated circuit according to claim 1, wherein the plurality of direct memory access circuits of each of the plurality of interface styles are configured to simultaneously transmit a plurality of data streams to the data processing array. **Claim 3** The integrated circuit according to claim 1, wherein each interface style includes a stream switch coupled to the plurality of direct memory access circuits. **Claim 4** Each stream switch is connected to at least one other stream switch disposed within an interface style adjacent to the west or an interface style adjacent to the east, Each stream switch is connected to another stream switch within an array tile adjacent to the north. The integrated circuit according to claim 3. **Claim 5** The direct memory access circuit is configured to convert a data stream received from the stream switch into a memory-mapped transaction transmitted to the NoC, The direct memory access circuit is configured to convert a memory-mapped transaction received from the NoC into a data stream transmitted to the stream switch. The integrated circuit according to claim 3. **Claim 6** Each interface style includes a stream multiplexer-demultiplexer coupled to each direct memory access circuit of the interface style. Each stream multiplexer-demultiplexer is configured to selectively pass the data streams received from the plurality of direct memory access circuits to the stream switch, and direct the streams received from the stream switch to a selected one of the plurality of direct memory access circuits. The integrated circuit according to claim 5. **Claim 7** Each interface style includes a plurality of NoC stream interfaces corresponding to the plurality of direct memory access circuits. Each NoC stream interface is configured to receive a data stream from the NoC, transmit the data stream to the stream multiplexer-demultiplexer, receive a data stream from the stream switch, and transmit the data stream to the NoC. The integrated circuit according to claim 6. **Claim 8** The grid of the plurality of computing tiles includes a plurality of columns, each of the plurality of columns including one or more of the plurality of computing tiles. Each column of the grid includes a selected one of the one or more of the plurality of computing tiles having a stream switch connected to the stream switch of the interface style aligned with the column. The integrated circuit according to claim 3. **Claim 9** The grid of the plurality of computing tiles includes a plurality of memory tiles. The grid includes a plurality of columns, each of the plurality of columns including one or more of the plurality of computing tiles and one or more of the plurality of memory tiles. Each column of the grid includes a selected one of the one or more of the plurality of memory tiles having a stream switch connected to the stream switch of the interface style aligned with the column. The integrated circuit according to claim 3. **Claim 10** The integrated circuit according to any one of claims 1 to 9, wherein each direct memory access circuit is connected to a NoC master circuit of the NoC and is configured to transmit a transaction through the NoC via the NoC master circuit. **Claim 11** An integrated circuit according to any one of claims 1 to 9, comprising a selector switch programmable to connect a selected direct memory access circuit within the same interface style or a selected direct memory access circuit within an adjacent interface style to a selected interface circuit among the plurality of interface circuits of the NoC.

12. The integrated circuit according to claim 11, wherein the selector switch is programmable to connect different interface circuits among the plurality of interface circuits of the NoC to different interface styles among the interface styles.

13. An array interface for an integrated circuit, the array interface comprising: a plurality of interface styles, each interface style comprising: a stream switch connected to an adjacent computation tile of a data processing array or another stream switch of an adjacent memory tile of the data processing array, the stream switch also being connected to the stream switch of at least one adjacent interface style of the array interface; a stream multiplexer-demultiplexer connected to the stream switch; a plurality of direct memory access circuits, each direct memory access circuit being connected to the stream multiplexer-demultiplexer and a network on chip (NoC); The direct memory access circuit is configured to convert a data stream received from the stream multiplexer-demultiplexer into a memory-mapped transaction transmitted to the NoC, and convert a memory-mapped transaction received from the NoC into a data stream transmitted to the stream multiplexer-demultiplexer.

14. A selector switch, wherein each interface style is programmable to connect a selected direct memory access circuit within the same interface style or a selected direct memory access circuit within an adjacent interface style to a selected interface circuit among the plurality of interface circuits of the NoC. The array interface according to claim 13. **Claim 15** The array interface according to claim 14, wherein the selector switch is programmable to connect different interface circuits among the plurality of interface circuits of the NoC to different interface styles among the interface styles.